IndustryWalmart store employees are gaining a new responsibility: identifying and correcting mistakes made by the retailer’s AI agents.IT之家 ↗
IndustryAnthropic models cost 4.4 times the provider average per token on Vercel, yet developers continue choosing them.The Decoder ↗
AgentsClaude Code added a /design command that generates and iterates on interface mockups directly inside the terminal.The Decoder ↗
ResearchResearchers still lack sufficiently representative data to determine how people actually use AI across populations and everyday contexts.MIT Technology Review ↗
AIGCRecraft showcased V4.1 generating poster typography whose distressed lettering, split effects, and lighting are integrated into the composition.Recraft on X ↗
ResearchAnalysis suggests recursive AI self-improvement may progress slower than expected because experiments, evaluation, and real-world feedback remain bottlenecks.MIT Technology Review ↗
SafetyTests found AI agents can silently lose user requirements when compressing long contexts, causing later actions to diverge from instructions.The Decoder ↗
AgentsA six-agent system demonstrated an automated game-development loop spanning generation, playtesting, bug detection, and iterative repair.量子位 ↗
QwenWork can now operate through WeCom chats, taking Alibaba’s workplace agent into the heart of Tencent’s enterprise ecosystem.
⚡ Quick hits
ModelsSakana AI made its Japanese business-focused Namazu reasoning model available through OpenRouter with web search and code execution.Sakana AI on X ↗
AIGCA Physical AI Hackathon project used Hyper3D Rodin to create assets for a user-imagined digital board game.Deemos Tech on X ↗
IndustryBaseten disclosed that its infrastructure powers model inference behind Notion’s AI-enabled product experiences.Baseten on X ↗
ResearchDan Luu argued that widespread benchmark contamination and optimization are making headline AI scores increasingly unreliable indicators of practical capability.Dan Luu ↗
IndustryGoogle reportedly paid $10 million for access to Spirit Airlines’ business data for AI training.SiliconANGLE ↗
AgentsArena added task categories and cost-to-completion data to its Agent Leaderboard, enabling performance-and-price comparisons.Arena on X ↗
IndustryAI workflow startup Relay is shutting down, with its staff joining Google’s Chrome team.TechCrunch ↗
IndustryFormer SpaceX engineers are building an automated factory designed to manufacture steel components with robotics.Ars Technica ↗
AIGCDreamina Seedance 2.5 reached first place for video editing on Arena’s crowdsourced Video Arena leaderboard.Arena on X ↗
AgentsReplit introduced black-box penetration tests that probe deployed apps externally, with Replit Agent able to address detected vulnerabilities.Replit on X ↗
ModelsNvidia’s Nemotron 3.5 Lightning became available through Amazon SageMaker JumpStart for streamlined model deployment.AWS Machine Learning Blog ↗
AIGCCartesia’s Sonic 3.6 took first place in both Provider Voice and Controlled Voice on Artificial Analysis’s Speech Arena.Artificial Analysis on X ↗
AIGCCartesia’s Sonic 3.6 generated 136.1 characters per second and cost $49 per million characters in Artificial Analysis testing.Artificial Analysis on X ↗
Cursor is expanding from AI coding into repository infrastructure, bringing code, agents and deployment integrations under one roof.
⚡ Quick hits
AgentsAWS published a reference workflow enabling OpenClaw agents to conduct transactions through Amazon Bedrock AgentCore’s payment capabilities.AWS Machine Learning Blog ↗
ResearchLlamaIndex said ExtractBench only credits extracted fields when both the value and its source citation are correct.LlamaIndex on X ↗
The generative-media company has secured a major round to scale its video and image platform amid intensifying AIGC competition.
⚡ Quick hits
ResearchOpenAI said retained reasoning and compaction raised GPT-5.6 Sol’s ARC-AGI-3 score from 13.3% to 38.3% while using roughly sixfold fewer output tokens.OpenAI Developers on X ↗
Groq’s latest official financing was a $650 million round co-led by Disruptive and Infinitum, supporting expansion of its global inference cloud.
⚡ Quick hits
AgentsElevenLabs launched a Claude-compatible MCP integration for securely managing voice and conversational agents without running a separate server.ElevenLabs on X ↗
FundingVoice-input startup Wispr raised $280 million at a $2 billion valuation as it expands beyond dictation.TechCrunch ↗
FundingPalona raised $20 million to develop AI automation products for physical retail and other brick-and-mortar businesses.SiliconANGLE ↗