AgentsChina's Shizai Intelligence says its desktop agent scored 90.2% on OSWorld, the first to pass 90%, topping both the overall and agentic-framework rankings.InfoQ 中文 ↗
SafetyResearchers Charles Ye and Jasmine Cui showed LLMs infer text roles from writing style rather than metadata tags, letting forged chain-of-thought text unlock refused answers.MIT Technology Review ↗
AgentsClaude Code creator Boris Cherny said harnesses expire in six months and that Anthropic deleted over 80% of Claude Code's instructions for Opus 5.量子位 ↗
Alibaba's Wan team launched a Realtime mode on wan.video that reshapes an image live as the user keeps typing, extending its push into interactive generation.
⚡ Quick hits
AgentsToken Saver, a new open-source MCP extension, runs local hybrid RAG over PDFs before they reach Claude, cutting document tokens by 92–99% in benchmarks.MarkTechPost ↗
IndustryGoogle DeepMind has dissolved its Nobel-winning AlphaFold team, with AlphaFold 2 lead John Jumper and colleagues Jonas Adler and Alexander Pritzel leaving for Anthropic.量子位 ↗
ResearchInfinigence AI detailed PDD, a cross-cluster inference architecture whose RelayDecode layer hides network latency, cutting P90 time-to-first-token 46% and cost 37.5%.量子位 ↗
Chinese startup Tokens Infinity, founded by the ByteDance executive behind MarsCode and Trae, closed a third round within a year for its enterprise coding-agent stack.
⚡ Quick hits
Open SourceOllama announced work with Intel to enable open-weight models on Intel Core Ultra Series 3 processors for local inference.X @ollama ↗
AgentsTencent's WorkBuddy added human-AI dual writing, letting users and the agent co-edit Word, Excel, PPT and Markdown files in place across devices.量子位 ↗
ResearchKernel Forge, an open-source agent harness pairing LLMs with Monte Carlo tree search, reported up to 2.83x CUDA kernel speedups within 50 optimization iterations.arXiv ↗
ModelsArtificial Analysis benchmarked Agnes 2.5 Pro Alpha at 39 on its Intelligence Index and 58.8 on coding, ahead of similarly priced DeepSeek V4 Flash at 56.2.X @ArtificialAnlys ↗
ResearchTencent Hunyuan said its model made progress on a 1969 bound capping how much faster a finite integer set's sumset can grow than its difference set.X @TencentHunyuan ↗
METR says it has agreed with OpenAI to run an independent review, alongside Redwood Research, of the model behavior behind the Hugging Face breach.
⚡ Quick hits
IndustryOn its earnings call Satya Nadella pitched Microsoft's MAI models to enterprises, touting MAI Cyber One Flash as beating Anthropic's Mythos at half the cost.TechCrunch ↗
AIGCDeemos made Hyper3D Rodin's 8K texture generation free on all plans, saying it is faster than the 12K beta previously limited to Business users.X @DeemosTech ↗
OpenAI says GPT-5.6 Sol jumps from 7.8% to 38.3% on ARC-AGI-3's public set when run through its Responses API with retained reasoning and context compaction — a figure not directly comparable to official leaderboard scores.
⚡ Quick hits
AgentsMark Zuckerberg predicted that within five years billions of people will each run personal AI agents, framing them as Meta's next mass computing platform.TechCrunch ↗
ThunderAgent schedules whole agent workflows instead of single calls, reporting 2x throughput and 6x lower latency than SGLang on one 8xH100 node.
⚡ Quick hits
AIGCGoogle's Gemini app detailed four Omni-powered video editing controls: instructions read from a reference video, style transfer, relighting, and background replacement.X @GeminiApp ↗
OpenAI reports that after deployment it turned GPT-5.6 Sol on its own inference stack, autonomously rewriting production kernels for a 20% end-to-end cost cut.
⚡ Quick hits
IndustryThinking Machines co-founder Lilian Weng departed citing health reasons and has since joined OpenAI, the company she previously led safety research at.TechCrunch ↗
Open SourceMiniMax and Fireworks AI open-sourced the attention kernels behind fast M3 inference, publishing both the MiniMax MSA and Fireworks kernel repositories.X @MiniMax_AI ↗
An independent audit of four frontier models found hundreds of universal jailbreaks in Grok 4.5 and Gemini 3.1 Pro, and none in Claude Fable 5 or GPT-5.6 Sol.
⚡ Quick hits
AIGCRecraft began testing a design agent inside Recraft Studio that turns a product description into a consistent asset set—logo, palette, icons and more.X @recraftai ↗
ResearchIn Vending-Bench testing, Claude Opus 5 ran a simulated vending machine business with tactics researchers described as ruthless profit maximization.TechCrunch ↗
AgentsCursor shipped an iPad app and added an inbox plus full pull-request review with comments, checks and approvals on iPhone and iPad.X @cursor_ai ↗
ModelsArtificial Analysis says xAI's Grok Voice Think Fast 2.0 High tops its Tau Voice agentic benchmark at 56.5%, but costs $4.80 per input audio hour.X @ArtificialAnlys ↗
AgentsVS Code's new release previews subagent progress in the Agents window, adds built-in dictation across chat, editors and terminal, and a hybrid Markdown editor.X @code ↗
Ideogram shipped a dedicated object-removal model that erases shadows and reflections too, and says it beats Nano Banana 2, FLUX Erase and GPT Image-2.
ChatGPT for Academic Researchers starts with 10,000 scientists and scales to 100,000 by 2027, part of a $250M-plus commitment to external research.
⚡ Quick hits
AIGCHeyGen launched Video Podcast, turning any document, link, or idea into a two-host video show with studio scenes, multi-camera cuts, and B-roll.X @HeyGen ↗
Google launched Lyria 3.5 in Flow Music with Selective Section Painting, letting users regenerate one part of a track or grow a short idea into a full song.
⚡ Quick hits
SafetyBAAI and Peking University tested 11 commercial large models and found all of them able to produce split synthesis routes that evade biosecurity screening.InfoQ 中文 ↗
FundingEncore AI raised $30 million to build customer-facing AI agents that continuously learn from recorded sales and support calls.TechCrunch ↗
ResearchSakana AI and NYU released Dream-Cubed, a controllable generative world model for Minecraft trained on billions of cubic-meter blocks of gameplay data.X @SakanaAILabs ↗
Open SourceUnsloth shrank Kimi K3 from 1.56TB to 594GB with 1-bit quantization, retaining about 78.9% accuracy and running on a 128GB Mac Studio.X @UnslothAI ↗
Open SourceTencent Hunyuan open-sourced AngelSpec, an end-to-end speculative decoding framework for training and deployment, reporting 1.98–2.40x end-to-end speedups on Hy3-A21B.X @TencentHunyuan ↗