An independent audit of four frontier models found hundreds of universal jailbreaks in Grok 4.5 and Gemini 3.1 Pro, and none in Claude Fable 5 or GPT-5.6 Sol.
The first reproducible cross-lab audit of frontier safeguards puts a dollar price on breaking each model — and two of the four leading systems fail cheaply.
ThunderAgent schedules whole agent workflows instead of single calls, reporting 2x throughput and 6x lower latency than SGLang on one 8xH100 node.
⚡ Quick hits
AIGCGoogle's Gemini app detailed four Omni-powered video editing controls: instructions read from a reference video, style transfer, relighting, and background replacement.X @GeminiApp ↗
OpenAI reports that after deployment it turned GPT-5.6 Sol on its own inference stack, autonomously rewriting production kernels for a 20% end-to-end cost cut.
⚡ Quick hits
IndustryThinking Machines co-founder Lilian Weng departed citing health reasons and has since joined OpenAI, the company she previously led safety research at.TechCrunch ↗
Open SourceMiniMax and Fireworks AI open-sourced the attention kernels behind fast M3 inference, publishing both the MiniMax MSA and Fireworks kernel repositories.X @MiniMax_AI ↗
AIGCRecraft began testing a design agent inside Recraft Studio that turns a product description into a consistent asset set—logo, palette, icons and more.X @recraftai ↗
ResearchIn Vending-Bench testing, Claude Opus 5 ran a simulated vending machine business with tactics researchers described as ruthless profit maximization.TechCrunch ↗
AgentsCursor shipped an iPad app and added an inbox plus full pull-request review with comments, checks and approvals on iPhone and iPad.X @cursor_ai ↗
ModelsArtificial Analysis says xAI's Grok Voice Think Fast 2.0 High tops its Tau Voice agentic benchmark at 56.5%, but costs $4.80 per input audio hour.X @ArtificialAnlys ↗
AgentsVS Code's new release previews subagent progress in the Agents window, adds built-in dictation across chat, editors and terminal, and a hybrid Markdown editor.X @code ↗
Ideogram shipped a dedicated object-removal model that erases shadows and reflections too, and says it beats Nano Banana 2, FLUX Erase and GPT Image-2.
ChatGPT for Academic Researchers starts with 10,000 scientists and scales to 100,000 by 2027, part of a $250M-plus commitment to external research.
⚡ Quick hits
AIGCHeyGen launched Video Podcast, turning any document, link, or idea into a two-host video show with studio scenes, multi-camera cuts, and B-roll.X @HeyGen ↗
Google launched Lyria 3.5 in Flow Music with Selective Section Painting, letting users regenerate one part of a track or grow a short idea into a full song.
⚡ Quick hits
SafetyBAAI and Peking University tested 11 commercial large models and found all of them able to produce split synthesis routes that evade biosecurity screening.InfoQ 中文 ↗
FundingEncore AI raised $30 million to build customer-facing AI agents that continuously learn from recorded sales and support calls.TechCrunch ↗
ResearchSakana AI and NYU released Dream-Cubed, a controllable generative world model for Minecraft trained on billions of cubic-meter blocks of gameplay data.X @SakanaAILabs ↗
Open SourceUnsloth shrank Kimi K3 from 1.56TB to 594GB with 1-bit quantization, retaining about 78.9% accuracy and running on a 128GB Mac Studio.X @UnslothAI ↗
Open SourceTencent Hunyuan open-sourced AngelSpec, an end-to-end speculative decoding framework for training and deployment, reporting 1.98–2.40x end-to-end speedups on Hy3-A21B.X @TencentHunyuan ↗
The Chinese security vendor reports its agent reproduced 1,301 of 1,507 vulnerabilities on CyberGym, claiming fourth place globally atop open GLM-5.2 weights.
Seoul's benchmark fell about 6% on Wednesday after a 10.84% Tuesday collapse, as SK Hynix missed estimates and China's first domestic DUV tools entered service.
The Apache-2.0 tool scans repositories, tracks findings across runs and verifies fixes, and reached the Hacker News front page before OpenAI announced it.
The data-security firm signed a letter of intent for the non-human identity startup, its third acquisition this year, as enterprises scramble to govern agent access.
⚡ Quick hits
PolicyThe UK's CMA opened an investigation into whether Microsoft misled Microsoft 365 Personal and Family customers when it bundled Copilot and rolled them onto pricier plans.GOV.UK (CMA) ↗