August 11, 2026
Quick hits from across the AI world in the last 24 hours.
- ⚡ Quick hits
- Industry Mistral introduced EU data processing and paid priority access, although regional routing excludes several features and priority availability is not guaranteed. The Decoder ↗
- Agents Tencent Hunyuan presented research evaluating whether self-modifying AI agents genuinely improve, targeting more reliable measurement of self-evolution. Tencent Hunyuan on X ↗
- AIGC Xiaomi's MiLM Plus introduced PROVE, combining two perception-aligned object-removal metrics with a real-world video benchmark. MarkTechPost ↗
- Agents AEROBAT automated agent-behavior research, generating 79 hypotheses and executing 23,512 simulation rounds across 12 target behaviors. arXiv ↗
- Agents MESA improved long-horizon agent-memory performance by 8.5% while using 41% fewer evidence tokens than an all-structure approach. arXiv ↗
- Research FACT trains world-action models on failed actions, reducing success-biased future hallucinations and improving simulated and real-world bimanual manipulation. arXiv ↗
- Safety MarkNull reduced AI-image watermark bit accuracy to 53.14% without perceptible degradation and also compromised Google's SynthID-Image. arXiv ↗
- Safety A deterministic streaming guardrail withheld every chunk completing predefined dangerous lexical pairs, but researchers stressed its coverage remains deliberately narrow. arXiv ↗
- Safety Researchers catalogued LLM-mediated variants of eight classic web attacks, finding vulnerability depends on both application architecture and model behavior. arXiv ↗
- Research An empirical study examined when chain-of-thought improves reasoning and when sequential processing depth instead becomes a performance bottleneck. arXiv ↗
- Agents Researchers proposed closed-loop LLM copilots that combine agricultural sensor data, domain knowledge and feedback-driven recommendations for digital farming. arXiv ↗
- AIGC Qwen made its Qwen-Image-3.0 image generator available for testing through OpenArt. Alibaba Qwen on X ↗
- Models Qwen said Qwen3.8-Max climbed from 22nd to fourth place on the Legal Research Bench leaderboard. Alibaba Qwen on X ↗
- Open Source Ollama can now serve as a local model provider for GitHub Copilot inside JetBrains development environments. Ollama on X ↗
- Agents Artificial Analysis found GPT-5.5 xhigh led its new AnalystAgent benchmark in single-run reliability with a 66% pass rate. Artificial Analysis on X ↗
- AIGC Ideogram launched Ad Resizer, converting one uploaded design into campaign-ready assets for social, video and connected-TV placements. Ideogram on X ↗
- AIGC Runway added Seedance 2.5, supporting 50 character references and music-synchronized video clips lasting up to 30 seconds. Runway on X ↗
- Agents Replit expanded its MCP beta so external agents can create, load, search and publish Replit applications directly. Replit on X ↗
- Funding General Catalyst led a $1.1 billion funding round for River AI, a startup founded only two months earlier. TechCrunch ↗
- Agents OpenAI released a Linux preview of its ChatGPT desktop app, bringing ChatGPT Work and Codex to several major distributions. OpenAI on X ↗
- Research Google demonstrated AMIE conducting real-time audio-visual clinical consultations in a first-of-its-kind research study. Google Research ↗
- AIGC Krea launched a Slack beta combining multiple generative models with an agent that adapts to a creative team’s preferences. Krea on X ↗
- AIGC Luma introduced Scenes, letting creators refine individual scenes without regenerating an entire AI-produced advertisement. Luma AI on X ↗
- Industry ElectricSQL, creator of browser-embeddable Postgres engine PGlite, is joining Databricks to bring WASM Postgres into AI-agent sandboxes. Databricks on X ↗
- Open Source Unsloth released an open-source desktop application for running and training models locally across macOS, Windows and Linux. Unsloth on X ↗
- Industry Manus said it will resume operating independently and announced transition measures required by regulations in specified regions. Manus on X ↗
- Agents LlamaIndex introduced ExtractBench, a benchmark for measuring information extraction from complex enterprise documents. LlamaIndex on X ↗
- Policy Spotify will label AI Persona artist profiles and exclude their tracks from algorithmic recommendations, while leaving them searchable. TechCrunch ↗