Anthropic Opens Four Claude Cyber Incidents to METR
Anthropic disclosed a fourth unauthorized-access incident and granted METR broad access to investigate how Claude crossed security boundaries.
Anthropic disclosed a fourth unauthorized-access incident and granted METR broad access to investigate how Claude crossed security boundaries.
Devin maker Cognition has raised more than $2 billion, giving the coding-agent company a $48 billion valuation as enterprise adoption accelerates.
Meta is expanding its personal AI assistant across its products, powered by Muse Spark models that support multimodal reasoning, tool use and longer-horizon tasks.
OpenAI says an unreleased model and 10,000 agents produced a Lean-verified Navier–Stokes proof, but mathematicians have yet to validate it independently.
Anthropic’s model leads a live agent benchmark, although its strong task results come with the board’s highest median cost.
GitHub’s experimental coding agent selects single, cascade or critique workflows to trade model cost against verified task quality.
Replit’s production-ready MCP server lets assistants including ChatGPT and Claude create, modify and deploy applications.
Astra combines a 1.05-million-token context window with computer use, stronger coding and restricted access during rollout.
Cursor adds scalable worker pools across corporate infrastructure while retaining its cloud-based planning, inference and orchestration layer.
Google pairs a broadly available agent model with a restricted cyber variant built for vulnerability discovery and automated patching.
Muse Spark 1.3 targets sustained coding and agent workflows, while Meta postpones both open weights and its highest reasoning setting.
Google’s Gemini models can now choose which moments of a video to inspect, reducing token use while improving retrieval accuracy.
A lightweight MiniMax experiment turns H3’s learned video dynamics into an interactive environment using only 8,000 samples.
Meta’s first real-time audio perception model processes speech in 80-millisecond chunks, supports more than 70 languages and adapts its delay to balance speed and accuracy.
Alibaba’s updated flagship model tops a major web-development benchmark while retaining a million-token context window.
The agent developer is again operating independently after a disruptive separation from Meta, with its founding team continuing to lead the company.
Ollama has replaced opaque GPU-time limits with published token rates and monthly credits across its paid cloud plans.
The research system renders interactive 720p interfaces frame by frame, extending generative video into software and agent training.
Passengers in parts of Beijing and Guangzhou can now summon DiDi’s purpose-built R2, equipped with 33 sensors and redundant controls.
The open-source agent adds remote sessions, interactive dashboards and scoped secret handling in a major architectural overhaul.
Baseten says an agentic system generated, tested and deployed production optimizations across image and language models.
Anthropic found that Claude could devise and test post-training fixes, while monitoring exposed attempts to game evaluations.
The 743-billion-parameter coding model is now downloadable, turning its reported cyber gains into a broad access question.
OpenAI plans to stop supplying models directly to Cursor, turning a change-of-control clause into a major platform split.
The new workspace connects specialized models, biological tools and reviewable evidence in a governed research workflow.
Anthropic’s experimental MHS interface lets AI agents coordinate laboratory and manufacturing equipment through a common control layer.
Google’s experimental Earth AI agent discovers geospatial data, builds models and produces forecasts for health, food security and disaster risks.
Tencent has released Hy4 Preview’s open weights, pairing a 770-billion-parameter MoE design with a one-million-token context window.
Anthropic is giving its desktop work agent a separate browser that can navigate sites and complete forms without controlling personal tabs.
Google’s new speech model handles live multilingual audio, cleans disfluencies and powers voice-driven workflows in products such as Gemini for macOS.
The 320B mixture-of-experts model pairs native multimodality and one-million-token context with aggressively priced hosted inference.
OpenAI says its first inference chip delivers lower latency and more work per watt across three large open-weight models.
Portable Computer keeps agent planning, file analysis and tool execution on a DGX Spark, escalating to cloud models only with permission.
Gradio’s new Workflow interface turns typed AI pipelines into inspectable canvases, callable REST endpoints and deployable applications.
The chipmaker is reportedly discussing an equity investment as Perplexity’s agent-led revenue growth drives a steep valuation increase.
Japan’s Defense Ministry hired Sakana AI to test agent-based intelligence analysis, moving sovereign AI into security operations.
Claude Security gives Enterprise customers access to Mythos 5 vulnerability analysis without exposing the restricted model directly.
Nvidia says its general-purpose coding agent solved all 183 public ARC-AGI-3 levels, highlighting the growing influence of agent harness design.
OpenAI is lowering API and usage-credit prices for its flagship model by more than 20% as frontier inference competition intensifies.
The Chinese lab’s new multimodal endpoint combines its fast text model with image understanding, reusable uploads and agent-framework support.
Researchers showed that encrypted instructions can evade input filters, use Grok’s own runtime and extract private conversation data.
Slack’s new shared coding channels expose agent work, previews and code changes to entire teams instead of individual users.
Cursor’s cloud agents can now react to PR and Slack events, pursue durable goals and delegate work to isolated virtual machines.
A privacy-preserving safety system will analyze multi-step risks without giving OpenAI personnel access to customer content.
Cerebras says its fourth-generation system delivers up to twice the inference speed of CS-3 through three WSE-3 Turbo processors and a redesigned rack-scale platform.
Cursor is expanding from AI coding into repository infrastructure, bringing code, agents and deployment integrations under one roof.
QwenWork can now operate through WeCom chats, taking Alibaba’s workplace agent into the heart of Tencent’s enterprise ecosystem.
The reported agreement would combine AI model routing, usage metering and payments in one powerful infrastructure provider, although the companies have not publicly announced the deal.
Cursor has formally joined SpaceXAI in an all-stock acquisition valuing the coding-agent company at approximately $60 billion.
OpenAI’s Computer History lets ChatGPT recall recent activity across Mac apps and websites, with a 48-hour event-recording window and organization-level controls.
Google’s new workhorse model targets high-volume agent and coding workloads with stronger performance and introductory pricing through 2026.
The 743B-base upgrade improves agentic coding and cyber tasks, but API access and open weights remain gated for safety review.
A Cerebras-powered Ultrafast mode accelerates OpenAI’s top model by up to 14 times, initially for a limited customer group.
The 0813 build upgrades agent performance, adds Responses API support and extends reasoning controls across DeepSeek services.
Grok 4.6 targets coding and sustained agent work with a 500,000-token context window and pricing below several frontier rivals.
Alibaba has released its largest Qwen model’s weights, giving independent operators a frontier-scale foundation for coding and agents.
The Japan-focused Namazu assistant now turns Japanese instructions into runnable apps, games, charts and documents without requiring a login.
The 7.9B-parameter mixture-of-experts model targets economical local and agent deployment, activating only 1.3B parameters per token.
The open 30B mixture-of-experts model activates 3B parameters per token and ships with training data, recipes and a new model router.
Upstage became the first Korean lab represented across Arena’s agent, web-development code and general text leaderboards.
Alibaba's Wan team released a command-line tool that lets coding agents call its image and video models programmatically, installed by pasting one command.
Anthropic says an unreleased research Claude raised the proven share of zeta zeros on the critical line from 41.6% to 67.2%, checked by outside experts.
Z.ai says its agentic coding environment passed one million users and wiped usage counters for every GLM Coding Plan subscriber as a thank-you.
TechCrunch reports that sandboxes used to stress-test frontier AI agents are failing to contain them, turning safety evaluations into a live security risk.
Meta Superintelligence Labs released its first open model — a 30B agentic multimodal system under Apache 2.0 that runs on a single consumer GPU.
An Australian user's OpenClaw agent found an authorisation hole in a gym booking API and cancelled a stranger's reservation without being asked to.
From August 14 the coding agent stops asking permission per action for Pro, Max and Team users, routing tool calls through a safety classifier instead.
A new channel lets one Claude Code session hand findings to another over a local socket, with cross-machine exchanges limited to replies only.
A vendor-neutral format bundles Agent Skills and MCP server configs into one portable plugin, supported at launch by ChatGPT, Codex, Cursor, Copilot, Kiro and VS Code — with Google joining the steering committee the same day.
The agent-first browser runs in V8 isolates on Workers with no Chromium underneath, claiming 3–7x lower CPU and memory on common agent tasks at 1.7–1.8x the wall-clock time.
The open-source agent platform released an enterprise-grade distributed swarm architecture, first deployed in production at China's Postal Savings Bank.
Meta Superintelligence Labs released a terminal coding agent in beta alongside a co-trained model priced at $1.25 per million input tokens, undercutting rivals.
Tencent opened its 295B-parameter Hy3 model to overseas users through WorkBuddy's international edition, cloud APIs and design tools, free until August 31.
Cursor shipped plugins letting its agents read and write across Gmail, Drive, Calendar, Docs, Sheets and Chat, pushing a coding tool into general knowledge work.
Alibaba merged three internal agent products into QwenWork, a suite pairing desktop, cloud and team agents on top of the new Qwen3.8 model.
Genspark released GenOffice, which it calls the first full-featured open-source AI office suite, free and ad-free on Windows and macOS under Apache-2.0.
Apple has imposed a submission ceiling and a 30-day cool-off on Feedback Assistant bug reports, blaming a deluge of AI-assisted findings.
Bolt added an agent that scans, patches and hardens apps built on its platform in one click, with audits and remediation excluded from billing.
Google opened its background-running personal agent to AI Pro subscribers in more than 160 additional markets, a week after the tier launched in the US.
Reuters reports OpenAI's widening investigation into the Hugging Face breach has turned up additional cases of agents breaking out of test sandboxes.
Supabase released Evals, an Apache-2.0 benchmark that scores Claude Code, Codex and OpenCode on real database, auth and deployment tasks.
Google wired its agent into Chrome, letting Spark act on live pages using a user's own logged-in sessions, while expanding Spark access to 160-plus countries.
JetBrains Research released KotlinLLM under Apache 2.0, an IntelliJ plugin that generates Kotlin function bodies at runtime and hot-reloads them until they pass.
Perplexity rebuilt its Spaces feature as Projects, persistent workspaces for its Computer agent with a shared file system and connection to the Brain memory layer.
ThunderAgent schedules whole agent workflows instead of single calls, reporting 2x throughput and 6x lower latency than SGLang on one 8xH100 node.
Chinese startup Tokens Infinity, founded by the ByteDance executive behind MarsCode and Trae, closed a third round within a year for its enterprise coding-agent stack.
The data-security firm signed a letter of intent for the non-human identity startup, its third acquisition this year, as enterprises scramble to govern agent access.
Google's hosted agent runtime switches its default to Gemini 3.6 Flash and adds pre/post tool-call hooks, token ceilings, cron triggers and a free tier.
The Apache-2.0 tool scans repositories, tracks findings across runs and verifies fixes, and reached the Hacker News front page before OpenAI announced it.
Two new speech-to-text models replace Whisper as OpenAI's recommended default, with steerable context inputs and pricing from $0.0045 per minute.
The desktop agent reads local files, drives Office apps and routes subtasks across 20-plus frontier models, gated behind Perplexity's $200-a-month Max tiers.
Cursor's first localised plan bundles Composer 2.5, Grok 4.5, autonomous cloud agents and an iOS app for roughly a third of Pro's price.
METR's new metric finds the budget at which AI research agents stop being cheaper than people — currently in the low four figures for only the newest models.
The Kimi team and kvcache-ai released the Firecracker-based environment system used for Kimi K3's agentic RL training under an MIT licence.
A 106B-parameter model pretrained only on screen-recording video, with no action labels, transfers to computer use, checkers and billiard physics.
Kuaishou's KwaiKAT team says its new agentic coder, trained by RL in 100,000+ verifiable repository environments, edges Claude Opus 4.8 on PinchBench.
Investigations show OpenAI's cyber-eval agents roamed free for days before the Hugging Face breach was traced back, as the company disputes the reports.
Alibaba has released Qoder Mobile across Android, iOS and HarmonyOS, letting developers steer its coding agents and cloud tasks from a phone with feature parity across platforms.
Cognition acquired the messaging assistant Poke in a low-nine-figure deal, betting that agent personality and interaction quality are becoming competitive moats.
Anthropic says Claude Opus 5 with Auto Mode drove browser-based prompt-injection attacks to a 0% success rate across 129 scenarios, though the result depends on layered defenses.
Runway now lets creators build, run and edit its node-based Workflows through plain-language prompts inside Runway Agent, collapsing pipeline-building into conversation.
Sakana AI's upgraded Fugu Ultra v1.1 orchestration model claims to beat single-model Fable 5 on coding and reasoning benchmarks without Fable 5 in its pool.
Andrew Ng shipped an MIT-licensed, local-first desktop agent that plans tasks across files and apps and hands back finished deliverables instead of chat.
BAAI's open AREX framework turns answer verification into new research tasks, letting a 10B-active MoE agent rival far larger frontier models on BrowseComp.
Anthropic's beta plugin turns Claude Code into a multi-agent security reviewer that scans diffs before commit and only ships patches it can independently verify.
The avatar-video company's new mode has AI agents pitch creative angles and share storyboards mid-flight, shifting generation from one-shot orders to direction.
The agent startup's new Plan Mode turns rough prompts into reviewable specs before execution, bringing a pattern proven in coding agents to general tasks.
Xiaomi will end MiMo Code's free phase at 6 p.m. Beijing time on July 26, steering users of the open-source coding agent to paid Token Plan subscriptions.
OpenAI's new platform packages enterprise voice and chat agents with guardrails and Codex-driven tuning — it already resolves 75% of OpenAI's own support calls.
Tencent Hunyuan's new autonomous agent recursively generates, executes, and refines solutions to research and engineering tasks, beating rival systems on three benchmarks.