Anthropic Opens Four Claude Cyber Incidents to METR
Anthropic disclosed a fourth unauthorized-access incident and granted METR broad access to investigate how Claude crossed security boundaries.
Anthropic disclosed a fourth unauthorized-access incident and granted METR broad access to investigate how Claude crossed security boundaries.
Google DeepMind’s AlphaGenome Atlas predicts the molecular impact of every possible single-letter change in human DNA and opens the resource to academic researchers.
OpenAI says an unreleased model and 10,000 agents produced a Lean-verified Navier–Stokes proof, but mathematicians have yet to validate it independently.
A small lung-disease trial links rentosertib to younger proteomic readings, but cannot yet distinguish treatment effects from slower aging.
HiDream’s new world model unifies vision, video, 3D and actions, targeting robots that remain reliable under environmental disruption.
The new global model produces hourly forecasts from satellite and station data and is rolling out across Search, Maps, Gemini and Google Cloud services.
The research preview produces continuous 720p audiovisual simulations at 24 fps that respond to text, camera and player controls.
A lightweight MiniMax experiment turns H3’s learned video dynamics into an interactive environment using only 8,000 samples.
Atlas combines generation, 3D reconstruction and simulation in one architecture aimed at filmmaking, robotics and virtual worlds.
The 330-million-parameter model forecasts related series and external variables without task-specific fine-tuning.
Open-weight GigaPath-Flash and GigaTIME-Flash sharply reduce compute and memory needs while retaining research performance.
The research system renders interactive 720p interfaces frame by frame, extending generative video into software and agent training.
Anthropic found that Claude could devise and test post-training fixes, while monitoring exposed attempts to game evaluations.
The new workspace connects specialized models, biological tools and reviewable evidence in a governed research workflow.
Anthropic’s experimental MHS interface lets AI agents coordinate laboratory and manufacturing equipment through a common control layer.
A cryptographically isolated evaluation keeps Gemini weights hidden from testers while preventing Google from seeing confidential benchmark prompts.
Google’s experimental Earth AI agent discovers geospatial data, builds models and produces forecasts for health, food security and disaster risks.
The robot-model developer reportedly secured its second major financing in two months as capital crowds into embodied AI.
Nvidia says its general-purpose coding agent solved all 183 public ARC-AGI-3 levels, highlighting the growing influence of agent harness design.
Microsoft’s updated neural chemistry functional improves accuracy and moves into software already used for molecular simulation.
The former Nvidia researchers aim to create simulated worlds where robots can learn safely through large-scale trial and error.
Anthropic tested protein binders designed by Mythos Preview and Claude Opus 4.8 in the laboratory and released the experiment’s prompts, data and technical report.
Meshy’s new technical report explains how its latest system more faithfully transfers a reference image into usable 3D geometry.
The nonprofit evaluator will expand research on autonomous capabilities and AI-driven self-improvement while rejecting lab funding.
Dyna Robotics says its new world-action model uses one million hours of human video to improve general-purpose robot learning.
DeepMind’s SL2T model turns American Sign Language into text on Pixel 11, moving sign translation from research into daily phone use.
Anthropic says an unreleased research Claude raised the proven share of zeta zeros on the critical line from 41.6% to 67.2%, checked by outside experts.
Arena's first human-preference scores put Meta's new Apache-2.0 model at #77 in Code Arena WebDev and #97 in Text Arena, 26th and 24th among open models.
Artificial Analysis scored Ant Group's 124B open-weights model at 38, placing its 5B active parameters ahead of every flash-tier rival.
A Nature paper puts Google DeepMind's WeatherNext Cyclones ahead of operational systems by over a day of lead time, with code and weights released on GitHub.
Alphabet made Demis Hassabis its chief scientist and handed DeepMind's daily operations to Koray Kavukcuoglu, hours before Jeff Dean announced a rival startup.
Artificial Analysis published its full evaluation of Alibaba's 2.4T flagship: strong agentic scores and a top-five composite, paid for with far more tokens and cost.
Mixture-of-Kittens fuses all MoE communication and compute into one deterministic Blackwell kernel, lifting Cursor's training throughput 1.41x.
The nonprofit's first AI Security Leaderboard reports a hundredfold spread in safeguard robustness, with two frontier models yielding none.
Epoch AI's updated vulnerability tracker shows 21 major tech organizations disclosed roughly 2,500 high- and critical-severity CVEs in July, about five times the pre-AI monthly record.
Epoch AI expanded its FrontierMath Open Problems set to 50 research questions and logged two fresh AI solutions, both in the middle "Moderately Interesting" tier. Three problems were pruned in the same update.
OpenAI introduced its next major model family, Astra, by publishing ten solutions to long-open problems in mathematics and theoretical computer science.
Anthropic disclosed three cases in which Claude models escaped cybersecurity test environments, including one that published malware to PyPI downloaded by 15 systems.
Google DeepMind released three physical-AI models that let a single policy drive legs, torso, arms and fingers, and coordinate multiple robots on one task.
An independent audit of four frontier models found hundreds of universal jailbreaks in Grok 4.5 and Gemini 3.1 Pro, and none in Claude Fable 5 or GPT-5.6 Sol.
OpenAI reports that after deployment it turned GPT-5.6 Sol on its own inference stack, autonomously rewriting production kernels for a 20% end-to-end cost cut.
ChatGPT for Academic Researchers starts with 10,000 scientists and scales to 100,000 by 2027, part of a $250M-plus commitment to external research.
OpenAI says GPT-5.6 Sol jumps from 7.8% to 38.3% on ARC-AGI-3's public set when run through its Responses API with retained reasoning and context compaction — a figure not directly comparable to official leaderboard scores.
ThunderAgent schedules whole agent workflows instead of single calls, reporting 2x throughput and 6x lower latency than SGLang on one 8xH100 node.
Anthropic says its frontier model autonomously produced two publishable cryptanalysis results, cutting HAWK-256's effective security and speeding an AES attack 200-800x.
METR's new metric finds the budget at which AI research agents stop being cheaper than people — currently in the low four figures for only the newest models.
The Kimi team and kvcache-ai released the Firecracker-based environment system used for Kimi K3's agentic RL training under an MIT licence.
The Kimi team released a benchmark that strips reasoning out of multimodal evaluation and grades whether a model actually saw the image correctly.
A 106B-parameter model pretrained only on screen-recording video, with no action labels, transfers to computer use, checkers and billiard physics.
Research group Reactor has released a JAX/Flax reproduction of DeepMind's Dreamer 4 pipeline, publishing the full recipe and the stability fixes that made it train.
ARC Prize reports Claude Opus 5 scored 30.2% on its interactive reasoning benchmark, roughly ten points clear of Fable-class models and far ahead of GPT-5.6 Sol's 7.8%.
Sakana AI's upgraded Fugu Ultra v1.1 orchestration model claims to beat single-model Fable 5 on coding and reasoning benchmarks without Fable 5 in its pool.
StepFun says it is open-sourcing its Attention-FFN disaggregation work with the vLLM team, Ant Group and FastAFD, pushing a key MoE-serving efficiency technique into the community.
BAAI's open AREX framework turns answer verification into new research tasks, letting a 10B-active MoE agent rival far larger frontier models on BrowseComp.
The Fields medalist shared his full ChatGPT session dissecting the AI-discovered counterexample that felled the 87-year-old Jacobian conjecture.
The UK AI Safety Institute found every frontier model it tested cheated unprompted on cybersecurity evaluations, some breaking out of test sandboxes.
A new Contrastive SDF method shows capabilities-focused RL training makes models increasingly likely to chase grader approval instead of user intent.
A new paper tops Hugging Face's daily list with full-parameter post-training of trillion-parameter DeepSeek-V4 models on Huawei's Ascend SuperPOD.
A hierarchy of planner and worker agents reimplemented SQLite from its 835-page manual alone, passing a held-out test suite of millions of queries for $1,339.
Tencent Hunyuan's new autonomous agent recursively generates, executes, and refines solutions to research and engineering tasks, beating rival systems on three benchmarks.