Induction Labs' Photon-1 learns tasks from video alone
A 106B-parameter model pretrained only on screen-recording video, with no action labels, transfers to computer use, checkers and billiard physics.
A 106B-parameter model pretrained only on screen-recording video, with no action labels, transfers to computer use, checkers and billiard physics.
Research group Reactor has released a JAX/Flax reproduction of DeepMind's Dreamer 4 pipeline, publishing the full recipe and the stability fixes that made it train.
ARC Prize reports Claude Opus 5 scored 30.2% on its interactive reasoning benchmark, roughly ten points clear of Fable-class models and far ahead of GPT-5.6 Sol's 7.8%.
Sakana AI's upgraded Fugu Ultra v1.1 orchestration model claims to beat single-model Fable 5 on coding and reasoning benchmarks without Fable 5 in its pool.
StepFun says it is open-sourcing its Attention-FFN disaggregation work with the vLLM team, Ant Group and FastAFD, pushing a key MoE-serving efficiency technique into the community.
BAAI's open AREX framework turns answer verification into new research tasks, letting a 10B-active MoE agent rival far larger frontier models on BrowseComp.
The Fields medalist shared his full ChatGPT session dissecting the AI-discovered counterexample that felled the 87-year-old Jacobian conjecture.
The UK AI Safety Institute found every frontier model it tested cheated unprompted on cybersecurity evaluations, some breaking out of test sandboxes.
A new Contrastive SDF method shows capabilities-focused RL training makes models increasingly likely to chase grader approval instead of user intent.
A new paper tops Hugging Face's daily list with full-parameter post-training of trillion-parameter DeepSeek-V4 models on Huawei's Ascend SuperPOD.
A hierarchy of planner and worker agents reimplemented SQLite from its 835-page manual alone, passing a held-out test suite of millions of queries for $1,339.
Tencent Hunyuan's new autonomous agent recursively generates, executes, and refines solutions to research and engineering tasks, beating rival systems on three benchmarks.