DeepSeek Releases V4.1 Flash With Million-Token Context
DeepSeek has launched V4.1 Flash with native multimodal API access, a 1-million-token context window and sharply lower serving costs.
DeepSeek has launched V4.1 Flash with native multimodal API access, a 1-million-token context window and sharply lower serving costs.
Tencent Hunyuan and research partners have released AuK, an open-source foundation model for controllable speech, audio and music generation.
Google DeepMind’s AlphaGenome Atlas predicts the molecular impact of every possible single-letter change in human DNA and opens the resource to academic researchers.
OpenAI has released Images 2.5 across ChatGPT and its API, combining faster generation with stronger reference fidelity and multi-turn editing.
OpenAI says an unreleased model and 10,000 agents produced a Lean-verified Navier–Stokes proof, but mathematicians have yet to validate it independently.
HiDream’s new world model unifies vision, video, 3D and actions, targeting robots that remain reliable under environmental disruption.
Anthropic’s model leads a live agent benchmark, although its strong task results come with the board’s highest median cost.
DeepSeek reportedly plans a vast Ascend deployment in Inner Mongolia, testing whether Chinese hardware can sustain frontier-scale inference.
Google has moved its latest music model into global consumer and developer channels, widening access beyond its earlier Flow Music launch.
MAI-Image-2.6-Flash pairs lower prices with production-focused editing, while the full model leads a major third-party benchmark.
AMD’s Threadripper Halo Station prototype pairs a 96-core Threadripper PRO processor with four Instinct accelerators and up to 2.6 TB of combined memory, with availability planned for 2027.
Astra combines a 1.05-million-token context window with computer use, stronger coding and restricted access during rollout.
Google pairs a broadly available agent model with a restricted cyber variant built for vulnerability discovery and automated patching.
Muse Spark 1.3 targets sustained coding and agent workflows, while Meta postpones both open weights and its highest reasoning setting.
Perplexity has released the Rust and Metal source, tests and reproduction materials for an engine it says runs Qwen3.6-35B-A3B up to 1.35 times faster than MLX-LM on an M5 Max.
Google’s Gemini models can now choose which moments of a video to inspect, reducing token use while improving retrieval accuracy.
Meta’s first real-time audio perception model processes speech in 80-millisecond chunks, supports more than 70 languages and adapts its delay to balance speed and accuracy.
OpenAI has begun a phased rollout of GPT-6 Astra, with tighter safeguards and additional restrictions on its most advanced cybersecurity capabilities.
Alibaba’s updated flagship model tops a major web-development benchmark while retaining a million-token context window.
Atlas combines generation, 3D reconstruction and simulation in one architecture aimed at filmmaking, robotics and virtual worlds.
The 330-million-parameter model forecasts related series and external variables without task-specific fine-tuning.
Open-weight GigaPath-Flash and GigaTIME-Flash sharply reduce compute and memory needs while retaining research performance.
The 743-billion-parameter coding model is now downloadable, turning its reported cyber gains into a broad access question.
OpenAI plans to stop supplying models directly to Cursor, turning a change-of-control clause into a major platform split.
Google’s production-ready video model adds 40-second extensions, keyframe interpolation, low-cost previews and output upscaling to 4K.
Midjourney’s first V8.2 editing model adds instruction-led changes, four-image references, inpainting and canvas expansion.
Tencent has released Hy4 Preview’s open weights, pairing a 770-billion-parameter MoE design with a one-million-token context window.
Google’s new speech model handles live multilingual audio, cleans disfluencies and powers voice-driven workflows in products such as Gemini for macOS.
The 320B mixture-of-experts model pairs native multimodality and one-million-token context with aggressively priced hosted inference.
Alibaba’s 125B-parameter multimodal model activates 6B parameters per token and introduces the architecture intended for Qwen4.
The design model can derive a reusable visual style from one to ten reference images without requiring a separate training process.
ModelBest says its open MiniCPM family has surpassed 50 million downloads as the Chinese lab prioritizes efficient, on-device AI.
The information-services group is taking direct control of a specialized model trained for legal, tax and regulatory work.
The Beijing company introduced a multimodal cooking model and two robots that adjust recipes by watching food change in real time.
OpenAI is lowering API and usage-credit prices for its flagship model by more than 20% as frontier inference competition intensifies.
The Chinese lab’s new multimodal endpoint combines its fast text model with image understanding, reusable uploads and agent-framework support.
Anthropic tested protein binders designed by Mythos Preview and Claude Opus 4.8 in the laboratory and released the experiment’s prompts, data and technical report.
OpenAI has slowed frontier development after its unreleased Astra model showed signs of critical offensive cyber capability.
Tripo’s new 3D preview generates native quad topology with up to 25,000 quads, targeting characters and real-time production assets.
Qwen’s download milestone signals that the center of gravity in open-weight AI is shifting toward China’s developer ecosystem.
Anthropic’s latest risk report raises its misalignment estimate and says a stronger internal model will not be released for now.
Apple has reportedly moved beyond simply embedding Qwen, developing a locally tailored model that could give it greater platform control.
Google’s new workhorse model targets high-volume agent and coding workloads with stronger performance and introductory pricing through 2026.
The 743B-base upgrade improves agentic coding and cyber tasks, but API access and open weights remain gated for safety review.
A Cerebras-powered Ultrafast mode accelerates OpenAI’s top model by up to 14 times, initially for a limited customer group.
North Micro Vision packages document and visual understanding into a 2.4B model with downloadable weights and an Apache 2.0 license.
The 0813 build upgrades agent performance, adds Responses API support and extends reasoning controls across DeepSeek services.
Dyna Robotics says its new world-action model uses one million hours of human video to improve general-purpose robot learning.
DeepMind’s SL2T model turns American Sign Language into text on Pixel 11, moving sign translation from research into daily phone use.
Grok 4.6 targets coding and sustained agent work with a 500,000-token context window and pricing below several frontier rivals.
Alibaba has released its largest Qwen model’s weights, giving independent operators a frontier-scale foundation for coding and agents.
The 7.9B-parameter mixture-of-experts model targets economical local and agent deployment, activating only 1.3B parameters per token.
LTX’s new open-weight world model adds adaptive rendering, native 4K HDR output and stronger multi-shot control for local video production.
Mistral will host rival open models, beginning with Z.ai’s GLM-5.2, while adding regional inference and reserved European capacity.
The open 30B mixture-of-experts model activates 3B parameters per token and ships with training data, recipes and a new model router.
Upstage became the first Korean lab represented across Arena’s agent, web-development code and general text leaderboards.
Anthropic scrapped the planned end of Claude Sonnet 5's introductory rate, locking $2 per million input and $10 per million output tokens past August 31.
Microsoft AI's newest image model entered Arena's text-to-image leaderboard at second with 1336 points, up from 2.5's tenth place and 45 behind GPT Image 2.
Arena's first human-preference scores put Meta's new Apache-2.0 model at #77 in Code Arena WebDev and #97 in Text Arena, 26th and 24th among open models.
OpenAI released a purpose-trained cyber model that answers 95% of advanced security prompts its general models refuse, behind identity-verified access tiers.
Meta Superintelligence Labs released its first open model — a 30B agentic multimodal system under Apache 2.0 that runs on a single consumer GPU.
A new channel lets one Claude Code session hand findings to another over a local socket, with cross-machine exchanges limited to replies only.
SpaceXAI's new image model entered both Arena leaderboards at No.2 behind GPT-Image-2; by 11 August it had been pushed to third in text-to-image by Microsoft's MAI-Image-2.6, while holding second in image editing.
Artificial Analysis scored Ant Group's 124B open-weights model at 38, placing its 5B active parameters ahead of every flash-tier rival.
OpenAI says preliminary evaluations cannot rule out that its unreleased Astra model reaches the Critical cyber tier, and has paused parts of its development.
The Qwen-Audio-based product bundles dictation, voice agents and audiobook generation, and ships free on iOS, Android, Mac and Windows.
The Toronto startup builds chips whose metal layers are finalised around one model's weights, promising inference without HBM, advanced packaging or liquid cooling.
A rewritten classifier constitution cut biology-related fallbacks by roughly 85%, after months of complaints that the model refused routine health questions.
Paid ChatGPT users get an updated GPT-5.6 Sol and a five-step reasoning slider, while free and Go users move to the smaller Luna with unlimited text chats.
ByteDance switched on public API access for its 30-second video model; Luma had it live a day earlier and Runway shipped it on the day, with Krea, Lovart, Pika and others announcing integrations still rolling out.
Anthropic posted roles for a custom silicon team, saying it intends to co-design hardware and models so Claude runs faster and more cheaply.
The Chinese lab that set the global floor for inference pricing told developers it will raise API rates substantially, without naming figures or a date.
Google began emailing users that Assistant stops working on phones, tablets, Wear OS watches, headphones and Android Auto on September 4, with Gemini taking over.
Meta Superintelligence Labs released a terminal coding agent in beta alongside a co-trained model priced at $1.25 per million input tokens, undercutting rivals.
Artificial Analysis published its full evaluation of Alibaba's 2.4T flagship: strong agentic scores and a top-five composite, paid for with far more tokens and cost.
Black Forest Labs' 20-second video model with native audio opened to the public on Runway, Krea and Replicate after a gated rollout.
JD.com released Apache-2.0 weights for JoyAI-Video-Edit, which edits 720p video at about 30 frames per second on a single NVIDIA B200.
The Apache-2.0 guard model takes safety policies as plain-language questions at inference and runs on a single 16GB GPU.
Beijing lab Sand.ai released Apache-2.0 weights for a 114-billion-parameter sparse video model that generates 1080p clips with synchronized audio.
Tencent opened its 295B-parameter Hy3 model to overseas users through WorkBuddy's international edition, cloud APIs and design tools, free until August 31.
Video Arena placed MiniMax-H3 first among open models in both text-to-video and image-to-video, some 280 Elo points clear of the next open weights entry. The weights are now public, but the licence restricts open-weight use in the US, EU, UK and South Korea.
Tencent's new speech model reports about 3% word error rates in Mandarin, English and Cantonese, and is live on Tencent Cloud and in Yuanbao.
Wan launched a Real-time feature suite whose first capability turns a live phone camera feed into stylised imagery, starting with anime looks.
Cogent AI released VR-1, a security-specialised model that composes and verifies multi-stage enterprise attack paths, with weights withheld from the public.
MiniMax published downloadable weights for H3, its omni-modal 2K video model, making a top-ranked video generator locally runnable for the first time.
Alibaba's new flagship is a 2.4-trillion-parameter mixture-of-experts model aimed at long-horizon agent work, with open weights promised next week.
Tokyo lab Sakana AI opened commercial access to Namazu, a Japanese-specialised model built by fine-tuning Moonshot's open-weight Kimi K2.6.
Epoch AI expanded its FrontierMath Open Problems set to 50 research questions and logged two fresh AI solutions, both in the middle "Moderately Interesting" tier. Three problems were pruned in the same update.
Google opened its background-running personal agent to AI Pro subscribers in more than 160 additional markets, a week after the tier launched in the US.
OpenAI introduced its next major model family, Astra, by publishing ten solutions to long-open problems in mathematics and theoretical computer science.
DeepSeek's re-post-trained V4-Flash entered public beta on Thursday, scoring 50 on Artificial Analysis' index at unchanged $0.14/$0.28 pricing.
Google DeepMind released three physical-AI models that let a single policy drive legs, torso, arms and fingers, and coordinate multiple robots on one task.
P-Image-Ideogram launches with four quality tiers, native 1K and 2K output and three-to-eight-second latency, aimed squarely at production image pipelines.
OpenAI slashed API pricing across its GPT-5.6 line, dropping Luna to $0.20 per million input tokens and adding a paid Fast mode for Sol.
ByteDance's Seedance 2.5 left closed enterprise beta on Friday, launching on Jimeng in China and Dreamina globally with 30-second single-pass generation.
Two weeks after its 975B flagship, Thinking Machines released a 276B Apache-2.0 sibling with 12B active parameters and full text, image and audio reasoning.
Alibaba's Wan team launched a Realtime mode on wan.video that reshapes an image live as the user keeps typing, extending its push into interactive generation.
OpenAI reports that after deployment it turned GPT-5.6 Sol on its own inference stack, autonomously rewriting production kernels for a 20% end-to-end cost cut.
Ideogram shipped a dedicated object-removal model that erases shadows and reflections too, and says it beats Nano Banana 2, FLUX Erase and GPT Image-2.
Google launched Lyria 3.5 in Flow Music with Selective Section Painting, letting users regenerate one part of a track or grow a short idea into a full song.
OpenAI says GPT-5.6 Sol jumps from 7.8% to 38.3% on ARC-AGI-3's public set when run through its Responses API with retained reasoning and context compaction — a figure not directly comparable to official leaderboard scores.
Business Insider reports Amazon has halted development on Nova Premier, Omni, Reel and Canvas, shifting resources to a research group under Pieter Abbeel.
Anthropic says its frontier model autonomously produced two publishable cryptanalysis results, cutting HAWK-256's effective security and speeding an AES attack 200-800x.
Google's hosted agent runtime switches its default to Gemini 3.6 Flash and adds pre/post tool-call hooks, token ceilings, cron triggers and a free tier.
LFM2.5-Encoder-230M and 350M bring an 8,192-token window to BERT-class models and cut CPU latency roughly threefold at full length.
Two new speech-to-text models replace Whisper as OpenAI's recommended default, with steerable context inputs and pricing from $0.0045 per minute.
Alibaba Cloud says its Zhenwu M890 supernode is the first Chinese system to run Moonshot's 2.8-trillion-parameter model; Huawei Ascend also claims day-0 support.
Moonshot AI has published Kimi K3's full 2.8-trillion-parameter weights under the Kimi K3 License — the largest open-weight model release to date.
Microsoft's first in-house security model handles up to 90% of vulnerability-hunting work inside MDASH, with OpenAI models still reserved for the hardest cases.
The Kimi team released a benchmark that strips reasoning out of multimodal evaluation and grades whether a model actually saw the image correctly.
Kuaishou's KwaiKAT team says its new agentic coder, trained by RL in 100,000+ verifiable repository environments, edges Claude Opus 4.8 on PinchBench.
ARC Prize reports Claude Opus 5 scored 30.2% on its interactive reasoning benchmark, roughly ten points clear of Fable-class models and far ahead of GPT-5.6 Sol's 7.8%.
Anthropic's Opus 5 approaches Fable 5-class results and triples the field on ARC-AGI-3, all at unchanged Opus pricing and lower token cost.
Midjourney promoted V8.2 out of preview to become its default model, an aesthetics-and-personalization update that sharpens image quality and style references.
Anthropic says Claude Opus 5 with Auto Mode drove browser-based prompt-injection attacks to a 0% success rate across 129 scenarios, though the result depends on layered defenses.
Sakana AI's upgraded Fugu Ultra v1.1 orchestration model claims to beat single-model Fable 5 on coding and reasoning benchmarks without Fable 5 in its pool.
DeepSeek's flagship V4 line exits preview: legacy chat and reasoner endpoints shut down today as new peak and off-peak API pricing takes effect.
Black Forest Labs' new flagship generates 20-second clips with native audio, claims wins over Kling and Runway, and promises open weights later in 2026.
Alibaba Cloud says its Zhenwu M890 supernode, built on in-house T-Head chips, now runs 2.4T-parameter Qwen3.8 in production, cutting inference costs over 40%.
Cursor's new request-level router sends each coding task to the best-fit model, with early enterprise users reporting 30-50% savings and no quality drop.
A joint UK AISI and US CAISI evaluation finds Moonshot's Kimi K3 leads open-weight models on offensive cyber tasks yet sits far below US frontier systems.
Qwen-Image-3.0 renders dense infographics and ten-pixel text in one pass, while Qwen-Audio-3.0-TTS tops a speech leaderboard across 16 languages.
Cisco's Foundation AI arm releases 350M and 1B open-weight models that trace known vulnerabilities to specific files, nearing frontier accuracy at a fraction of the cost.
Poolside released Laguna S 2.1, a 118B MoE coding model with 1M context under an open license, hitting 78.5% on SWE-Bench Multilingual with only 8B active params.
A new paper tops Hugging Face's daily list with full-parameter post-training of trillion-parameter DeepSeek-V4 models on Huawei's Ascend SuperPOD.
Google's new Flash tier targets agentic workloads with fewer output tokens and lower prices, plus a cybersecurity model piloted with governments.
NVIDIA open-sourced Cosmos 3 Edge, a 4B-parameter world model that reasons and generates robot actions locally, hitting 15 Hz real-time control on Jetson Thor.
Moonshot AI halted new Kimi K3 subscriptions after demand surged sixfold, as third-party tests put the open-weight model within reach of Claude Fable 5.