DeepSeek Releases V4.1 Flash With Million-Token Context
DeepSeek has launched V4.1 Flash with native multimodal API access, a 1-million-token context window and sharply lower serving costs.
DeepSeek has launched V4.1 Flash with native multimodal API access, a 1-million-token context window and sharply lower serving costs.
Tencent Hunyuan and research partners have released AuK, an open-source foundation model for controllable speech, audio and music generation.
Mistral’s record European funding round gives its open-weight strategy unprecedented capital, while tying Samsung and Europe more closely to frontier AI.
The chipmaker is buying the leading open-model hub while promising continued support for rival hardware, clouds and models.
Perplexity has released the Rust and Metal source, tests and reproduction materials for an engine it says runs Qwen3.6-35B-A3B up to 1.35 times faster than MLX-LM on an M5 Max.
The 330-million-parameter model forecasts related series and external variables without task-specific fine-tuning.
Open-weight GigaPath-Flash and GigaTIME-Flash sharply reduce compute and memory needs while retaining research performance.
Ollama has replaced opaque GPU-time limits with published token rates and monthly credits across its paid cloud plans.
The partners say the project will nearly triple Together AI’s capacity and support more than $5 billion in annualized revenue.
The open-source agent adds remote sessions, interactive dashboards and scoped secret handling in a major architectural overhaul.
The 743-billion-parameter coding model is now downloadable, turning its reported cyber gains into a broad access question.
Tencent has released Hy4 Preview’s open weights, pairing a 770-billion-parameter MoE design with a one-million-token context window.
The 320B mixture-of-experts model pairs native multimodality and one-million-token context with aggressively priced hosted inference.
Conflicting reports say Nvidia has agreed, or is close, to acquire the central hub of the open-model ecosystem for $12.9 billion.
Alibaba’s 125B-parameter multimodal model activates 6B parameters per token and introduces the architecture intended for Qwen4.
ModelBest says its open MiniCPM family has surpassed 50 million downloads as the Chinese lab prioritizes efficient, on-device AI.
Gradio’s new Workflow interface turns typed AI pipelines into inspectable canvases, callable REST endpoints and deployable applications.
The open-model hub is reportedly testing buyer interest, raising questions over who could own a critical layer of AI infrastructure.
Microsoft’s updated neural chemistry functional improves accuracy and moves into software already used for molecular simulation.
Qwen’s download milestone signals that the center of gravity in open-weight AI is shifting toward China’s developer ecosystem.
North Micro Vision packages document and visual understanding into a 2.4B model with downloadable weights and an Apache 2.0 license.
Alibaba has released its largest Qwen model’s weights, giving independent operators a frontier-scale foundation for coding and agents.
The 7.9B-parameter mixture-of-experts model targets economical local and agent deployment, activating only 1.3B parameters per token.
LTX’s new open-weight world model adds adaptive rendering, native 4K HDR output and stronger multi-shot control for local video production.
Mistral will host rival open models, beginning with Z.ai’s GLM-5.2, while adding regional inference and reserved European capacity.
The open 30B mixture-of-experts model activates 3B parameters per token and ships with training data, recipes and a new model router.
Arena's first human-preference scores put Meta's new Apache-2.0 model at #77 in Code Arena WebDev and #97 in Text Arena, 26th and 24th among open models.
Z.ai says its agentic coding environment passed one million users and wiped usage counters for every GLM Coding Plan subscriber as a thank-you.
In a 6,500-word essay Zuckerberg recommitted Meta to open models, attacked closed rivals for concentrating power, and floated auctioning superintelligence compute.
Meta Superintelligence Labs released its first open model — a 30B agentic multimodal system under Apache 2.0 that runs on a single consumer GPU.
Artificial Analysis scored Ant Group's 124B open-weights model at 38, placing its 5B active parameters ahead of every flash-tier rival.
A vendor-neutral format bundles Agent Skills and MCP server configs into one portable plugin, supported at launch by ChatGPT, Codex, Cursor, Copilot, Kiro and VS Code — with Google joining the steering committee the same day.
A Nature paper puts Google DeepMind's WeatherNext Cyclones ahead of operational systems by over a day of lead time, with code and weights released on GitHub.
The open-source agent platform released an enterprise-grade distributed swarm architecture, first deployed in production at China's Postal Savings Bank.
Mixture-of-Kittens fuses all MoE communication and compute into one deterministic Blackwell kernel, lifting Cursor's training throughput 1.41x.
JD.com released Apache-2.0 weights for JoyAI-Video-Edit, which edits 720p video at about 30 frames per second on a single NVIDIA B200.
The Apache-2.0 guard model takes safety policies as plain-language questions at inference and runs on a single 16GB GPU.
Beijing lab Sand.ai released Apache-2.0 weights for a 114-billion-parameter sparse video model that generates 1080p clips with synchronized audio.
Video Arena placed MiniMax-H3 first among open models in both text-to-video and image-to-video, some 280 Elo points clear of the next open weights entry. The weights are now public, but the licence restricts open-weight use in the US, EU, UK and South Korea.
Genspark released GenOffice, which it calls the first full-featured open-source AI office suite, free and ad-free on Windows and macOS under Apache-2.0.
MiniMax published downloadable weights for H3, its omni-modal 2K video model, making a top-ranked video generator locally runnable for the first time.
Tokyo lab Sakana AI opened commercial access to Namazu, a Japanese-specialised model built by fine-tuning Moonshot's open-weight Kimi K2.6.
Supabase released Evals, an Apache-2.0 benchmark that scores Claude Code, Codex and OpenCode on real database, auth and deployment tasks.
JetBrains Research released KotlinLLM under Apache 2.0, an IntelliJ plugin that generates Kotlin function bodies at runtime and hot-reloads them until they pass.
Two weeks after its 975B flagship, Thinking Machines released a 276B Apache-2.0 sibling with 12B active parameters and full text, image and audio reasoning.
ThunderAgent schedules whole agent workflows instead of single calls, reporting 2x throughput and 6x lower latency than SGLang on one 8xH100 node.
LFM2.5-Encoder-230M and 350M bring an 8,192-token window to BERT-class models and cut CPU latency roughly threefold at full length.
The Apache-2.0 tool scans repositories, tracks findings across runs and verifies fixes, and reached the Hacker News front page before OpenAI announced it.
Responding to months of speculation, Anthropic published its open-weights position: no bans, but mandatory pre-release safety testing for any sufficiently capable model.
Alibaba Cloud says its Zhenwu M890 supernode is the first Chinese system to run Moonshot's 2.8-trillion-parameter model; Huawei Ascend also claims day-0 support.
Moonshot AI has published Kimi K3's full 2.8-trillion-parameter weights under the Kimi K3 License — the largest open-weight model release to date.
The Kimi team and kvcache-ai released the Firecracker-based environment system used for Kimi K3's agentic RL training under an MIT licence.
The Kimi team released a benchmark that strips reasoning out of multimodal evaluation and grades whether a model actually saw the image correctly.
Nvidia's Open Secure AI Alliance lists Microsoft, IBM, Hugging Face and others among 50-plus inaugural partners; separately, the open-weights policy letter has grown to 50 signatories including Google and OpenAI, with Anthropic and Amazon still out.
Kuaishou's KwaiKAT team says its new agentic coder, trained by RL in 100,000+ verifiable repository environments, edges Claude Opus 4.8 on PinchBench.
Research group Reactor has released a JAX/Flax reproduction of DeepMind's Dreamer 4 pipeline, publishing the full recipe and the stability fixes that made it train.
Reporting indicates the White House is leaning toward targeting specific Chinese open-weight models on security grounds rather than a blanket prohibition.
The July 24 open-weights letter has expanded beyond Nvidia, Microsoft and Meta, with updated signatories now including Google and OpenAI while Anthropic remains absent.
OpenAI and Google have joined the expanding open-weights appeal to Washington, making Anthropic the most visible major frontier-lab holdout.
StepFun says it is open-sourcing its Attention-FFN disaggregation work with the vLLM team, Ant Group and FastAFD, pushing a key MoE-serving efficiency technique into the community.
Andrew Ng shipped an MIT-licensed, local-first desktop agent that plans tasks across files and apps and hands back finished deliverables instead of chat.
BAAI's open AREX framework turns answer verification into new research tasks, letting a 10B-active MoE agent rival far larger frontier models on BrowseComp.
DeepSeek's flagship V4 line exits preview: legacy chat and reasoner endpoints shut down today as new peak and off-peak API pricing takes effect.
Cisco's Foundation AI arm releases 350M and 1B open-weight models that trace known vulnerabilities to specific files, nearing frontier accuracy at a fraction of the cost.
Poolside released Laguna S 2.1, a 118B MoE coding model with 1M context under an open license, hitting 78.5% on SWE-Bench Multilingual with only 8B active params.
NVIDIA open-sourced Cosmos 3 Edge, a 4B-parameter world model that reasons and generates robot actions locally, hitting 15 Hz real-time control on Jetson Thor.
Moonshot AI halted new Kimi K3 subscriptions after demand surged sixfold, as third-party tests put the open-weight model within reach of Claude Fable 5.