DeepSeek Releases V4.1 Flash With Million-Token Context
DeepSeek has launched V4.1 Flash with native multimodal API access, a 1-million-token context window and sharply lower serving costs.
DeepSeek has launched V4.1 Flash with native multimodal API access, a 1-million-token context window and sharply lower serving costs.
OpenAI appointed alignment researcher Paul Christiano to its Foundation board and safety committee as frontier-model oversight faces sharper scrutiny.
Devin maker Cognition has raised more than $2 billion, giving the coding-agent company a $48 billion valuation as enterprise adoption accelerates.
Meta is expanding its personal AI assistant across its products, powered by Muse Spark models that support multimodal reasoning, tool use and longer-horizon tasks.
Mistral’s record European funding round gives its open-weight strategy unprecedented capital, while tying Samsung and Europe more closely to frontier AI.
A small lung-disease trial links rentosertib to younger proteomic readings, but cannot yet distinguish treatment effects from slower aging.
South Korea expects AI data centers and chip fabs to add up to 30 gigawatts of demand, pushing nuclear power back into the national AI strategy.
Anthropic’s model leads a live agent benchmark, although its strong task results come with the board’s highest median cost.
DeepSeek reportedly plans a vast Ascend deployment in Inner Mongolia, testing whether Chinese hardware can sustain frontier-scale inference.
Gimlet’s large Series B backs software that pools heterogeneous accelerators as inference buyers seek alternatives to fixed GPU clusters.
GitHub’s experimental coding agent selects single, cascade or critique workflows to trade model cost against verified task quality.
MAI-Image-2.6-Flash pairs lower prices with production-focused editing, while the full model leads a major third-party benchmark.
Replit’s production-ready MCP server lets assistants including ChatGPT and Claude create, modify and deploy applications.
AMD’s Threadripper Halo Station prototype pairs a 96-core Threadripper PRO processor with four Instinct accelerators and up to 2.6 TB of combined memory, with availability planned for 2027.
The new global model produces hourly forecasts from satellite and station data and is rolling out across Search, Maps, Gemini and Google Cloud services.
The chipmaker is buying the leading open-model hub while promising continued support for rival hardware, clouds and models.
The global Daybreak program will subsidize advanced AI tools, training and technical support for essential-service operators.
Cursor adds scalable worker pools across corporate infrastructure while retaining its cloud-based planning, inference and orchestration layer.
Perplexity has released the Rust and Metal source, tests and reproduction materials for an engine it says runs Qwen3.6-35B-A3B up to 1.35 times faster than MLX-LM on an M5 Max.
Alibaba’s updated flagship model tops a major web-development benchmark while retaining a million-token context window.
The agent developer is again operating independently after a disruptive separation from Meta, with its founding team continuing to lead the company.
Ollama has replaced opaque GPU-time limits with published token rates and monthly credits across its paid cloud plans.
The partners say the project will nearly triple Together AI’s capacity and support more than $5 billion in annualized revenue.
ChatGPT Search must undergo systemic-risk reviews and independent audits after crossing the EU’s 45-million-user threshold.
Passengers in parts of Beijing and Guangzhou can now summon DiDi’s purpose-built R2, equipped with 33 sensors and redundant controls.
Baseten says an agentic system generated, tested and deployed production optimizations across image and language models.
The short-dated private debt will finance Nvidia chips leased to Microsoft, deepening AI infrastructure’s reliance on asset-backed credit.
OpenAI plans to stop supplying models directly to Cursor, turning a change-of-control clause into a major platform split.
The new workspace connects specialized models, biological tools and reviewable evidence in a governed research workflow.
More than 150 organizations warn that AI-enabled attacks will spread within months and urge immediate support for critical-infrastructure defenders.
Anthropic’s experimental MHS interface lets AI agents coordinate laboratory and manufacturing equipment through a common control layer.
Google’s production-ready video model adds 40-second extensions, keyframe interpolation, low-cost previews and output upscaling to 4K.
The six-year agreement ties Anthropic’s next growth phase to 460 megawatts of Nvidia Vera Rubin infrastructure in West Virginia.
Amazon will deploy two million additional Blackwell Ultra and Rubin-generation GPUs while extending Nvidia technology across AWS and robotics.
Conflicting reports say Nvidia has agreed, or is close, to acquire the central hub of the open-model ecosystem for $12.9 billion.
ModelBest says its open MiniCPM family has surpassed 50 million downloads as the Chinese lab prioritizes efficient, on-device AI.
OpenAI says its first inference chip delivers lower latency and more work per watt across three large open-weight models.
Portable Computer keeps agent planning, file analysis and tool execution on a DGX Spark, escalating to cloud models only with permission.
EA and three global music groups participated in Stability AI’s Series B, tying its next phase more closely to licensed creative production.
OpenAI banned accounts tied to a covert Russian campaign that used ChatGPT to promote a fabricated institute and stolen scholarship.
The information-services group is taking direct control of a specialized model trained for legal, tax and regulatory work.
The open-model hub is reportedly testing buyer interest, raising questions over who could own a critical layer of AI infrastructure.
The chipmaker is reportedly discussing an equity investment as Perplexity’s agent-led revenue growth drives a steep valuation increase.
Japan’s Defense Ministry hired Sakana AI to test agent-based intelligence analysis, moving sovereign AI into security operations.
The Beijing company introduced a multimodal cooking model and two robots that adjust recipes by watching food change in real time.
Claude Security gives Enterprise customers access to Mythos 5 vulnerability analysis without exposing the restricted model directly.
OpenAI is lowering API and usage-credit prices for its flagship model by more than 20% as frontier inference competition intensifies.
Runway’s new model converts uploaded or generated SDR footage into high-bit-depth HDR formats designed for professional post-production.
Slack’s new shared coding channels expose agent work, previews and code changes to entire teams instead of individual users.
Cursor’s cloud agents can now react to PR and Slack events, pursue durable goals and delegate work to isolated virtual machines.
A privacy-preserving safety system will analyze multi-step risks without giving OpenAI personnel access to customer content.
The former Nvidia researchers aim to create simulated worlds where robots can learn safely through large-scale trial and error.
Cerebras says its fourth-generation system delivers up to twice the inference speed of CS-3 through three WSE-3 Turbo processors and a redesigned rack-scale platform.
OpenAI launched a dedicated experience that automatically places identified 13-to-17-year-olds behind stronger safety and learning controls.
Tripo’s new 3D preview generates native quad topology with up to 25,000 quads, targeting characters and real-time production assets.
A tracked shipment exposed an Amazon operation that buys, disassembles and scans physical books to obtain AI-training material.
Cursor is expanding from AI coding into repository infrastructure, bringing code, agents and deployment integrations under one roof.
Groq’s latest official financing is a $350 million Series A led by Disruptive, with planned participation from Nvidia, bringing its recent funding to $1 billion.
The generative-media company has secured a major round to scale its video and image platform amid intensifying AIGC competition.
QwenWork can now operate through WeCom chats, taking Alibaba’s workplace agent into the heart of Tencent’s enterprise ecosystem.
Meshy’s new technical report explains how its latest system more faithfully transfers a reference image into usable 3D geometry.
The reported agreement would combine AI model routing, usage metering and payments in one powerful infrastructure provider, although the companies have not publicly announced the deal.
Qwen’s download milestone signals that the center of gravity in open-weight AI is shifting toward China’s developer ecosystem.
Users in most markets can remove visible marks from generated images, video and music, while Google retains SynthID and C2PA provenance; legally required marks remain in some jurisdictions.
Cursor has formally joined SpaceXAI in an all-stock acquisition valuing the coding-agent company at approximately $60 billion.
Apple has reportedly moved beyond simply embedding Qwen, developing a locally tailored model that could give it greater platform control.
OpenAI’s Computer History lets ChatGPT recall recent activity across Mac apps and websites, with a 48-hour event-recording window and organization-level controls.
Investor demand pushed Databricks far beyond its initial funding target as AI infrastructure spending continues to concentrate.
Google’s new workhorse model targets high-volume agent and coding workloads with stronger performance and introductory pricing through 2026.
A Cerebras-powered Ultrafast mode accelerates OpenAI’s top model by up to 14 times, initially for a limited customer group.
Suno is turning its music generator into an editable production environment, giving creators finer control over AI-made tracks.
The 0813 build upgrades agent performance, adds Responses API support and extends reasoning controls across DeepSeek services.
DeepMind’s SL2T model turns American Sign Language into text on Pixel 11, moving sign translation from research into daily phone use.
Grok 4.6 targets coding and sustained agent work with a 500,000-token context window and pricing below several frontier rivals.
The Japan-focused Namazu assistant now turns Japanese instructions into runnable apps, games, charts and documents without requiring a login.
Twitch now lets creators refuse generative-AI training, but its default setting leaves eligible channel content available to Amazon.
Claude is adding model-level marks to generated text and signed C2PA provenance to supported files worldwide as EU transparency rules take effect.
Google’s AI assistant crossed one billion monthly users, matching ChatGPT and becoming its fastest-growing product to reach that scale.
Mistral will host rival open models, beginning with Z.ai’s GLM-5.2, while adding regional inference and reserved European capacity.
Upstage became the first Korean lab represented across Arena’s agent, web-development code and general text leaderboards.
Alibaba's Wan team released a command-line tool that lets coding agents call its image and video models programmatically, installed by pasting one command.
Anthropic scrapped the planned end of Claude Sonnet 5's introductory rate, locking $2 per million input and $10 per million output tokens past August 31.
Z.ai says its agentic coding environment passed one million users and wiped usage counters for every GLM Coding Plan subscriber as a thank-you.
In a 6,500-word essay Zuckerberg recommitted Meta to open models, attacked closed rivals for concentrating power, and floated auctioning superintelligence compute.
The Economist reports UK employment tribunals are buried under chatbot-drafted claims, with open cases up 55 percent to 64,000 and non-existent statutes cited.
Meta Superintelligence Labs released its first open model — a 30B agentic multimodal system under Apache 2.0 that runs on a single consumer GPU.
A gas plant permitted to serve Amazon's Pecos County data center would emit more carbon dioxide annually than any other US power station.
From August 14 the coding agent stops asking permission per action for Pro, Max and Team users, routing tool calls through a safety classifier instead.
The Qwen-Audio-based product bundles dictation, voice agents and audiobook generation, and ships free on iOS, Android, Mac and Windows.
The Toronto startup builds chips whose metal layers are finalised around one model's weights, promising inference without HBM, advanced packaging or liquid cooling.
The agent-first browser runs in V8 isolates on Workers with no Chromium underneath, claiming 3–7x lower CPU and memory on common agent tasks at 1.7–1.8x the wall-clock time.
Paid ChatGPT users get an updated GPT-5.6 Sol and a five-step reasoning slider, while free and Go users move to the smaller Luna with unlimited text chats.
The open-source agent platform released an enterprise-grade distributed swarm architecture, first deployed in production at China's Postal Savings Bank.
ByteDance switched on public API access for its 30-second video model; Luma had it live a day earlier and Runway shipped it on the day, with Krea, Lovart, Pika and others announcing integrations still rolling out.
Anthropic posted roles for a custom silicon team, saying it intends to co-design hardware and models so Claude runs faster and more cheaply.
The Chinese lab that set the global floor for inference pricing told developers it will raise API rates substantially, without naming figures or a date.
The Ulanqab Xinghe base entered production with a 120,000-square-metre hall, million-card parallel design and direct green-power supply, phase one of a 5GW plan.
Google began emailing users that Assistant stops working on phones, tablets, Wear OS watches, headphones and Android Auto on September 4, with Gemini taking over.
Alphabet made Demis Hassabis its chief scientist and handed DeepMind's daily operations to Koray Kavukcuoglu, hours before Jeff Dean announced a rival startup.
Reddit is rolling out Rules Hub, which lets moderators write rules in prose for a language model to enforce by intent, and will reduce reliance on karma gates.
The six-year deal routes Anthropic through a 133MW hydro-powered Norwegian site running Nvidia Vera Rubin systems, operated with Bitdeer.
Mixture-of-Kittens fuses all MoE communication and compute into one deterministic Blackwell kernel, lifting Cursor's training throughput 1.41x.
Tencent opened its 295B-parameter Hy3 model to overseas users through WorkBuddy's international edition, cloud APIs and design tools, free until August 31.
Governor Greg Abbott ordered ERCOT and state regulators to audit every new data center before it can connect, with 474GW now queued.
Cursor shipped plugins letting its agents read and write across Gmail, Drive, Calendar, Docs, Sheets and Chat, pushing a coding tool into general knowledge work.
Tencent's new speech model reports about 3% word error rates in Mandarin, English and Cantonese, and is live on Tencent Cloud and in Yuanbao.
Wan launched a Real-time feature suite whose first capability turns a live phone camera feed into stylised imagery, starting with anime looks.
Apple's highest-volume laptop is now short of supply as DRAM makers divert wafers to AI data centres, with custom orders slipping into September.
Alibaba merged three internal agent products into QwenWork, a suite pairing desktop, cloud and team agents on top of the new Qwen3.8 model.
Cogent AI released VR-1, a security-specialised model that composes and verifies multi-stage enterprise attack paths, with weights withheld from the public.
Genspark released GenOffice, which it calls the first full-featured open-source AI office suite, free and ad-free on Windows and macOS under Apache-2.0.
Alibaba's new flagship is a 2.4-trillion-parameter mixture-of-experts model aimed at long-horizon agent work, with open weights promised next week.
Tokyo lab Sakana AI opened commercial access to Namazu, a Japanese-specialised model built by fine-tuning Moonshot's open-weight Kimi K2.6.
Apple has imposed a submission ceiling and a 30-day cool-off on Feedback Assistant bug reports, blaming a deluge of AI-assisted findings.
Bolt added an agent that scans, patches and hardens apps built on its platform in one click, with audits and remediation excluded from billing.
Google opened its background-running personal agent to AI Pro subscribers in more than 160 additional markets, a week after the tier launched in the US.
Google rolled back a Nano Banana 2-powered generator that let users paint fabricated imagery onto satellite maps, citing policy-violating outputs being shared online.
Reuters reports OpenAI's widening investigation into the Hugging Face breach has turned up additional cases of agents breaking out of test sandboxes.
Snap has adjusted Spotlight recommendations to exclude entirely AI-generated videos while still allowing AI as an editing and enhancement tool.
DeepSeek's re-post-trained V4-Flash entered public beta on Thursday, scoring 50 on Artificial Analysis' index at unchanged $0.14/$0.28 pricing.
Google wired its agent into Chrome, letting Spark act on live pages using a user's own logged-in sessions, while expanding Spark access to 160-plus countries.
OpenAI slashed API pricing across its GPT-5.6 line, dropping Luna to $0.20 per million input tokens and adding a paid Fast mode for Sol.
Perplexity rebuilt its Spaces feature as Projects, persistent workspaces for its Computer agent with a shared file system and connection to the Brain memory layer.
OpenAI reports that after deployment it turned GPT-5.6 Sol on its own inference stack, autonomously rewriting production kernels for a 20% end-to-end cost cut.
ChatGPT for Academic Researchers starts with 10,000 scientists and scales to 100,000 by 2027, part of a $250M-plus commitment to external research.
Business Insider reports Amazon has halted development on Nova Premier, Omni, Reel and Canvas, shifting resources to a research group under Pieter Abbeel.
The data-security firm signed a letter of intent for the non-human identity startup, its third acquisition this year, as enterprises scramble to govern agent access.
Google's hosted agent runtime switches its default to Gemini 3.6 Flash and adds pre/post tool-call hooks, token ceilings, cron triggers and a free tier.
Seoul's benchmark fell about 6% on Wednesday after a 10.84% Tuesday collapse, as SK Hynix missed estimates and China's first domestic DUV tools entered service.
The desktop agent reads local files, drives Office apps and routes subtasks across 20-plus frontier models, gated behind Perplexity's $200-a-month Max tiers.
The Chinese security vendor reports its agent reproduced 1,301 of 1,507 vulnerabilities on CyberGym, claiming fourth place globally atop open GLM-5.2 weights.
Alibaba Cloud says its Zhenwu M890 supernode is the first Chinese system to run Moonshot's 2.8-trillion-parameter model; Huawei Ascend also claims day-0 support.
Cursor's first localised plan bundles Composer 2.5, Grok 4.5, autonomous cloud agents and an iOS app for roughly a third of Pro's price.
Microsoft's first in-house security model handles up to 90% of vulnerability-hunting work inside MDASH, with OpenAI models still reserved for the hardest cases.
Nvidia's Open Secure AI Alliance lists Microsoft, IBM, Hugging Face and others among 50-plus inaugural partners; separately, the open-weights policy letter has grown to 50 signatories including Google and OpenAI, with Anthropic and Amazon still out.
A reported 81.5% of Samsung foundry staff want out within two years as rival SK Hynix pays bonuses more than three times larger on the back of AI memory profits.
A Vectoral investigation maps the 'relay' proxy ecosystem, where abused free trials and hijacked endpoints feed a gray market for cut-price AI tokens.
Clem Delangue urges OpenAI to publish the traces of the agents that breached Hugging Face and asks for $100M in compute to build open security defenses.
The Chinese lab halted its second raise near a $71B valuation after a leaked four-hour investor briefing by founder Liang Wenfeng went viral.
Library workshops teaching people how AI works — and how to switch it off — are drawing waitlists and viral engagement, a data point on consumer AI fatigue.
Alibaba has released Qoder Mobile across Android, iOS and HarmonyOS, letting developers steer its coding agents and cloud tasks from a phone with feature parity across platforms.
Anthropic's Opus 5 approaches Fable 5-class results and triples the field on ARC-AGI-3, all at unchanged Opus pricing and lower token cost.
Cognition acquired the messaging assistant Poke in a low-nine-figure deal, betting that agent personality and interaction quality are becoming competitive moats.
Midjourney promoted V8.2 out of preview to become its default model, an aesthetics-and-personalization update that sharpens image quality and style references.
The July 24 open-weights letter has expanded beyond Nvidia, Microsoft and Meta, with updated signatories now including Google and OpenAI while Anthropic remains absent.
OpenAI and Google have joined the expanding open-weights appeal to Washington, making Anthropic the most visible major frontier-lab holdout.
StepFun says it is open-sourcing its Attention-FFN disaggregation work with the vLLM team, Ant Group and FastAFD, pushing a key MoE-serving efficiency technique into the community.
Shanghai humanoid maker AgiBot has formally begun its Hong Kong listing process, backed by world-leading 2025 shipments and a valuation target near US$6 billion.
AMD put its Helios rack into production — 72 2nm MI455X GPUs plus 256-core Venice CPUs per rack — with OpenAI, Microsoft, Meta and Oracle lining up.
OpenAI's health service exits its waitlist: US adults can link Apple Health and hospital records, while critics flag HIPAA gaps and a free-tier quality gap.
DeepSeek's flagship V4 line exits preview: legacy chat and reasoner endpoints shut down today as new peak and off-peak API pricing takes effect.
Alibaba Cloud says its Zhenwu M890 supernode, built on in-house T-Head chips, now runs 2.4T-parameter Qwen3.8 in production, cutting inference costs over 40%.
Alphabet posted $119.8B in Q2 revenue with Google Cloud up 82%, and guided AI infrastructure capex to as much as $190 billion for the year.
AMD will invest up to $5 billion in Anthropic, which plans to deploy 2 gigawatts of Instinct MI450-series GPUs for Claude starting in early 2027.
Cursor's new request-level router sends each coding task to the best-fit model, with early enterprise users reporting 30-50% savings and no quality drop.
Google will supply DOE labs with AlphaFold 3, AlphaEvolve, WeatherNext and Gemini seats as its contribution to the US drive to double scientific discovery.
OpenAI confirms its first self-built data center: a 3.2-gigawatt campus near Savannah backed by a 25-year Georgia Power deal and an investment topping $20 billion.
After 'unlimited' access drew some 19,000 users in about 45 days, the Army's Ask Sage-powered LLM workspace ran dry and usage caps are back.
Xiaomi will end MiMo Code's free phase at 6 p.m. Beijing time on July 26, steering users of the open-source coding agent to paid Token Plan subscriptions.
The French lab expanded its Microsoft partnership to build sovereign European AI infrastructure, while Samsung reportedly weighs up to €1B at a €20B valuation.
OpenAI launched 'Advertise in ChatGPT,' a self-serve ads manager targeting users by conversational context, with Best Buy, Lowe's and VistaPrint among early buyers.
OpenAI's new platform packages enterprise voice and chat agents with guardrails and Codex-driven tuning — it already resolves 75% of OpenAI's own support calls.
Tencent is said to be consolidating the teams behind its two OpenClaw-era AI agent products, ending an internal race and unifying its agent strategy.
A hierarchy of planner and worker agents reimplemented SQLite from its 835-page manual alone, passing a held-out test suite of millions of queries for $1,339.
Moonshot AI halted new Kimi K3 subscriptions after demand surged sixfold, as third-party tests put the open-weight model within reach of Claude Fable 5.