Universal Music and ElevenLabs Form AI Music Partnership
Universal Music and ElevenLabs announce a multi-year partnership to develop licensed AI music tools with artists, labels and publishers.
Universal Music and ElevenLabs announce a multi-year partnership to develop licensed AI music tools with artists, labels and publishers.
Tencent Hunyuan and research partners have released AuK, an open-source foundation model for controllable speech and audio generation and editing.
OpenAI has released Images 2.5 across ChatGPT and its API, combining faster generation with stronger reference fidelity and multi-turn editing.
Google has moved its latest music model into global consumer and developer channels, widening access beyond its earlier Flow Music launch.
MAI-Image-2.6-Flash pairs lower prices with production-focused editing, while the full model leads a major third-party benchmark.
The research preview produces continuous 720p audiovisual simulations at 24 fps that respond to text, camera and player controls.
A lightweight MiniMax experiment turns H3’s learned video dynamics into an interactive environment using only 8,000 samples.
Atlas combines generation, 3D reconstruction and simulation in one architecture aimed at filmmaking, robotics and virtual worlds.
The research system renders interactive 720p interfaces frame by frame, extending generative video into software and agent training.
Baseten says an agentic system generated, tested and deployed production optimizations across image and language models.
Google’s production-ready video model adds 40-second extensions, keyframe interpolation, low-cost previews and output upscaling to 4K.
Midjourney’s first V8.2 editing model adds instruction-led changes, four-image references, inpainting and canvas expansion.
The design model can derive a reusable visual style from one to ten reference images without requiring a separate training process.
EA and three global music groups participated in Stability AI’s Series B, tying its next phase more closely to licensed creative production.
Gradio’s new Workflow interface turns typed AI pipelines into inspectable canvases, callable REST endpoints and deployable applications.
Runway’s new model converts uploaded or generated SDR footage into high-bit-depth HDR formats designed for professional post-production.
Tripo’s new 3D preview generates native quad topology with up to 25,000 quads, targeting characters and real-time production assets.
The generative-media company has secured a major round to scale its video and image platform amid intensifying AIGC competition.
Meshy’s new technical report explains how its latest system more faithfully transfers a reference image into usable 3D geometry.
Users in most markets can remove visible marks from generated images, video and music, while Google retains SynthID and C2PA provenance; legally required marks remain in some jurisdictions.
Suno is turning its music generator into an editable production environment, giving creators finer control over AI-made tracks.
LTX’s new open-weight world model adds adaptive rendering, native 4K HDR output and stronger multi-shot control for local video production.
Alibaba's Wan team released a command-line tool that lets coding agents call its image and video models programmatically, installed by pasting one command.
Microsoft AI's newest image model entered Arena's text-to-image leaderboard at second with 1336 points, up from 2.5's tenth place and 45 behind GPT Image 2.
SpaceXAI's new image model entered both Arena leaderboards at No.2 behind GPT-Image-2; by 11 August it had been pushed to third in text-to-image by Microsoft's MAI-Image-2.6, while holding second in image editing.
The Qwen-Audio-based product bundles dictation, voice agents and audiobook generation, and ships free on iOS, Android, Mac and Windows.
ByteDance switched on public API access for its 30-second video model; Luma had it live a day earlier and Runway shipped it on the day, with Krea, Lovart, Pika and others announcing integrations still rolling out.
The music generator will embed tamper-resistant audio watermarks, add fingerprinting, restrict bulk downloads and license Musixmatch's Sentinel for lyric copyright detection.
Black Forest Labs' 20-second video model with native audio opened to the public on Runway, Krea and Replicate after a gated rollout.
JD.com released Apache-2.0 weights for JoyAI-Video-Edit, which edits 720p video at about 30 frames per second on a single NVIDIA B200.
Beijing lab Sand.ai released Apache-2.0 weights for a 114-billion-parameter sparse video model that generates 1080p clips with synchronized audio.
Video Arena placed MiniMax-H3 first among open models in both text-to-video and image-to-video, some 280 Elo points clear of the next open weights entry. The weights are now public, but the licence restricts open-weight use in the US, EU, UK and South Korea.
Tencent's new speech model reports about 3% word error rates in Mandarin, English and Cantonese, and is live on Tencent Cloud and in Yuanbao.
Wan launched a Real-time feature suite whose first capability turns a live phone camera feed into stylised imagery, starting with anime looks.
MiniMax published downloadable weights for H3, its omni-modal 2K video model, making a top-ranked video generator locally runnable for the first time.
SB 942 became operative on 2 August, forcing large generative-AI providers to embed provenance data and publish a free detection tool.
From 2 August the AI Act's Article 50 transparency duties bite across the EU: machine-readable marking, chatbot disclosure and deepfake labels.
Google rolled back a Nano Banana 2-powered generator that let users paint fabricated imagery onto satellite maps, citing policy-violating outputs being shared online.
Munich Regional Court I ruled that Suno infringed copyrights in both training and generated music, rejecting the company's US-style fair use defence. Suno says it disagrees and is weighing an appeal.
Snap has adjusted Spotlight recommendations to exclude entirely AI-generated videos while still allowing AI as an editing and enhancement tool.
P-Image-Ideogram launches with four quality tiers, native 1K and 2K output and three-to-eight-second latency, aimed squarely at production image pipelines.
ByteDance's Seedance 2.5 left closed enterprise beta on Friday, launching on Jimeng in China and Dreamina globally with 30-second single-pass generation.
Alibaba's Wan team launched a Realtime mode on wan.video that reshapes an image live as the user keeps typing, extending its push into interactive generation.
Ideogram shipped a dedicated object-removal model that erases shadows and reflections too, and says it beats Nano Banana 2, FLUX Erase and GPT Image-2.
Google launched Lyria 3.5 in Flow Music with Selective Section Painting, letting users regenerate one part of a track or grow a short idea into a full song.
Two new speech-to-text models replace Whisper as OpenAI's recommended default, with steerable context inputs and pricing from $0.0045 per minute.
European nonprofit AI Forensics found seven of the nine most popular image-editing models on Hugging Face complied with simple undressing prompts.
Midjourney promoted V8.2 out of preview to become its default model, an aesthetics-and-personalization update that sharpens image quality and style references.
Runway now lets creators build, run and edit its node-based Workflows through plain-language prompts inside Runway Agent, collapsing pipeline-building into conversation.
Black Forest Labs' new flagship generates 20-second clips with native audio, claims wins over Kling and Runway, and promises open weights later in 2026.
Chinese video startup PixVerse released VibeMV on web, Android and API, generating rhythm-synced, subtitle-ready music videos from one uploaded song.
ElevenLabs' new Vocals and Styles features fine-tune its Music v2 model on an artist's own tracks, generating songs in their voice and sound with copyright screening.
HeyGen's new Audio to Video feature turns a spoken script plus a single image into a finished video through one API call, with batch support for bulk output.
The avatar-video company's new mode has AI agents pitch creative angles and share storyboards mid-flight, shifting generation from one-shot orders to direction.