⚡ AI Focus Bulletin
AIGC

HeyGen Adds Audio to Video: One API Call Turns Voice Into Video

HeyGen's new Audio to Video feature turns a spoken script plus a single image into a finished video through one API call, with batch support for bulk output.

AI video company HeyGen has launched Audio to Video, a feature that flips the usual avatar-video workflow: instead of typing a script for a synthetic voice to read, creators speak the script themselves. The company announced the capability on X late on July 22, pitching it in three beats — "your voice plus an image becomes a video," a single API call does the conversion, and the same image can be paired with many narrations in batch to mass-produce variants.

Voice-first, API-first

The input pairing is deliberately minimal: one audio track and one still image are enough to produce a talking video, with HeyGen's models handling lip sync and speech-driven animation. That removes two steps that have defined avatar tools so far — writing a script and configuring a text-to-speech voice — and preserves the timing, emphasis and personality of a real recording, which synthetic narration still struggles to match. HeyGen's developer documentation lists an Audio to Video API in beta alongside its existing video-generation endpoints, which have long accepted audio files as a voice source for pre-built avatars; the new feature extends that to arbitrary images in a single call.

The batch angle is aimed squarely at programmatic use. HeyGen's example — one image, a hundred narrations — describes localization, personalized outreach and A/B-tested ad creative, workflows where the visual stays constant and only the voice track changes.

Why it matters

The launch continues the migration of AI video from creative tool to infrastructure. HeyGen has been building out its developer platform against rivals such as Synthesia and D-ID, and a one-call voice-plus-image endpoint is the kind of primitive that gets embedded into other products — CRMs, learning platforms, marketing automation — rather than used in an editor. It also lowers the floor for non-writers: anyone who can talk into a phone can now generate presenter-style video at scale, which will accelerate both legitimate content pipelines and the ongoing debate about synthetic media provenance.

Sources