ElevenLabs Dubbing v2 Adds Accent and Audio Controls for Video Localization
ElevenLabs introduces a rearchitected AI dubbing model that preserves original emotion and timing across more than 90 languages.

Stock photo for illustration only, not from the actual event
- ElevenLabs launched Dubbing v2 supporting translation across over 90 languages
- Preserves original speaker emotion, delivery cadence, and timing accuracy
- Adds localized accent controls for regional dialects like Castilian and Latin American Spanish
- Manages multi-speaker scenes and background audio with upcoming API support
ElevenLabs has introduced Dubbing v2, a rearchitected AI dubbing model built to retain a speaker's emotion, delivery, and timing while translating content across more than 90 languages. The update expands the company’s localization proposition beyond a basic language replacement: its documented controls cover dialect-specific accents, multi-speaker material, and background audio management for more complex video and audio scenes.
According to ElevenLabs’ Dubbing v2 announcement, the model is designed for creators, marketers, studios, and broadcasters that need to localize video at production scale. It is integrated with ElevenCreative for one-click video localization and ElevenProductions, the company’s professional localization service.

Stock photo for illustration only, not from the actual event
The central goal is preserving the original performance rather than simply generating translated speech. That matters for material in which pacing, vocal emphasis, and emotional delivery are part of the message, including marketing campaigns, creator videos, and professionally produced programming. Dubbing v2 is intended to synchronize translated dialogue with the source speaker’s timing and delivery across its supported languages.
The evolution of AI voice dubbing increasingly focuses on emotional nuance and cultural context rather than literal translation. Models capable of separating background noise and applying regional accents significantly reduce traditional post-production friction, enabling smaller creators and studios to deploy international content efficiently.
The most practical additions are the controls documented for ElevenLabs' dubbing workflow. They give teams more ways to shape a dub around the source material and the intended audience, particularly when a project includes regional language variation or a mix of dialogue and sound.
For Spanish-language localization, the accent capability is especially relevant. ElevenLabs highlights sharper locale accuracy for regional accents such as Castilian Spanish and Latin American Spanish. The documented target_accent field provides a mechanism for applying dialect-specific accents, although ElevenLabs labels that control experimental. Teams should therefore treat it as a useful production option that still merits review against their own editorial and brand requirements.
The update also addresses a common limitation in automated dubbing: real source media rarely consists of one clean voice track. Interviews, shows, advertisements, and social content can include multiple speakers, music, effects, and ambient audio. The num_speakers control and separate foreground and background audio inputs indicate that Dubbing v2 is designed to accommodate these more complicated inputs. The drop_background_audio option gives users a documented way to remove background sound when that is appropriate for the output.
Source: Dev.to
Found something wrong in this article? Report an issue with this article
Comments
Leave a Comment