17/01/2023
Artificial intelligence (AI) and machine learning technologies will likely play a crucial role in the future of dubbing films. Multiple procedures are required to dub a movie, including eliminating the original dialogue, translating the script, imitating the actors' voices, synchronizing the video with the new audio, and verifying that the tone, accent, and rate of speech are identical to the original.
Step 1: Eliminating script words from a movie using the original audio file and the written script as inputs.
Reconstructing audio without the original text for dubbing is now a laborious and time-consuming process. AI might be used to search for and eliminate specific words from an audio track without affecting the background sounds. This would be a big improvement over the existing method, which includes copying and pasting portions of background noises from silent moments when dialogue is present.
Step 2: Translate the film script (why not use AI?).
The translation services of today are quite competent, but they will continue to improve until they reach perfection. Using AI, the process of translating a film script will become more precise and efficient. In around five years, it is predicted that this technique will be mastered by humans.
Step three involves imitating the actor's voices and pronouncing the translated script.
This is accomplished by using the voices of the actors in the original audio to train an algorithm to imitate their voices, and then feeding the translated script to the new voice to produce the translated audio track. The startup Lyrebird can already mimic your voice if you feed it a few sentences, but it is still a proof-of-concept and not accurate enough for this type of application - yet.
Step four involves applying the created phrases to the original tape while ensuring that the intonation, accent, speed of speech, etc. are identical.
This is the most difficult step in the entire procedure. Obtaining the timestamps of every beginning and finish of a sentence in the original file and using them to fit the identical translated sentences in the new track is one possibility; however, this would require significant adjustment to sound realistic.
Step 5 involves synchronizing the video to the new audio track.
This includes producing extra frames to adapt the length of speech in different languages and use existing deep-fake video technology algorithms to synchronize the performers' lips with the new track. This is the final step in ensuring the accuracy of the dubbing, and it ensures that the translated film is identical to the original.
While the process of dubbing a film using AI is still in its infancy, this technology has the potential to completely transform the business. With the breakthroughs in AI and machine learning, future dubbing of films is anticipated to be significantly more accurate and efficient. Additionally, this technology might be used to subtitle media on platforms such as YouTube, making it easy to watch videos in any language.