For the OneSixthual dolls I need to assign each doll a particular voice, which fits to its character. I have experimented with several online tools, and here are my results.
Vidnoz for TTS Voices
I have used Vidnoz since 2 years, it was my first venture into AI-generated animations. At first I used their avatar mode: uploaded image of a doll, then made the doll speak by selecting a voice and added the text prompt. The result was then that the avatar spoke.
This approach had the disadvantage that the avatars were quite unflexible, a bit stiff, and only the head moved together with the speaking. Mouth movement was OK.
So then I devised another more complicated method: Vidnoz also has a mode where audio-only can be generated, without the video. I used this to select from the voices the ones I liked, then entered the text prompt. The resulting wave file I then merged with DreamFace to an existing video snippet.
Recording the voice
And one can filter by language. It is actually great fun to select a language other than English, e.g. French, and then enter an English text - the voice then speaks English with an accent!
And one can customize each voice:
A great feature: one can preview the voice recording without needing to pay - this allows experimentation before the paid final version is recorded.
Using DreamFace for merging video with audio
The app DreamFace costs $30 per year, and credits can be purchased: 30 cr for $5, 1000 cr for $100. 5cr are added automatically weekly, but they expire after 1 week. I forgot how many credits a merging of the video with the audio costs... So I upload the video snippet, then the audio wave file, and I let the app do its magic.
There are a few issues:
- DreamFace does not support NSFW videos, and it refuses the upload of a video which does contain such content. Therefore, one must only upload just a video close-up of the face that i supposed to talk.
- Multiple faces in a video are not supported. There must only be one clearly visible face.
- When the audio is longer than the video, the video at its and is added playing backwards. This ensures a smooth overall video, but it makes the video unusable in a context where its ending is pre-determined. Therefore, the video snippet should always be longer than the audio snippet.
Dreamface also does have its own voices, and they appear to have a great selection. But I have not used them yet, as I like the choices offered by Vidnoz.
Creating the final video
I then paste / merge the resulting video+audio into my overall video timeline, overwriting the previously non-speaking part with this new speaking part. For this I am currently using Microsoft ClipChamp, in the free version. Works fine for me.

