Published:
Tagged: AI Voice Python Self-Hosting Presentations
TLDR; Many open voice-cloning models don’t allow commercial use, so read the licence before you listen. Of the ones that do, Qwen3-TTS (Apache-2.0) sounded best on my voice, and Chatterbox (MIT) is a smaller download with fewer moving parts. For my voice, preparing the ten-second sample mattered at least as much as choosing the model: level it, end it in a pause, and give Qwen3-TTS its transcript. Then mark everything you make as synthetic.
I can record my talks now without the need of any specialised software package. Sipario, the deck engine I wrote about last month, has a recorder in its notes window, in a future version: press V, give the talk, and it keeps my voice along with the moment each slide appeared, then makes a video of the slides over the audio.
The problem is what happens next. I change one slide’s script and the recording is out of date, and I’m not re-recording a whole talk for one paragraph. So I wanted the machine to read the changed paragraph in my voice, on my own laptop, because a model of my voice is the last thing I want sitting on somebody else’s server. It’s not quite there yet: sipario can read a whole talk’s script in my voice, and putting a changed paragraph back into my original recording comes next.
The licences ruled out three of the models I’d read about before I’d heard a single clip. XTTS v2 clones a voice from six seconds of audio, but its licence “allows only non-commercial use of a machine learning model and its outputs”, and “outputs” means the audio you make with it. F5-TTS is MIT for the code and CC-BY-NC for the trained models, because of the data they were trained on. Fish Speech puts code and weights under one research licence that grants no commercial rights.
My talks are part of my work, so that left three: Chatterbox from Resemble AI (MIT), Qwen3-TTS from Alibaba’s Qwen team (Apache-2.0), and MeloTTS with OpenVoice on top (both MIT). Licences change, so check the current ones yourself rather than trusting my list.
Chatterbox went first, and it sounded a bit like me but too American. I think I have a slight British accent 😀.
Here’s me giving the talk, and then Chatterbox as installed saying the same words in my cloned voice, starting “For the last 18 months, agents have been changing who reads our docs”. Every clip in this post is at the same loudness, and every synthetic one carries the watermark described below.
Me, recorded giving the talk.
Synthetic: Chatterbox as installed, cloned from ten seconds of the same recording.
It’s hard to improve on “a bit like me” without something to measure, so I tried two ways of measuring it:
say voices it gave the expected labels: “england” for me, and “england” and “us” for the Mac’s British and American voices. On clones it couldn’t be trusted. It rated clones as English that I could hear were American, and one clip flipped between two near-identical generations.I made 117 variants, every one of them commercially usable, each reading the same three passages of the talk. They came from three families of models: five from Chatterbox (the original, Turbo, Nano and two multilingual ones), Qwen3-TTS in two sizes, and MeloTTS with OpenVoice. Then I picked five of them to listen to. They aren’t the five highest scores: one is Chatterbox as installed, for comparison, and the others are the best I got from Chatterbox, from Chatterbox Turbo and from Qwen3-TTS with a ten-second sample, plus a Turbo variant that talked at my pace.
I also tried it the other way round: start with a voice that already has a British accent, then make it sound like me. MeloTTS has a British English speaker, and OpenVoice shifts its tone towards a reference voice. The classifier called it English every time, and it scored 0.26 against me, closer to a stranger than to me. It had the accent, but it wasn’t my voice.
| Model and settings | Similarity | Pace vs me | |
|---|---|---|---|
| 1 | Chatterbox, as installed | 0.77 | 1.51 |
| 2 | Chatterbox, sample peak at -3 dBFS, exaggeration=0.3 |
0.85 | 1.21 |
| 3 | Chatterbox Turbo | 0.87 | 1.13 |
| 4 | Qwen3-TTS 1.7B Base, sample levelled and transcribed | 0.85 | 1.29 |
| 5 | Chatterbox Turbo, slower settings | 0.82 | 1.00 |
Both numbers are averages over the three passages. Similarity is each clip’s score against the average of my three real recordings, and pace compares how fast the clone talks with how fast I say the same words, so 1.51 means 51% faster. Rows 1 and 2 cloned from the ten seconds sipario cuts automatically, and 3 to 5 from ten seconds I picked by hand.
Here are 2 to 5 saying the same words as the two clips above, so you can listen the way I did:
Synthetic, row 2: Chatterbox, sample peak at -3 dBFS, exaggeration=0.3.
Synthetic, row 3: Chatterbox Turbo.
Synthetic, row 4: Qwen3-TTS 1.7B Base, the one I picked.
Synthetic, row 5: Chatterbox Turbo, slower settings.
I listened to all five against a recording of myself, at the same loudness, and picked 4. By similarity Qwen3-TTS is level with 2 and behind 3, but IMHO it sounds the closest to how I speak, given I’m not a native English speaker. My accent and tone are replicated pretty well. The scores didn’t tell me which accent I’d hear, so I made the final choice by listening.
Every clone but 5 talked faster than I do, and I wondered whether that’s what made them sound American. It doesn’t seem to be the only reason: on the one passage I listened to, 2 and 3 ran at my pace, and 4, the one I picked, was 18% faster than me, against the three-passage averages in the table.
This is one experiment on one voice, so read the numbers with that in mind. Most of them are single runs, and running the same settings again moved a score by as much as 0.035. That’s more than the 0.02 between the three models’ best scores (0.85 to 0.87), which also came from different samples and settings, so they don’t compare like for like. The changes below are fairer comparisons, because each one changes a single thing:
| What changed | Model | Before | After |
|---|---|---|---|
| A transcript of the sample | Qwen3-TTS, 30 s sample | 0.70 | 0.80 |
| Sample levelled to -20 LUFS | Qwen3-TTS | 0.82, 0.81 | 0.84, 0.85 |
| Sample’s loudest moment brought to -3 dBFS | Chatterbox | 0.77, 0.80 | 0.83, 0.81 |
Then exaggeration from 0.5 to 0.3 |
Chatterbox | 0.83, 0.81 | 0.85, 0.84 |
Where there are two numbers, they’re two runs with different seeds. For my voice, preparing the sample mattered at least as much as choosing the model.
Record it raw. Browsers clean up a microphone by default, with echo cancellation, noise suppression and automatic gain. That’s right for a call, but for a sample I wanted the model to hear my voice rather than the browser’s processing, so sipario’s recorder sets echoCancellation, noiseSuppression and autoGainControl to false.
Ten seconds is enough. It worked well for me with both models. In Chatterbox’s source, the stage that turns speech tokens into sound uses the first ten seconds of a sample (DEC_COND_LEN = 10 * S3GEN_SR) and the speech prompt the first six, although its speaker encoder hears all of it. Qwen3-TTS cloned me about as closely from ten levelled seconds as from thirty, 0.85 against 0.86. What matters is ten seconds of you talking without a long pause, and ffmpeg can find them:
ffmpeg -i talk.webm -af silencedetect=noise=-35dB:d=1 -f null -
Anything quieter than -35 dB for a second or more counts as a pause. I started at 0.4 seconds, and a real talk had no ten-second stretch without one, because you stop for breath every few seconds.
Level it. My laptop’s microphone, across the room, recorded me at about -44 LUFS (the loudness unit broadcasters use). Both models cloned closer from a louder sample: bring it up to -20 LUFS for Qwen3-TTS, and bring its loudest moment up to -3 dBFS for Chatterbox. One ffmpeg line cuts and levels a sample, with -ss where your ten seconds start:
ffmpeg -ss 203.5 -t 10 -i talk.webm -af loudnorm=I=-20:TP=-2 -ar 24000 -ac 1 sample.wav
Give it the words. Qwen3-TTS can clone from the sound alone, and its README warns that “cloning quality may be reduced” if you do. In my tests it was, by more than anything else I changed (the first row of the table above). faster-whisper (MIT) makes the transcript on the CPU.
End it in a pause. A sample cut in the middle of a word gets transcribed with the whole word, and Qwen3-TTS then says that word before your line. My sample ended inside “explicit”, and this came out:
Explicit, nothing you hear in this clip was ever sent to a server.
Synthetic: Qwen3-TTS. The line it was given starts at "Nothing".
In sipario, two of a test deck’s seven scripted steps opened the same way. End the sample before that word and it goes away.
I ran everything below on an Apple M4 Max with 128 GB of memory, and haven’t measured anything smaller. You’ll need uv (brew install uv) and ffmpeg, and room for the models: the first run downloads about 6 GB for Qwen3-TTS, 4.5 GB of it the model, or 3.2 GB for Chatterbox.
Qwen3-TTS, with the versions I measured:
uv venv --python 3.12 ~/voice-clone
VIRTUAL_ENV=~/voice-clone uv pip install qwen-tts==0.1.1 \
torch==2.14.1 torchaudio==2.11.0 transformers==4.57.3 \
faster-whisper==1.2.1 av==18.1.0 resemble-perth==1.0.1 'setuptools<81'
Two of those pins matter whatever else changes. av==18.1.0, because av 19 dropped an argument faster-whisper 1.2.1 still passes, and transcription stops with TypeError: open() got an unexpected keyword argument 'metadata_errors'. And setuptools<81, because the watermarker imports pkg_resources, which setuptools 81 removed. That torch has wheels for Apple Silicon on macOS 14 or later, Linux and Windows, and none for an Intel Mac.
Save this as clone.py, next to your sample.wav:
import os
import sys
import numpy as np
import perth
import soundfile as sf
import torch
from faster_whisper import WhisperModel
from qwen_tts import Qwen3TTSModel
SAMPLE = "sample.wav"
TRANSCRIPT = "sample.txt"
# Qwen3-TTS clones much closer when it knows the words in the sample. The
# first run writes them to sample.txt and stops, so you can check them.
if not os.path.exists(TRANSCRIPT):
whisper = WhisperModel("medium.en", device="cpu", compute_type="int8")
segments, _ = whisper.transcribe(SAMPLE, language="en", beam_size=5)
with open(TRANSCRIPT, "w") as f:
f.write("".join(s.text for s in segments).strip() + "\n")
print(f"Check {TRANSCRIPT} against {SAMPLE}, fix anything it got wrong, then run this again.")
sys.exit()
words = open(TRANSCRIPT).read().strip()
device = ("cuda" if torch.cuda.is_available()
else "mps" if torch.backends.mps.is_available() else "cpu")
model = Qwen3TTSModel.from_pretrained(
"Qwen/Qwen3-TTS-12Hz-1.7B-Base", device_map=device,
dtype=torch.float32 if device == "cpu" else torch.bfloat16,
attn_implementation="sdpa")
voice = model.create_voice_clone_prompt(ref_audio=SAMPLE, ref_text=words)
marker = perth.PerthImplicitWatermarker()
for i, line in enumerate(sys.argv[1:], 1):
wavs, sr = model.generate_voice_clone(
text=line, language="English", voice_clone_prompt=voice)
# Qwen3-TTS doesn't mark what it makes, so mark it here
wav = marker.apply_watermark(np.asarray(wavs[0], dtype=np.float32), sample_rate=sr)
sf.write(f"qwen-line-{i}.wav", wav, sr)
print(f"qwen-line-{i}.wav: {len(wav) / sr:.1f}s")
Run it once to get the transcript:
~/voice-clone/bin/python clone.py
It writes what it heard to sample.txt and stops. Listen to the sample and check sample.txt word by word, especially the last one: if it’s a word you can’t hear whole, cut the sample shorter and delete sample.txt to start again. Fix any other word it got wrong. Whenever you replace or change sample.wav, delete sample.txt and repeat the transcript step. Then run it with the lines you want it to say:
~/voice-clone/bin/python clone.py "The first line to say." "And a second one."
It warns that SoX isn’t installed, and runs fine without it. Once the models are downloaded, it takes about seven seconds to load and then about 0.85 seconds of work for each second of speech. On the CPU, one 13-second line took 27 seconds.
Chatterbox is a single package:
uv venv --python 3.11 ~/voice-clone-cb
VIRTUAL_ENV=~/voice-clone-cb uv pip install chatterbox-tts==0.1.7 'setuptools<81'
Save this as clone_chatterbox.py, next to sample-raw.wav, which is the ffmpeg cut from above without the loudnorm filter:
import sys
import numpy as np
import soundfile as sf
import torch
import torchaudio
from chatterbox.tts import ChatterboxTTS
# Bring the sample's loudest moment to -3 dBFS
raw, rate = sf.read("sample-raw.wav", dtype="float32")
sf.write("sample-peak.wav", raw * 10 ** (-3 / 20) / np.abs(raw).max(), rate)
device = ("cuda" if torch.cuda.is_available()
else "mps" if torch.backends.mps.is_available() else "cpu")
model = ChatterboxTTS.from_pretrained(device=device)
model.prepare_conditionals("sample-peak.wav", exaggeration=0.3)
for i, line in enumerate(sys.argv[1:], 1):
wav = model.generate(line, exaggeration=0.3)
torchaudio.save(f"chatterbox-line-{i}.wav", wav, model.sr)
print(f"chatterbox-line-{i}.wav: {wav.shape[-1] / model.sr:.1f}s")
~/voice-clone-cb/bin/python clone_chatterbox.py "The first line to say." "And a second one."
Without setuptools<81, Chatterbox stops at its first line with TypeError: 'NoneType' object is not callable rather than speak without its watermark. It took about 1.3 seconds of work for each second of speech, so a 25-minute talk takes over half an hour to voice.
One good clip isn’t enough when you’re voicing a whole talk. These are the four problems sipario has to deal with:
loudnorm, brings it to -20 LUFS with one gain, and holds the peaks at -2 dBFS with alimiter.max_new_tokens=1000 in its source). Split long text at sentence ends: sipario sends Chatterbox at most 300 characters at a time. Qwen3-TTS said a 1581-character paragraph in one go, and sipario still splits at 800.Every clip both setups make carries Perth, Resemble AI’s imperceptible watermark, and Perth’s detector can find it later, so these clips can be told apart from a real recording. It only finds its own mark: it can’t tell you whether some other audio is synthetic. Chatterbox applies it to everything and has no switch for it. Qwen3-TTS marks nothing, which is why the script above does it. Perth’s detector reads 1.0 on every clone in this post, after their conversion to MP3, and 0.0 on the recording of me. The Qwen3-TTS clip among the five came from an experiment before I’d added the mark to that setup, so I marked it for this post. The mark also survived the AAC audio in sipario’s exported video.
Then there’s consent. Clone your own voice and nobody else’s. Sipario enforces what it can: it only clones from a recording made in its own notes window, never from a file handed to it, and never from one of its own voiced recordings. What it can’t check is who’s speaking into the microphone, or whether what the microphone heard was a real person at all: synthetic audio played out loud would get recorded like anything else.
sipario voice reads a deck’s script in my voice, slide by slide, and makes a recording the video export treats like one I made myself. It isn’t released yet. The first full run was the whole of Your Next Customer Is a Robot, 73 steps and a 21:52 video, with the picture and the voice ending together. Next, I want it to re-read only the slides whose script has changed and drop them into my original recording of the talk.
If you try this yourself, check the licence first and put most of your effort into the ten-second sample. The numbers will get you to a shortlist, but pick the model by listening to it.