Six voices, every one consensual, every one paid for. Why the harder path was the only path.
When you train a voice the easy way, you scrape. YouTube monologues, audiobook torrents, podcast feeds. You hide what you took, you wave at the synthesizer, you ship.
We chose to not do that. Not because we’re virtuous — because the voices we built that way kept being somebody. Recognizably somebody. And that somebody never said yes.
What we did instead
We paid six voice actors. We signed real contracts. We gave them an unusual one: AGirl owns the voice model, the actor owns their identity, and any rendered output that resembles their natural voice gets royalties forever.
Then we recorded — not the standard “read 300 phonemes” — but emotional content. Stories. Whispered lines. Loud lines. Tired lines. Lines they made up. Twenty hours per voice. The conversational training data alone took six months.
Why it works
The downside is obvious: six is a small number. We can’t scale to a hundred voices the way a scraping shop can.
The upsides are less obvious but more important.
- No model drift. When a base model updates, our voices don’t accidentally start sounding like a Marvel actor. The training corpus is locked.
- Real emotion. Synthesized voices trained on read-aloud content sound like audiobook narrators. Ours sound like people. The difference is enormous and you hear it within the first five seconds.
- Defensible. No lawsuits coming. We sleep at night.
The hard part
The hard part is that we have to be careful with how the voice changes over time. We retrain quarterly with new sessions from the same actors. Each new voice version is A/B tested against the old one — and we ship only if listeners prefer the new voice without being able to articulate why.
It’s slow. It’s expensive. It’s the entire point.