Articles · In this section
NOMOTO MEDIA

I’ve Been Building a Voice Factory With AI.

By Niklas S. Osterman

For about a year I’ve been circling a single obsession: make ideas audible. Not in the vague, “AI will change everything” way—literally audible. Words becoming voices, voices becoming episodes, episodes becoming archives. A kind of home studio that doesn’t just record audio, but manufactures it—repeatably, cleanly, at scale—without turning the process into soulless sludge.

’ve lived long enough inside institutions to know what happens when tools become systems. Systems choose defaults. Defaults become policy. Policy becomes culture. And culture becomes a cage. That’s part of why I’ve been pushing so hard on this: because we’re at the stage where the “voice layer” of AI is becoming real, and if ordinary people don’t shape what it becomes, corporations and governments will.

I’m not saying that as a conspiracy. I’m saying it the way you say the weather is coming.

The Work: Turning Text Into Finished Audio Like It’s a Print Job

A lot of people think “AI voice” is just clicking a button and getting a clip. That’s not what I’ve been building.

I’ve been building a pipeline. A factory.

Text in. Finished MP3 out. Not just a sample—something you could publish. Something you could batch. Something you could iterate on. Something you could do again tomorrow without reinventing the wheel.

Over time, I’ve built workflows that:

  • take long scripts (books, essays, episodes),
  • chunk them safely,
  • render them to voice through an API,
  • stitch segments together without pops and timing bugs,
  • normalize output and control pacing,
  • and then append “real world” production pieces: intros, outros, marketing bumpers, music beds.

It’s not glamorous. It’s the plumbing. And plumbing is what turns “cool demos” into actual work.

The most exciting part hasn’t been the narrator voice. Narration is the easy win—one voice, one tone, long-form pacing.

The real breakthrough is casting.

That’s when the system stops being “text-to-speech,” and starts acting like a stage: multiple characters, each with a distinct presence, each with consistent delivery, each with their own personality. Not just changing the voice name—changing the performance.

I’ve been testing a casting setup where each character gets:

  • a voice choice,
  • a speed,
  • a style (documentary vs theatre, etc.),
  • and most importantly: an instruction file.

That instruction file is where it becomes theatre.

You describe a character the way a director does: pacing, tone, emotional temperature, how they end sentences, what they do before reveals, how they handle urgency, whether they “punch” verbs, whether they linger. It’s not about faking accents for cheap laughs. It’s about giving a voice a believable intention.

And here’s the truth: when it works, it’s uncanny. You can hear the difference. You can hear a persona, not just a waveform.

The Human Layer Is the Product

I’m not doing this because I want to flood the world with content. I’m doing it because I want to preserve the human layer.

AI is getting better at voice—fast. But the danger is that “better” becomes “default.” And default becomes a style: smooth, corporate, neutral, unthreatening. A voice designed to offend nobody and therefore move nobody.

That’s why I keep pushing characterization. Acting. The messy specifics. The human fingerprint.

Because the future won’t be decided by who has the best model. It will be decided by who defines the use of the model—who builds the interfaces, the workflows, the creative habits, the cultural expectations.

The reason I care about this isn’t technical. It’s personal.

After brain surgery, the idea of going back into healthcare like nothing happened doesn’t feel realistic. And honestly, I don’t want to spend what time I have left inside systems that pretend people are simple.

I want to make things. I always have.

I’ve built creative businesses before. I’ve done design. I’ve done writing. I have a Master’s in Dramatic Arts. I’ve made podcasts that nobody around me even wants to touch—because “AI” makes people nervous, or because the future makes them tired, or because it’s easier to scroll than to listen.

But I can’t shake the feeling that this work matters—not because it’s trendy, but because it’s a way for ordinary people to participate in shaping what AI becomes.

The Most Important Lesson: AI Needs Direction Like a Stage Needs a Director

Working with these systems has reminded me of something theatre already knows:

A performer without direction will still perform—but it will drift. It will default to habit. It will slide toward cliché. It will smooth out.

AI is the same.

The difference between “cool” and “useful” is structure. The difference between “voice” and “acting” is intent. The difference between output and art is human taste.

That’s what I’ve been building: a way to put taste back into the loop. A pipeline where you can iterate like a director—not once, but hundreds of times—until a character lands, until a narrator breathes correctly, until the pacing feels like someone meant it.

Where This Is Going: Audio First, Then Video

People keep talking about video like it’s the endgame. But the truth is: if you can’t direct voices, you can’t direct scenes.

Audio is the rehearsal space.

If a system can take a script, assign roles, preserve consistent character delivery, and output a publishable episode—then the jump to video is “just another layer.” Hard, yes. But conceptually the same: casting, performance, pacing, staging.

And that’s why I’ve been obsessed with this. Because once this works as a real tool—once it’s a UI that normal people can use—it stops being my private experiment and becomes a creative instrument.

I don’t know if the companies building these models ever see the weird edge-case work happening in spare bedrooms and home studios. They probably don’t. But the truth is: that’s where culture is made.

Not in corporate meetings. Not in policy decks. In the hands of people who refuse to let the future be decided without them.

So yes—I’ve been circling this for a year. And now, for the first time, it feels like the pieces are finally clicking: the voices are better, the gaps are smaller, the workflow is real. I can render long books. I can cast multi-character scenes. I can stitch full episodes with intros and outros and all the boring parts that make it publishable.

It’s not perfect. But it’s real.

And the point isn’t that AI will replace human creativity.

The point is that human creativity needs tools that scale without erasing the human.

That’s what I’m building.

Published by NOMOTO MEDIA

Support independent work

Help fund what comes next.

NOMOTO MEDIA publishes essays, investigations, fiction, audio, and films without a paywall. If the work is valuable to you, help support the next piece.