Local Kokoro Narration — Listening Gate 001
Status · SUPPORTED
RESEARCH NOTE
July 27, 2026
Research question
Is a local Kokoro voice promising enough to justify separate long-form narration tests for articles and short stories?
Prior hypothesis
A lightweight local Kokoro model might produce narration natural enough to deserve longer, work-specific evaluation without replacing the existing production renderer.
Method
The same 54-word reflective passage was rendered through the existing Marin control, two Azure voices, and two local Kokoro voices. Wording, durations, hashes, and deliberate paragraph pauses were recorded. The owner then listened to the samples and supplied the missing qualitative judgment.
Observations
- Both Kokoro samples were generated locally with no speech API charge.
- The local model preserved the complete 54-word comparison text.
- The two deliberate thought pauses were present in the local outputs.
- The owner listened and described the Kokoro result as sounding great.
- The owner identified articles and short stories as potentially requiring different narration setups.
Supporting evidence
- A manifest records five directly comparable samples using the same source passage.
- The Kokoro samples retain deterministic source and audio hashes.
- The local female and male samples have non-zero, decodable audio durations.
- The local rendering lane completed without an external speech-provider request.
- The owner’s explicit listening judgment supplies direct human evidence that Kokoro deserves a longer proof.
Counterevidence
- The comparison used only one 54-word reflective passage.
- No complete article or short story has been rendered with Kokoro.
- Long-form listening fatigue has not been evaluated.
- The current test does not establish which Kokoro voice suits which work type.
- Kokoro does not provide the same prompted performance-direction controls as the current OpenAI renderer.
Conclusion
Kokoro passed the initial human listening gate and is strong enough to advance to work-specific long-form auditions. This does not yet establish a production replacement or prove that articles and short stories require different voices or settings.
Confidence
HIGH that Kokoro warrants further testing; LOW that long-form suitability or the final work-type profiles have been established
Limits
- The human judgment covers a short audition rather than a finished work.
- Only one passage style was tested.
- Voice choice, pacing, pauses, and direction still need to be evaluated separately by work type.
- No published narration should be replaced from this result alone.
Production consequence
Keep the current production renderer as the control. Next, create one reflective-article profile and one dramatic-short-story profile using local Kokoro, compare each against the existing production version, and adopt only the settings that win through listening.
Next unresolved question
Do separate Kokoro profiles for a reflective article and a dramatic short story improve naturalness, pacing, and fit over one shared local narration setup without increasing listening fatigue?