Local Kokoro Long-Form Narration — Run 001
Status · SUPPORTED
RESEARCH REPORT
July 27, 2026
Research question
Can the local Kokoro production lane render and deliver one complete long-form N’OMOTO article without a speech API request?
Prior hypothesis
A lightweight local Kokoro model could render a complete single-narrator article with resumable paragraph units, exact source lineage, deliberate paragraph pauses, and standard web playback.
Method
One published 2,872-word Maimonides article with no existing audio was divided into 24 authoritative paragraph units after excluding one obsolete promotional opening. Each unit was rendered locally with the Kokoro af_heart voice at a restrained rate, preserved with text and audio hashes, assembled into one MP3, fully decoded, uploaded, range-tested, attached to the canonical article, and played through the public browser interface before and after refresh.
Observations
- All 24 paragraph units rendered locally and remained resumable.
- The complete narration contains 2,872 source words and runs for approximately 18 minutes 36 seconds.
- The assembled MP3 is non-empty, decodable, and served with HTTP partial-content support.
- The public article exposes Read and Listen without a YouTube action for that work.
- Browser playback started, paused, and started again after refresh.
- The obsolete Selenius Media promotional opening was excluded while the substantive article text remained the narration source.
Supporting evidence
- The source and narration hashes are recorded and the 24 segment indexes are sequential.
- The final MP3 hash matches the uploaded WordPress media file.
- Full-file decoding completed without an audio error.
- The public media route returned HTTP 206 for a byte-range request.
- The browser reported a non-zero 1,116-second duration and ready playback state.
- The canonical public article returns HTTP 200 and links the verified audio asset.
Counterevidence
- The owner has not yet listened to the entire 18-minute result.
- Only one reflective educational work has been rendered long-form.
- No dramatic short story has completed the work-specific comparison.
- No multi-speaker, cloning, speech-to-speech, dubbing, or translation capability was tested.
- Kokoro provides less direct expressive prompting than Eleven v3 or the current OpenAI renderer.
- This run establishes technical production and delivery, not equality with every commercial voice system.
Conclusion
Local Kokoro can complete the technical single-narrator Article → Listen workflow for one 2,872-word work without a speech API request. Long-form artistic adoption remains conditional on full human listening, and this result does not establish ElevenLabs-equivalent expressive, cloning, dialogue, or dubbing capabilities.
Confidence
HIGH for complete local rendering, delivery, and browser playback; LOW for long-form artistic preference and untested commercial-provider capabilities
Limits
- One educational article cannot establish a general production default.
- Complete-owner listening remains pending.
- The work used one voice and one narration profile.
- The proof did not test dramatic acting, dialogue, cloning, dubbing, or multilingual production.
Production consequence
The local Kokoro lane is now eligible for bounded single-narrator article tests. Azure remains a parallel credit-backed production lane. Batch replacement and the short-story profile remain on hold until full listening and a dramatic-work proof pass.
Next unresolved question
Does a complete dramatic short story require a different Kokoro voice or pacing profile, and does the owner prefer both work-specific profiles after full-length listening?