Technical Deep-Dive
How Kokoro-82M TTS Runs
on Your iPhone
Updated September 2026 ยท Written by the LocalReader team
Most text-to-speech apps work by sending your text to a remote server, running a large AI model in a data center, and streaming audio back to your phone. LocalReader does something different: it runs a neural text-to-speech model called Kokoro-82M directly on your iPhone's hardware, generating speech locally without ever contacting a server.
This article explains how that works, what Kokoro-82M is, why it can run on a phone, and what the trade-offs are. This is written from our first-hand experience building LocalReader.
What is Kokoro-82M?
Kokoro-82M is an open-weight neural text-to-speech model with 82 million parameters. It was designed to be small enough for edge deployment while producing speech quality that competes with much larger cloud models. The "82M" in the name refers to the parameter count โ for comparison, large cloud TTS models can have billions of parameters.
The key insight behind Kokoro is that you don't need billions of parameters to produce natural-sounding speech for reading tasks. An 82M-parameter model, properly trained and quantized, produces clear, natural voices for reading documents aloud โ which is what LocalReader is built for.
The on-device inference pipeline
When you press play in LocalReader, here's what happens under the hood:
- Text preprocessing โ The raw text (from a PDF, EPUB, or pasted content) goes through a phonemizer that converts written English (or Spanish, French, etc.) into phoneme sequences. This handles abbreviations, numbers, punctuation pauses, and language-specific pronunciation rules.
- Neural inference on Apple Neural Engine โ The phoneme sequence is fed into the Kokoro-82M model, which runs on the Apple Neural Engine (ANE) โ a dedicated hardware accelerator for machine learning built into every iPhone since the A12 chip (iPhone XS, 2018). The ANE is separate from the CPU and GPU, designed specifically for running neural networks efficiently.
- Audio waveform generation โ The model outputs raw audio at 24kHz, sentence by sentence. Playback starts as soon as the first sentence is ready while subsequent sentences render ahead in the background.
- Synchronized highlighting โ The timing data from phoneme alignment is used to drive real-time word highlighting at up to 120Hz on ProMotion displays, keeping visual text in sync with the spoken audio.
Why Apple Neural Engine matters
Running an 82M-parameter model on a phone's CPU would be slow and drain the battery quickly. The Apple Neural Engine changes this equation:
- Dedicated hardware โ The ANE is purpose-built for matrix multiplications and tensor operations that neural networks depend on. It runs these operations 10-100x more efficiently than the CPU.
- Energy efficiency โ The ANE uses a fraction of the power that CPU or GPU inference would consume, so reading a long document doesn't kill your battery.
- Low latency โ LocalReader achieves sub-50ms time-to-first-audio because the ANE processes the first sentence almost instantly. Cloud TTS adds network round-trip time (typically 200-800ms) on top of the processing time.
The A12 Bionic (2018) was the first chip with a practical ANE for this workload, which is why LocalReader requires iPhone XS or later.
Model optimization for mobile
You can't just take a server-side model and drop it onto a phone. Getting Kokoro-82M to run smoothly on iPhone required several optimizations:
- Quantization โ The model weights are quantized from 32-bit floating point to more compact representations, reducing memory footprint without meaningful quality loss. The quantized model fits within the app's memory budget on devices with 3GB+ RAM.
- CoreML conversion โ The original model is converted to Apple's CoreML format, which is the native format the ANE understands. This conversion includes graph optimizations, operator fusion, and memory layout adjustments specific to Apple silicon.
- Streaming synthesis โ Rather than generating audio for an entire chapter at once (which would require too much memory and create a long initial delay), LocalReader synthesizes sentence by sentence. Playback starts immediately while the next sentences are being generated.
- Voice weight bundling โ All 50+ voice configurations are bundled inside the app binary. There's no "download voices" step โ everything is available from the first launch.
The trade-offs of on-device TTS
On-device TTS isn't strictly "better" than cloud TTS โ it's a different set of trade-offs. We chose on-device for LocalReader because the benefits aligned with what a reading app needs:
| Dimension | Cloud TTS | On-Device (LocalReader) |
|---|---|---|
| Voice quality ceiling | Higher (larger models) | Very good (Kokoro-82M) |
| Latency | 200-800ms + network | Sub-50ms |
| Offline support | No | Full |
| Privacy | Text sent to servers | Nothing leaves device |
| Language coverage | 50+ languages typical | 8 languages |
| Device requirements | Any device with internet | iPhone XS+ (A12 chip) |
| Cost model | Per-character or subscription | Flat subscription, no metering |
For a reading app โ where you're processing long documents privately, often without internet โ the on-device trade-offs are the right ones. You get privacy, offline access, and instant playback in exchange for fewer languages and a slightly lower ceiling on voice expressiveness.
50+ voices from one model
Kokoro-82M is a multi-speaker model โ it can produce different voices from the same model weights by conditioning on a speaker embedding. Each of LocalReader's 50+ voices is a different speaker configuration, not a separate model. This means:
- Switching voices is instant โ no model reload.
- All voices share the same quality baseline.
- Adding new voices doesn't increase the app's binary size significantly.
The voices span American English, British English, Spanish, French, Japanese, Chinese (Mandarin), Hindi, Italian, and Portuguese โ with natural gender and accent variation.
What this means for you
You don't need to understand any of this to use LocalReader. The point is: when you press play, real AI is running on your phone โ the same class of neural network technology that powers cloud TTS services โ but it's doing it locally, privately, and instantly. No server in between.
If you're curious, try opening a long PDF in Airplane Mode. It reads the entire thing, every page, with natural voices and word highlighting. That's Kokoro-82M on your Apple Neural Engine, doing its job.
FAQ
What is Kokoro-82M TTS?
Kokoro-82M is an open-weight neural text-to-speech model with 82 million parameters. It generates natural human-like speech from text input and is small enough to run on mobile devices without requiring cloud servers.
Can Kokoro TTS run on iPhone without internet?
Yes. LocalReader bundles the Kokoro-82M model weights inside the app and runs inference on the Apple Neural Engine. No internet connection is needed โ speech is generated entirely on your iPhone in real time.
How does on-device TTS compare to cloud TTS quality?
Kokoro-82M produces natural, human-like speech comparable to many cloud TTS services. The trade-off is that cloud services can run larger models (billions of parameters) with potentially more expressiveness, while on-device gives you privacy, offline access, and sub-50ms latency.
Kokoro-82M on-device TTS ยท No account ยท No internet required