# Paradee: Kokoro's voice in 8M parameters, and a buzz that lived in the phase

> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/paradee-tiny-tts
> date: 2026-10-08
> tags: tts, speech, audio, distillation, on-device, small-models, quantization

Every film on this site is narrated by Kokoro-82M, voice `af_heart`, on a CPU. So when Hugging Face's apps account posted "Kokoro TTS, but 10× smaller", with "just 9 MB of weights" that "can speak faster than real time on a single CPU thread and runs on your toaster", I had a selfish reason to look. Paradee is the same voice, `af_heart`, distilled into 8.07M parameters.

I expected a story about squeezing weights. Size turned out to be the easy part. Kokoro's text side shrinks by a factor of seven with almost no fuss. What took the work was a faint buzz the small decoder put on every voiced sound, and the fix for it has no weights at all. It is a phase correction applied after synthesis, in one frequency band, chosen because that is exactly where Kokoro stops helping its own vocoder.

The work is by Sahil Mahendrakar, who built it to ship inside his Chrome reading extension. There is a [paper](https://arxiv.org/abs/2610.06817), an [interactive write-up](https://www.sahilmahendrakar.com/writings/paradee), [the code](https://github.com/sahilmahendrakar/paradee) (training included) and [the model files](https://huggingface.co/sahilmahendrakar/Paradee-8M-v1.0), all Apache-2.0. I read all of it and counted the weights in the released checkpoints. I did not run the model on my machine; more on that at the end.

<ModelCard repo="sahilmahendrakar/Paradee-8M-v1.0" />

<RepoCard repo="sahilmahendrakar/paradee" />

## Kokoro is mostly decoder, and the decoder is mostly arithmetic

Kokoro is StyleTTS 2 without the diffusion and the reference encoder. A sentence goes through three stages. The [misaki](https://github.com/hexgrad/misaki) library turns text into phonemes, mostly by dictionary lookup. A **text side** reads those phonemes and decides how to say them: a phoneme-level ALBERT, a prosody predictor that outputs a duration per phoneme plus pitch (F0) and energy contours, and a text encoder that produces a 512-channel feature vector per phoneme. Then a **decoder** turns all of that into a 24 kHz waveform. The author's own diagram is the clearest picture of it I have seen:

<Figure
  src="https://ai.thesatyajit.com/articles/paradee-tiny-tts/fig3.png"
  alt="Block diagram. Text goes to a text-to-phonemes box, then into the text side (28M in Kokoro, 4.2M in Paradee) containing Phoneme BERT, Prosody predictor and Text encoder. Timing, pitch, loudness and sound descriptions flow to the decoder (53M in Kokoro, 3.9M in Paradee), which has decoder blocks and a helper tone fed by pitch, both feeding the waveform generator, which outputs audio. A voice input feeds both halves."
  caption="Kokoro's two halves, with the sizes of each in Kokoro and in Paradee. The 'helper tone' is a fixed sine excitation computed from the predicted pitch, not a learned part. (Sahil Mahendrakar's Paradee write-up, architecture diagram.)"
/>

The decoder has two pieces. A stack of residual 1-D convolutions with adaptive instance normalisation (AdaIN) mixes the phoneme features with pitch, energy and the voice's style vector, one vector per 25 ms frame. Then an [iSTFTNet](https://arxiv.org/abs/2203.02395) generator, a HiFi-GAN that stops short and lets a tiny inverse STFT finish the job, upsamples those frames into audio. Kokoro's `config.json` has the generator's shape: upsampling rates 10 and 6, then an inverse STFT with `n_fft` 20 and hop 5. The generator also gets a "helper tone": sine waves at the predicted pitch and its first eight overtones, computed by formula, so it does not have to invent the periodicity of a voice from nothing.

Before training anything, the paper measures where Kokoro spends its parameters and where it spends its compute. They do not line up at all.

<Figure
  src="https://ai.thesatyajit.com/articles/paradee-tiny-tts/fig2.png"
  alt="Bar chart of three parts of Kokoro. Text side: 35% of parameters, 8% of FLOPs. Decoder AdaIN blocks: 41% of parameters, 3% of FLOPs. Generator: 24% of parameters, 89% of FLOPs."
  caption="Share of Kokoro's parameters and of its compute (FLOPs per second of audio) in each part. The waveform generator is a quarter of the parameters but almost all of the arithmetic. (Paradee paper, Figure 2.)"
/>

Kokoro needs 55 GFLOP per second of audio, and 89% of that happens in the generator, because its last stages run convolutions at 60 times the frame rate. The AdaIN blocks hold 41% of the weights and do 3% of the work. The text side is 35% of the weights and 8% of the work.

That chart decides the whole project. If you want to fit in a small file, you shrink the AdaIN blocks and the text side. If you want to be fast on a CPU, only the generator matters. Paradee has to do both, and the second one is where the quality goes.

## Distilling against the teacher's own homework

The recipe is the part I would steal. StyleTTS 2 is trained with a learned text-to-speech alignment, adversarial losses, a diffusion style sampler and more. None of that is needed to copy a model that already exists, because the teacher will tell you everything it computed.

So the first step is to run Kokoro over 12,000 English sentences from WikiText-103 (23.9 hours of audio) and save every intermediate value, not just the waveform. From `training/scripts/gen_teacher.py`, lines 19 to 30, trimmed:

```python
# training/scripts/gen_teacher.py (paradee, commit 9c8b4d7)
dur = torch.sigmoid(m.predictor.duration_proj(x)).sum(-1)
pred_dur = torch.round(dur).clamp(min=1).long().squeeze(0)
...
en = d.transpose(-1, -2) @ aln; F0, Nn = m.predictor.F0Ntrain(en, s)
t_en = m.text_encoder(ids, L, mask); asr = t_en @ aln
audio = m.decoder(asr, F0, Nn, r[:, :128]).squeeze()
return dict(ids=..., dur=..., pred_dur=..., d=..., t_en=...,
            F0=..., N=..., audio=...)
```

Durations, F0, energy (`N`), the 512-channel phoneme features (`t_en`) and the audio, 6.9 GB in all. With that on disk, the two halves of the student never have to meet during training.

<Figure
  src="https://ai.thesatyajit.com/articles/paradee-tiny-tts/fig1.png"
  alt="Two rows. Top: Kokoro-82M teacher, voice af_heart: Phonemes to Teacher text side (28M params, 8% of FLOPs) to Teacher decoder (53M params, 92% of FLOPs) to Audio, with durations, F0, energy and features between them. Bottom: Paradee student: Phonemes to Student text side (4.2M params) to Student decoder (3.9M params) to Audio. Dashed arrows from each teacher half to its student half: trained to match teacher outputs, and trained to match teacher audio, fed teacher signals."
  caption="Kokoro's two halves (top) and Paradee's (bottom). Each student half learns from the frozen teacher separately (dashed arrows). At inference the two student halves are connected and turn phonemes into audio on their own. (Paradee paper, Figure 1.)"
/>

The student text side reads phonemes and is trained, with the teacher's durations forced in, to predict the teacher's log durations, pitch, energy and phoneme features. Four L1 losses, nothing clever. From `train_text.py`, lines 55 to 58:

```python
l_dur = (F.l1_loss(torch.log(pd.clamp_min(1e-3)), torch.log(dur.clamp_min(1e-3)), reduction="none") * tm).sum() / tm.sum()
l_f0  = (F.l1_loss(pF0, F0, reduction="none") * fm).sum() / fm.sum() / 100
l_n   = (F.l1_loss(pN, N, reduction="none") * fm).sum() / fm.sum()
l_asr = (F.l1_loss(pasr, ten, reduction="none") * tm[:, None]).sum() / tm.sum() / 512 * A.asr_weight
```

The student decoder is fed the teacher's saved features, pitch and energy, and has to produce the teacher's saved audio, on 1.6 s segments. Because both halves train against fixed targets, there is no alignment to learn and no joint training. After training, the student text side is plugged into the student decoder and that is the model.

There is a detail in the corpus script that tells you the author was careful. Kokoro's output differs slightly between CPU and Apple's GPU backend, so `regen_audio.py` re-renders every waveform on the CPU with a fixed seed per sentence. The helper tone has a random initial phase, and if the target audio was made with a different tone than the one the student sees, a waveform-level comparison is meaningless. A comment in `common.py` explains why: one phase feature sits on the plus-or-minus pi branch cut and flips sign between devices.

## Where the parameters went

Paradee is Kokoro's own code at smaller widths. The student modules import `CustomAlbert`, `ProsodyPredictor`, `TextEncoder` and the iSTFTNet `Decoder` straight from the `kokoro` package and pass smaller numbers. The paper puts the text side at 4.23M and the decoder at 3.85M, and I wanted to see that in the files, so I summed the tensor shapes in the two released PyTorch checkpoints. `text_side.pt` holds 4,231,828 parameters and `decoder.pt` holds 3,848,578, so 8,080,406 in total. Of those, 9,050 are weight-norm gains, which disappear when weight norm is folded into the convolutions for inference; fold them and you get 8,071,356, the paper's 8.07M. The vocoder is in there: the decoder's `generator.*` tensors come to 1,185,974, which matches the paper's 1.19M.

Click a stage to see what changed in it:

<ParamMap />

Read top to bottom, the widget shows the pattern. ALBERT drops from 768 hidden units to 256, and from 12 layers to 6, but ALBERT shares one block across all its layers, so the layer cut is cosmetic: the student's whole encoder stack is 691,456 parameters, one block's worth. The prosody predictor goes from 512 to 192 channels. The AdaIN blocks go from 1,024 channels to 256 and come out 12.6 times smaller. The generator goes from 512 initial channels to 128 and comes out 16.6 times smaller, keeping its kernel sizes (3, 7, 11), its upsampling, its excitation module and its iSTFT head.

Two small things fell out of the count. ALBERT's pooler is in the student checkpoint, 65,792 parameters that no code path calls; the ONNX export drops it. And the voice is gone: in place of Kokoro's style vector, which is looked up per voice and per utterance length, each half has one learned constant, 32 numbers for the text side and 16 for the decoder. The student has to absorb the length dependence from the data.

The ONNX file tells the same story with one twist. The fp32 graph carries 9,056,129 numbers, which looks like more than the model. Of those, 1,049,633 are the DFT tables of the phase filter I describe below, and the remaining 8,006,496 are the model's weights after weight norm is folded and the pooler is gone. In the int8 file the export script throws those tables away and rebuilds them at load time from a handful of ops, which is a large part of how the file gets down to 9,037,971 bytes. So the "9 MB" is the int8 ONNX, with weights quantised per output channel in int8 with fp16 scales, everything else in fp16. The paper's 8.45 MB is the same quantisation saved from PyTorch.

Compute fell harder than parameters. The student decoder needs 3.36 GFLOP per second of audio against the teacher decoder's 51.9, and the whole model 3.4 against 55: 15 times less, against 10 times fewer parameters. That gap is the generator.

## The text side was easy, with one surprise

With a 4.2M text side plugged into the frozen teacher decoder, the paper gets UTMOS 4.36 against the teacher's 4.52. (UTMOS is a network trained on human ratings that predicts a mean opinion score from 1 to 5; it is the paper's main metric, and I come back to its limits below.) A 7.1M text side scored 4.35. Capacity was not the limit.

The projection was. The teacher's phoneme features are 512 channels, but about 190 dimensions hold 90% of their variance and 430 hold 99%. A 192-channel student with a linear projection to 512 can express at most 192 of them, and it came within reach of the best any linear map allows. Replacing the projection with a two-layer MLP lifted UTMOS from 4.15 to 4.36. That MLP is 361,472 parameters, about 9% of the text side, and in the paper's account it is the best-spent 9%.

The surprise is the experiment that failed. The author also tried training the text side through the frozen teacher decoder, with a log-mel loss on the audio it produced. That student reached 4.52, the teacher's own score. Then it was connected to Paradee's own decoder and fell to 3.78, while the plainly supervised text side scored 4.39 with the same decoder. Features learned by matching the teacher's features transfer to a new decoder, and features learned by pleasing one particular decoder do not. I think this is the most useful sentence in the paper for anyone distilling a two-stage model: optimise the interface, not the downstream score.

The two halves also did not compound their errors. The student text side through the teacher decoder scores 4.36, the teacher text side through the student decoder 4.37, and the full student 4.39, better than either. The paper's explanation is that the teacher decoder, trained on exact features, amplifies small feature errors, while the student decoder, trained on audio from the start, is more forgiving.

## The decoder that sounded like two voices

Trained on spectral losses alone (a log-mel L1 and a multi-resolution STFT loss at FFT sizes 512, 1024 and 2048), the 3.85M decoder converged nicely on paper, a DTW log-mel distance of 0.50, and scored 2.98 on UTMOS. The author describes it as the `af_heart` voice "with a second, robotic voice speaking at the same time". The spectrograms show why: the student kept crisp harmonic stripes up to 6–8 kHz, where the teacher's harmonics dissolve into breathy noise above 3–4 kHz. The helper tone was leaking straight through.

Adversarial training is the usual fix, and the code has the standard kit: a multi-period discriminator with periods 2, 3, 5, 7 and 11, a multi-resolution spectrogram discriminator, least-squares GAN losses and feature matching. It did nothing at first. With HiFi-GAN's default spectral weight of 45, the spectral term outweighed the adversarial terms about 15 to 1, and 3,000 steps moved UTMOS from 2.98 to 3.02. Lowering the weight to 10 for 5,000 steps took the decoder to 4.29, and then to 3 for 5,000 more took it to 4.37. Making the decoder wider or twice as large did not help. In this regime, the balance of the losses mattered more than the size of the network, and I have not seen that put so bluntly elsewhere.

## The buzz is in the phase, above 2 kHz

After all that, one audible flaw remained: a slight buzz on voiced sounds. This is the section that makes the paper worth reading.

A spectrogram has two halves for every frequency at every moment: a magnitude, how loud that frequency is, and a phase, where its wave is in its cycle. The spectral losses only ever looked at magnitude. To find the buzz, the author swapped halves between student and teacher, with identical excitation. Student magnitude with teacher phase: no audible buzz, UTMOS 4.49. Teacher magnitude with student phase: still buzzing, 4.45. Then the teacher's phase was swapped in for one band at a time, and swapping only 2 to 8 kHz was enough.

The reason for that band is the helper tone. It has nine harmonics, the pitch and eight overtones, and for this voice they end near 1.9 kHz. Below that line the generator is shaping a periodic signal it was handed. Above it, it has to make the harmonics itself, and the small student gets their timing slightly wrong, frame to frame. Incoherent phase across neighbouring frames is the classic "phasiness" of a phase vocoder, and it sounds like a buzz.

<Figure
  src="https://ai.thesatyajit.com/articles/paradee-tiny-tts/fig4.jpg"
  alt="Four spectrograms of the same sentence. Kokoro: harmonic stripes at the bottom fade into detailed texture higher up. First student decoder: similar below 2 kHz but smooth and blurry above it. After adversarial training: less blurry above 2 kHz. Paradee with the filter: closest to Kokoro."
  caption="The same sentence through Kokoro and through three stages of the student decoder. Above the 2 kHz line, where the helper tone stops, the first student is smooth where Kokoro is detailed. (Sahil Mahendrakar's Paradee write-up, spectrogram panel; the four stages are composed side by side here.)"
/>

Retraining did not fix it. Wider critics, waveform and phase losses, distilling the teacher's final iSTFT head and larger decoders all left the buzz audible, and the training script still carries the flags for every one of those attempts (`--wave-weight`, `--phase-weight`, `--head-weight`, `--src-bands`, `--full-phase`). So the author stopped training and corrected the phase after synthesis instead.

The filter uses the one signal in the decoder whose phase is known to be clean: the helper tone itself. In voiced frames and only in the 2–8 kHz band, it measures how far the output's phase has drifted from the tone's, smooths that drift over time, and resynthesises with the smoothed phase and the original magnitude. The whole idea fits in eight lines of `training/scripts/phase_lock.py` (lines 25 to 32):

```python
def lock(s, src, f0, L):
    """s: student waveform, src: its harmonic excitation, f0: F0 curve at 300-sample hops."""
    k = min(len(s), len(src)); s, src = s[:k], src[:k]
    S = torch.stft(s, n, hop, window=W, return_complex=True); E = torch.stft(src, n, hop, window=W, return_complex=True)
    v = torch.tensor([float(f0[min(j * hop // 300, len(f0) - 1)]) > 60 for j in range(S.shape[1])])[None]
    rel = S * E.conj() / E.abs().clamp_min(1e-6)
    ph = torch.where(BAND & v, E.angle() + smooth(rel, L, v).angle(), S.angle())
    return torch.istft(torch.polar(S.abs(), ph), n, hop, window=W, length=k)
```

`rel` is the output's spectrum rotated into the tone's frame of reference: its angle is the drift. `smooth` runs a Hann window of `L` frames over it, and the released export uses `L = 33`, which at a hop of 256 samples is about a third of a second. The new phase is the tone's phase plus the smoothed drift, so the output keeps its own slow phase relationships but loses the frame-to-frame jitter. Frames with F0 below 60 Hz count as unvoiced and pass through untouched, as does everything outside the band. It is close to identity phase locking from the phase-vocoder literature.

It costs about 5% of run time and zero parameters. UTMOS barely noticed: 4.39 to 4.41. The author's ears did, and the paper says plainly that it relied on listening for this artifact because UTMOS did not track it, at one point rating a larger decoder 0.5 lower than one that sounded the same. A slight buzz remains on some stressed syllables, where the student is 3–6 dB quieter than the teacher in that band, and no phase fix can repair a magnitude. You can judge it yourself on the model card's five held-out sentences:

<SamplePairs />

## The claims, checked

Three claims came with the post, and each rests on something different.

"9 MB of weights": the released `paradee_int8.onnx` is 9,037,971 bytes, and the tensors inside it are 7,958,144 int8 weights plus about 0.2 MB of fp16 parameters and scales. Fair.

"Faster than real time on a single CPU thread": yes, comfortably, but read the conditions. The paper's 25.0x is PyTorch on one thread of an Apple M4 Pro, over 8 held-out sentences, about 61 seconds of audio, and it excludes grapheme-to-phoneme conversion. The released ONNX file runs at 17.8x on the same machine in onnxruntime, slower than the PyTorch model because of how the model and filter are packaged. The teacher on the same thread runs at 7.6x, so Paradee is 3.3 times faster, not 15. The paper says why: at this size PyTorch's per-operation overhead on small convolutions and LSTMs dominates, so a 15x cut in FLOPs does not become a 15x cut in time.

I also pressed one example on the Space, which runs on Hugging Face's free `cpu-basic` hardware with the reference code unmodified: 9.8 seconds of audio in 1.11 s, 8.7x real time, and that figure includes misaki's phonemization. One run on a shared VM, possibly a cold one, so treat it as a sanity check rather than a benchmark.

<Figure
  src="https://ai.thesatyajit.com/articles/paradee-tiny-tts/fig5.jpg"
  alt="The Paradee Gradio Space. A waveform player for 9.8 s of speech, the text 'As of August 2015, there were 169 proposed targets for these goals and 304 proposed indicators to show compliance.', speed set to 1, a status line '9.8s of audio in 1.11s on CPU (8.7x real time) · 24 kHz', and the misaki phoneme string for the sentence."
  caption="The Paradee Space after one example: 9.8 s of audio in 1.11 s on the Space's free CPU, phonemization included, and the misaki phonemes the model actually reads. (Screenshot of the Paradee Space on Hugging Face.)"
/>

The paper also lines Paradee up against the other small open models on the same 200 sentences, each in its own voice:

<TtsLadder />

Paradee has the smallest file and the highest UTMOS of the small models. [Kokoro-7M-Distill](https://huggingface.co/oddadmix/Kokoro-7M-Distill), an independent distil of the same teacher released on Hugging Face in September, is about 1.4 times faster and scores 0.23 lower; its own card says that without a WavLM feature-matching term its output "stays subtly robotic", which may well be the same phase problem. Piper's `en_US-lessac-medium` is close on UTMOS and has the worst WER. KittenTTS has the lowest WER. Two caveats from the paper itself: Piper and KittenTTS were timed from text, so their speeds include phonemization and Paradee's do not; and every row is one model in one voice, judged by a predictor.

"Runs on your toaster" is where I would push back. The model is 9 MB. The thing you install is not. `pyproject.toml` pulls onnxruntime, misaki, spaCy, `phonemizer-fork`, `espeakng-loader` and num2words. From PyPI's own listings for Linux x86_64 and Python 3.12, the onnxruntime wheel is 23.6 MB, spaCy 35.5 MB, its `blis` dependency 11.4 MB, numpy 16.7 MB, `espeakng-loader` 10.1 MB (it bundles eSpeak NG) and misaki 3.6 MB, of which 6.1 MB uncompressed is the American English dictionaries. On first run misaki also downloads spaCy's `en_core_web_sm` if it is missing; I did not measure that. The front end outweighs the model by an order of magnitude. It is the same front end Kokoro uses, and Paradee has to use it, because misaki's phoneme spelling is the only spelling the model ever saw.

The browser path makes that point sharper. Browser TTS libraries such as kokoro-js phonemize with eSpeak NG, which spells some sounds differently from misaki. Fed raw eSpeak phonemes, Paradee mumbles: the README puts Whisper's word error rate at about 31%. The repo ships `web/misaki.js`, a 37-line table that rewrites eSpeak's spelling into misaki's (`aɪ` to `I`, `oʊ` to `O`, length marks dropped, flaps to `T`), adapted from misaki's own fallback table, and with it the error rate drops to about 2%. (The file's own comment says 3%; the README says 2%.) If you put any Kokoro-family model in a browser, that table is worth reading.

## The licence chain

Paradee's code and weights are Apache-2.0, like Kokoro's. The training sentences in `training/data/` come from WikiText-103 and are CC BY-SA 3.0, which the README states. The training audio is Kokoro's own output.

Two things sit further up the chain. Kokoro's card says it was trained on "permissive/non-copyrighted audio data", which it describes as public domain audio, Apache- or MIT-licensed audio, and synthetic audio generated by closed TTS models from large providers, plus some CC BY audio it credits by name. Paradee inherits whatever that last category means for you; it is a judgement about Kokoro, not about Paradee, and it is the same judgement this site already made by using Kokoro. And the runtime path loads eSpeak NG, which is GPL-3.0, through `espeakng-loader` and `phonemizer-fork` (also GPL-3.0). `paradee/tts.py` imports misaki's eSpeak fallback unconditionally, so the library is loaded even when every word is in the dictionary. Kokoro's own pipeline does the same, but if you are shipping a closed product, the weights are not where your licence question is.

What Paradee keeps from Kokoro is exactly one thing: the `af_heart` voice, American English, `british=False`. Kokoro v1.0 has 54 voices across 8 languages. The style input is gone from the network, so there is no flag to flip; a second voice means a second distillation.

## Would I swap it into this site's films?

No, and the reason is specific to how the films are made. The narration is voiced offline on a 16-core machine, cached by the text of each line, and painting the frames takes far longer than voicing them, so Kokoro's speed costs this site nothing worth saving. What a switch would cost is quality on exactly the material the films are: two minutes of narration that a viewer chooses to play with sound on. A 0.11 drop in UTMOS sounds small, but the paper itself says UTMOS barely registers the buzz on stressed syllables that the author can still hear, and a narrated film is where a listener would. The pieces that would carry over are encouraging: Paradee reads misaki phonemes and takes a `speed` input, so this site's pronunciation table and its 1.1x reading speed would work unchanged.

Where I would use it is anywhere the voice has to run on the reader's side: reading an article aloud in the browser, an offline device, a laptop without a usable GPU. It was built for exactly that, as the light voice in the author's Chrome extension, and 9 MB with no WebGPU requirement is the right shape for it.

## What would make me trust it more

The quality evidence is an automatic predictor over 200 sentences and the author's own listening, which the paper says itself. UTMOS is also the paper's main instrument for most of its ablations, and the same paper shows UTMOS missing the artifact the paper is about. A small blind listening test between Paradee, the teacher and Kokoro-7M-Distill would settle more than another table. Speed was measured on one Apple machine; a number on a low-end x86 core or a Raspberry Pi would test the toaster claim directly.

None of that changes what I take from it. Distilling a two-stage model half by half against its own saved intermediates is cheap enough to do on a laptop (the author reports about 30 hours of compute, against Kokoro's 1,000 A100 hours), and the census in Figure 2 tells you before you start where your speed has to come from. The phase result is the part I expect to outlive the model: when a small vocoder buzzes, look at the band just above where its excitation ends.

Related on this site: [Pocket TTS with drifting](/articles/pocket-tts-drifting) is the other CPU-first TTS I have taken apart this month, at 100M parameters; [Breeze TTS 2](/articles/breeze-tts-2) is the opposite end of the size range; [the Claude demo videos article](/articles/claude-demo-videos) reads Kokoro's duration predictor, the same predictor Paradee's text side learned to imitate; and [tiny browser models](/articles/tiny-browser-models) is about what it takes to ship a model this small to a page.

## How I checked

I read the paper (arXiv 2610.06817, v1, and the PDF in the repo), the author's write-up, the README and model card, and every file in the GitHub repository at commit `9c8b4d7`, including all the training scripts; code quotes above give file and line. Parameter counts are summed from the tensor shapes in `pytorch/text_side.pt` and `pytorch/decoder.pt` at the model repo's `v1.0` tag, read with a pickle loader that only accepts tensor and dictionary objects, so no code from the checkpoints ran; the `position_ids` buffer is excluded. ONNX counts come from the initializers of `onnx/paradee.onnx` and `onnx/paradee_int8.onnx`, parsed with the `onnx` library, not executed. Kokoro's widths are from its `config.json`, its per-stage parameter counts from the paper's Table 7, and its training-data terms from its model card. Dependency sizes are PyPI's listed wheel sizes. Every quality and speed number is the paper's, except the single Space run, which I triggered by clicking an example on the hosted demo; I did not run Paradee or Kokoro on my own machine. The audio clips on this page are the model card's samples, re-encoded to AAC.
