Cactus Whistle: speech-to-text in 16.9 MB, and a FLEURS bar the Whisper paper doesn't support
mdjsonmcp2026-10-06 · 18 min · speech · asr · on-device · quantization · small-models · benchmarks · explainer
Whistle is Cactus Compute's speech-to-text model. It ships as one file, whistle.cact, and runs on the CPU inside the same C++ engine as their Needle tool-calling models. The announcement on X says it "mostly beats Whisper base with 9x less file size and 6x speed", in seven languages: English, German, French, Spanish, Italian, Dutch and Polish. The headline numbers are 4.31 WER on LibriSpeech test-clean and 10.49 on test-other, against 4.9 and 11.0 for Whisper base at 145.3 MB.
Needle's blocks, engram tables and Cactus Quants are covered in Needle 2, Needle 3, its fine-tune path, the environments and MimiModel. This piece covers what is new for speech, then checks the claims. I read the checkpoint's header and four gate tensors over HTTP range requests, and Whisper base's figures from the Whisper paper's own tables. I did not run the engine, a closed binary, so every latency here is reported.
| Model | Cactus-Compute/whistle · Apache-2.0 · released October 2, 2026 |
| File | whistle.cact, 16,919,407 bytes (measured) · 2-4 bit Cactus Quants, group size 128 |
| Parameters | 55,151,433 (measured, from the checkpoint header) · 19.93M of them in two engram tables |
| Input | 16 kHz mono, up to 30 s per pass · 80 log-mel bins · 375 encoder frames, one per 80 ms |
| Output | text in 8,192 pieces + 7 language tokens, at most 320 · word times · encoder embedding |
| Decoding | 5 beams, length-normalised log prob, keyword bias by Aho-Corasick automaton |
| Engine | prebuilt for 17 targets, macOS to RISC-V to a WASI component (17 platform folders counted on the Hub) |
- license
- Apache-2.0
- branch
- main
- tests
- 19 files
- source
- 451.5 kB
- commit date
- 2026-10-05
by size of tracked source at this commit, file counts in brackets; docs, data and vendored trees excluded
local clone, 2026-10-06 at 7faf26b — branch, commit, commitDate, fileCount, hasTests, languages, license, licenseFile, shallow, testFileCount
shallow clone: counts describe the pinned tree, not the history
Encoder-decoder speech recognition, in one screen
Speech recognition has a shape problem: 16,000 samples a second in, a few tokens a second out, neither length known in advance. An encoder-decoder model listens to the whole clip, then writes.
- Front end. 25 ms windows every 10 ms, each turned into 80 log-mel energies. Thirty seconds is 480,000 samples and 3,000 frames.
- Stem. Strided convolutions halve the frame count three times, to 375 frames, one per 80 ms.
- Encoder. Bidirectional self-attention over all 375 frames at once: a matrix of what was heard where.
- Decoder. An autoregressive language model that, at every layer, queries that matrix through cross-attention.
- Search. Keep the best few partial transcripts (beams) and extend them all.
Whisper has this shape, and so does Whistle. What differs is what fills the boxes.

What is in the 16.9 MB
The size claim holds, which is worth saying because Needle 3's did not: its 20-layer file measured 35.34 MB against a stated 29 MB ceiling. whistle.cact is 16,919,407 bytes. That is 16.92 MB, or 16.14 MiB (measured, from the Hub's file listing).
The checkpoint beside it, checkpoints/whistle.safetensors, is 220,618,620 bytes. Its header lists 116 tensors, all F32, totalling 55,151,433 parameters. 55,151,433 × 4 bytes is the payload to the byte (measured). Here is where those parameters live:
| Part | Parameters | Share |
|---|---|---|
| Convolutional stem | 691,584 | 1.3% |
| Encoder, 8 blocks | 13,665,760 | 24.8% |
| Decoder self-attention, mixers, norms (8 blocks) | 7,223,008 | 13.1% |
| Decoder cross-attention (8 blocks) | 9,446,152 | 17.1% |
| Token embedding, 8,199 × 512 | 4,197,888 | 7.6% |
| Engram tables, 2 sites | 19,927,040 | 36.1% |
| Total | 55,151,433 |
All measured from the header. Three things in that table are not in the blog.
The encoder is Conformer-shaped. Each encoder block holds a pointwise projection from 512 to 1,024 channels, a depthwise convolution with kernel 9, and a projection back to 512. That is the convolution module of a Conformer, and it is 792,064 of each block's 1,658,977 parameters (measured). Each block also has two Monarch Hadamard MLPs, 29,184 parameters together; two half-step MLPs is the Conformer's macaron layout (reasoned from the names). The blog calls these "Simple Attention blocks, shared with Needle" and puts "kernel 9" on the stem. In the checkpoint the stem's convolutions are 3×3 and the kernel-9 convolution sits in every encoder block (measured). A labelling slip, but the encoder is more than Needle's block with the mask removed.
Cross-attention is the decoder's biggest part. Of each decoder block's 2,034,402 parameters, 1,180,257 are cross-attention (measured). The decoder's self-attention uses grouped queries, 8 query heads sharing 2 key-value heads. Cross-attention does not: its keys are 8 heads × 48 and its values 8 heads × 64.
Over a third of the model is lookup tables. The two engram tables are each 4 × 18,432 × 128 (orders 2 and 3, two hashes each). Their 18.87M rows of numbers are read by gather, not matmul. So the parameter count overstates the compute: about 36M parameters ever touch a multiply (reasoned).
The header's __metadata__ records the recipe: SpecAugment (2 frequency masks of width 17, 10 time masks), speed perturbation 0.9 to 1.1, an averaged checkpoint (avg230-250.pkl), then a quantization-aware stage qapt of 25,000 steps with weight_bits embedding=4,stack/mhc=4,encoder/mhc=4,default=2 and 8-bit KV and activations (measured). The training data is not disclosed anywhere I found.
Does that recipe produce 16.9 MB? Cactus Quants cost bits per weight: bits of index plus a 16-bit norm shared by 128 weights (derivation in the Needle 3 piece). Token embedding and mHC weights at 4 bits, every other matrix at 2, small vectors at 16: that predicts 16.29 MB, 0.63 MB under the real file, the rest being headers and tokenizer. The engram tables must be at 2 bits; at 4 the estimate passes 21 MB (both reasoned).
Why a 55M model can sit next to Whisper base
Whisper base is 74M parameters, 6 layers wide 512 (Whisper paper, Table 1), trained on "680,000 hours of multilingual and multitask supervision". So how is a smaller model level with it on LibriSpeech?
The 9x is mostly bits, not architecture. 145.3 MB for a 74M model is about two bytes per parameter: Whisper base as distributed, at 16-bit precision. Whistle has 1.34x fewer parameters. The other 6.4x of the 8.6x size ratio is that Whistle stores about 2 bits per weight. At CQ2's 2.125 bits, Whisper base's 74M parameters would weigh about 19.7 MB (all reasoned). Whisper was never trained to survive 2 bits, which is the point: Whistle's real trick is quantization-aware training to 2 bits, as in Needle, not a radically smaller network.
Whisper base spends its budget elsewhere. Its GPT-2-sized vocabulary, around 50,000 tokens, makes an embedding of around 26M parameters at width 512, a third of the model (reasoned), and it also translates and covers dozens of languages. Whistle's vocabulary is 8,199 rows, 4.2M parameters, for seven languages and one task.
Narrow wins on clean read speech. Whistle leads on LibriSpeech audiobooks and loses on TED talks, meetings and the multilingual averages (below). That is what a smaller model with a narrower target looks like, and also what training data closer to LibriSpeech would look like. The data is undisclosed, so I can't tell which. Cactus states that no test audio is in its training or validation data, checked by audio checksums and speaker IDs (reported).
Gated cross-attention, and eight gates that are wide open
Every decoder layer reads the encoder with
where is the text state, and are the normalised query and keys, the values from the audio, and a learned scalar per layer. The gate idea comes from Flamingo, which bolted cross-attention onto a frozen language model with a tanh gate initialised at zero. At step one the new layers contribute nothing; training opens the gates as cross-attention learns something useful.
The checkpoint has the gate as stack/layers/block/cross_gate, eight floats. I fetched them:
cross_gate logits: 22.806 22.427 22.022 22.011 22.314 22.843 22.760 22.842
sigmoid: 1.0000 to nine decimal places, in all eightσ(22) differs from 1 by about 3 × 10⁻¹⁰ (measured values, reasoned arithmetic). At inference the scalar gate does nothing: Whistle's "gated cross attention" is plain cross-attention, the gate having earned its keep in training and saturated. The decoder's self-attention gates sit at 19.5 to 22.8; the only partly closed gate in the model is the encoder's layer-3 self-attention, at 0.909. Each cross-attention also has a 512 × 512 gate_proj, an elementwise output gate the formula leaves out (measured presence; role reasoned from Needle's self-attention, which has the same tensor).
The systems part is the cross memory. The keys and values come from the audio, which does not change while the decoder writes. So they are projected once per clip and held: 375 frames × (384 + 512) numbers × 8 layers = 2.69M numbers for a 30-second clip (reasoned from the measured shapes). Five beams share it, so beam search costs five short text caches, not five passes over the audio, which helps explain 1,319 tokens/s with 5 beams on an M4 Pro CPU (reported).
The ladder is on the decoder
"Laddered like Needle's" means the same mechanism the Needle 3 piece traced to its training report. Blocks are added in a fixed order: both ends first, then the midpoint of the widest remaining gap. Depth runs the first blocks of that order in their original sequence. Needle's ladder_order() on eight blocks gives 0, 7, 3, 5, 1, 2, 4, 6 (computed from needle/model/architecture.py), so --audio-depth 2 keeps blocks 0 and 7. That Whistle uses this order is reasoned: its blocks "run Needle's code", but its training code is not published.
Two consequences follow from where the ladder sits.
- It doesn't touch time to first token. The encoder always runs all eight blocks over all 375 frames. The first token waits for it whatever the depth, so a shallower decoder buys decode speed and memory, not first-token latency (reasoned).
- Nobody has published what a shallow Whistle can hear. Blog, card and README give word error rates for the 8-layer decoder only. Needle 3's 2-layer slice scored 0.9% on Mobile Actions before fine-tuning; a depth-2 Whistle keeps one engram site and two of eight cross-attention reads. Whether it transcribes anything is unverified.
Keyword biasing: a trie walked beside the beams
A speech decoder is a language model, confident about common words. "Siobhan" sounds like shiv-AWN, and a model that rarely saw the spelling prefers "shivon" to Si-ob-han, whose rare sub-word pieces each cost probability. The audio is fine; the prior is wrong.
Biasing fixes the prior at search time. You pass keywords=["Siobhan", "Krzysztof"]. The engine builds an Aho-Corasick automaton over those phrases, a trie with failure links that tracks all of them at once. Each beam carries a state in it, and when a beam's next token advances that state, its log probability is lifted. That is shallow fusion with a contextual bonus. It needs no retraining, the phrases can change per call, and the cost is one table lookup per beam per token.
The catch is that the bonus has no idea whether the name was said. Push it hard enough and the model writes "Siobhan" over "show on". The widget below has made-up log probabilities: three clips, four finished beams each, scored by length-normalised log prob plus a bonus per keyword token.
- 1.call shivon at noon-0.420
- 2.call she von at noon-0.500
- 3.call Siobhan at noonwhat was said-0.600
- 4.call Joan at noon-0.650
transcript: "call shivon at noon" (wrong)
With those illustrative numbers, "Siobhan" overtakes "shivon" once the bonus passes 0.36 nats per token. "Krzysztof" needs 0.80 to beat "Christoph", a common spelling with a better prior. Past 1.22 the bonus inserts "Siobhan" into a clip that said "put the show on now". The useful setting is a window, and it moves with how rare the name is and what it sounds like.
The Python API takes the phrase list but not the bias strength, so the engine's setting is fixed and unpublished. A reply under the announcement asked how biasing affects false insertions; Cactus answered "pass in your surname and see if Whistle picks it up". That tests recall, not false alarms. No false-insertion rate is published.
One more detail sits in the same search. The detected language is emitted as one of the 7 language tokens in the vocabulary, so detection is a beam decision like any other, and language="de" forces it.
Word timestamps from the decoder's own attention
Cross-attention already says where the decoder looked to write each token: a distribution over 375 frames of 80 ms. Take those maps from heads that behave like alignments, find the monotonic path through them (openai-whisper uses dynamic time warping over chosen "alignment heads" for its word_timestamps), merge sub-word pieces into words, and each word gets a start and an end with no second model. The resolution floor is the 80 ms frame (reasoned from the frame rate). Which heads Whistle uses, and how it smooths them, is inside the closed engine.
Silence returns nothing
Whisper hallucinates on non-speech, writing plausible sentences over silence (the hallucination-projection piece measured how often). Whistle sidesteps it before the model runs: the engine measures the clip's loudness range and, below a threshold, returns an empty transcript and empty language without entering beam search.
A range rather than a level is what makes "steady noise" work. A fan is loud but flat; speech swings. The repository's test feeds one second of zeros and asserts the result is {"text": "", "language": "", "words": [], "ttft_ms": 0.0, "decode_tps": 0.0} (measured, from tests/test_whistle.py). The threshold isn't published, nor how it treats quiet speech in a noisy room, where energy gates usually fail.
The benchmarks, against the Whisper paper

The model card says Whisper's numbers are "as published (Whisper paper Tables 9, 10 and 13)". The chart labels only Whistle's bars, so I read Whisper's from the SVG geometry (6.87 px per WER point) and then from the tables themselves:
| Benchmark | Whistle (reported) | Whisper base, chart bar (measured) | Whisper base, paper (reported) | Lower |
|---|---|---|---|---|
| LibriSpeech test-clean | 4.31 | 4.9 | 4.9 (Table 9) | Whistle |
| LibriSpeech test-other | 10.49 | 11.0 | 11.0 (Table 9) | Whistle |
| TED-LIUM | 7.61 | 5.0 | 5.0 (Table 9) | Whisper |
| AMI | 26.07 | 21.5 | 21.5, AMI-IHM (Table 9) | Whisper |
| MLS, 6 languages | 24.9 | 23.1 | 23.2 mean (Table 10) | Whisper |
| FLEURS, 7 languages | 21.4 | 24.5 | 21.0 mean (Table 13) | Whisper |
The first four rows match to the decimal; Table 9 is Whisper with 5-beam search and temperature fallback, a fair match for Whistle's 5 beams. The MLS bar is within rounding of the six-language mean (reasoned).
FLEURS does not. Whisper's Table 13 gives base 17.9 (German), 8.9 (English), 9.9 (Spanish), 28.5 (French), 17.9 (Italian), 33.0 (Dutch) and 30.8 (Polish). Those seven sum to 146.9, and the mean is 21.0 (reasoned). The chart draws 24.5. The tweet repeats it, "21.4 on the FLEURS average against 24.5". As published, Whisper base is 0.4 points ahead on FLEURS. I can't reconstruct 24.5 from any Whisper base table. The six-language mean without English is 23.0, and the MLS means are 21.6 and 23.2. Whatever produced it, it isn't the table the card cites.
That changes the tally. Against Whisper base, Whistle is ahead on two of the six shared benchmarks: both LibriSpeech splits, by 0.59 and 0.51 points. To be fair to Cactus, their blog lists TED-LIUM, AMI and MLS as Whisper wins. The "mostly beats" is the tweet's, and it rests on the FLEURS bar.
The rows with no Whisper bar:
- SPGISpeech (7.65). Whisper published none. Moonshine tiny v2's bar reads about 7.70 (measured from the SVG), so "ahead" is by 0.05.
- Earnings-22 (19.01). The card says Whisper reports no Earnings-22 figure. Table 16 reports 18.4 for base, but on long-form whole calls, not the segmented clips leaderboards score. Leaving the bar off is defensible; saying no figure exists is not.
Speed is reported, not re-run; the ratios are mine:
- Size 145.3 / 16.9 is 8.6x. First token 73.2 / 11.1 ms is 6.6x, the tweet's "6x speed". Decode 1,319 / 266 tokens/s is 5.0x.
- Whisper pads every input to 30 s, while Whistle's first token tracks the clip: 5.9 ms at 5 s, 11.1 at 10, 36.3 at 30. On a 30-second clip the gap is about 2x, if Whisper's stays near its flat 73.2 (reasoned).
- The published comparison command,
needle whistle compare, calls openai-whisper'stranscribe()with onlylanguageandfp16=False, which decodes greedily. So the speed panel runs Whisper greedy in fp32 PyTorch while the WER panel quotes Whisper with 5 beams. That flatters Whisper's speed, not Whistle's.
Whistle's own word error rates are reported over 86,174 utterances, scored with the Whisper normalizers.
One engine, two models

This is the product argument, and a good one: a device already running Needle gets speech input without a second runtime, quantization format or transcript round-trip through app code. The tool-calling half is Needle 3, whose confidence head is covered in the fine-tune piece and the environments audit. The 17-platform claim checks out: the engine repository on the Hub has 17 platform folders, android-arm64 to wasm-component (counted). For other speech stacks here, see Nemotron 3 Diarization, Qwen-Audio-3.1 and Index-Translate.
What is open and what isn't
- Open: the weights (Apache-2.0,
.cactand fp32 checkpoint), the Python package, and Needle's JAX training code with its ladder. - Closed: the C++ engine, shipped prebuilt; the Linux x86-64 wheel's
libneedle3.so(package 3.2.0) is 1,581,448 bytes (measured). Whistle's model and training code are not in the repository, so the keyword automaton, timestamp alignment, silence threshold and beam scoring can't be checked from source. - Telemetry: Python
transcribe()sends an anonymous usage event per call, without audio or text;NEEDLE_TELEMETRY=0turns it off (measured, fromneedle/_telemetry.py).
The ledger
| Claim | Verdict |
|---|---|
| One 16.9 MB file | Holds. 16,919,407 bytes (measured) |
| 4.31 / 10.49 on LibriSpeech vs 4.9 / 11.0 for Whisper base | Whistle's figures reported; Whisper's match its paper's Table 9 |
| FLEURS 21.4 against 24.5 | Does not hold. Table 13's seven-language mean for base is 21.0 |
| "Mostly beats Whisper base" | Does not hold. Two of six comparable benchmarks |
| "9x less file size" | 8.6x, and mostly 2-bit storage vs 16-bit, not a smaller network (1.34x fewer parameters) |
| "6x speed" | 6.6x to first token on 10 s, 5.0x decode (reported); Whisper ran greedy in the speed test |
| Gated cross-attention at every layer | Present; all eight gates are saturated at σ = 1 (measured) |
| Every decoder depth from 2 up is deployable | Loads by design; no accuracy published below 8 layers |
| Keyword biasing rescues rare names | Mechanism is standard and sound; no false-insertion rate published |
| Silence and steady noise return empty | Silence verified by the repo's test; steady noise and quiet speech untested here |
Whistle is real engineering: a 2-bit quantization-aware encoder-decoder that matches Whisper base on clean English and fits beside a tool-calling model in one runtime. It is not ahead of Whisper base on most benchmarks, and the one bar that makes it look that way doesn't match the paper it cites.