# Cactus Whistle: speech-to-text in 16.9 MB, and a FLEURS bar the Whisper paper doesn't support

> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/whistle-stt
> date: 2026-10-06
> tags: speech, asr, on-device, quantization, small-models, benchmarks, explainer

[Whistle](https://cactuscompute.com/blog/whistle) is Cactus Compute's speech-to-text model. It ships as one file, `whistle.cact`, and runs on the CPU inside the same C++ engine as their [Needle](/articles/needle-3) tool-calling models. The announcement on X says it "mostly beats Whisper base with 9x less file size and 6x speed", in seven languages: English, German, French, Spanish, Italian, Dutch and Polish. The headline numbers are 4.31 WER on LibriSpeech test-clean and 10.49 on test-other, against 4.9 and 11.0 for Whisper base at 145.3 MB.

Needle's blocks, engram tables and Cactus Quants are covered in [Needle 2](/articles/needle-finetune), [Needle 3](/articles/needle-3), [its fine-tune path](/articles/needle-3-finetune), [the environments](/articles/needle-environments) and [MimiModel](/articles/mimimodel). This piece covers what is new for speech, then checks the claims. I read the checkpoint's header and four gate tensors over HTTP range requests, and Whisper base's figures from the Whisper paper's own tables. I did not run the engine, a closed binary, so every latency here is reported.

| | |
|---|---|
| Model | [Cactus-Compute/whistle](https://huggingface.co/Cactus-Compute/whistle) · Apache-2.0 · released October 2, 2026 |
| File | `whistle.cact`, **16,919,407 bytes** (measured) · 2-4 bit Cactus Quants, group size 128 |
| Parameters | **55,151,433** (measured, from the checkpoint header) · 19.93M of them in two engram tables |
| Input | 16 kHz mono, up to 30 s per pass · 80 log-mel bins · 375 encoder frames, one per 80 ms |
| Output | text in 8,192 pieces + 7 language tokens, at most 320 · word times · encoder embedding |
| Decoding | 5 beams, length-normalised log prob, keyword bias by Aho-Corasick automaton |
| Engine | prebuilt for 17 targets, macOS to RISC-V to a WASI component (17 platform folders counted on the Hub) |

<RepoCard repo="cactus-compute/needle" />

## Encoder-decoder speech recognition, in one screen

Speech recognition has a shape problem: 16,000 samples a second in, a few tokens a second out, neither length known in advance. An encoder-decoder model listens to the whole clip, then writes.

- **Front end.** 25 ms windows every 10 ms, each turned into 80 log-mel energies. Thirty seconds is 480,000 samples and 3,000 frames.
- **Stem.** Strided convolutions halve the frame count three times, to 375 frames, one per 80 ms.
- **Encoder.** Bidirectional self-attention over all 375 frames at once: a matrix of what was heard where.
- **Decoder.** An autoregressive language model that, at every layer, queries that matrix through cross-attention.
- **Search.** Keep the best few partial transcripts (beams) and extend them all.

Whisper has this shape, and so does Whistle. What differs is what fills the boxes.

<Figure
  src="https://ai.thesatyajit.com/articles/whistle-stt/fig1.png"
  alt="Whistle's architecture as drawn on Cactus's blog, two stacked columns. Encoder: waveform (16 kHz mono, up to 30 s, 480,000 samples), log-mel (80 bins, 25 ms window, 10 ms hop, 250-3500 Hz, 3,000 frames), convolutional stem (128 channels, three halvings to 375 frames, one per 80 ms), Simple Attention blocks times 8 shared with Needle. Decoder: cross memory (K and V projected once per clip, 375 frames times 8 layers, shared by every beam), Laddered Simple Attention blocks times 8 (GQA 8 query to 2 KV heads, 48 qk and 64 v, 3-tap causal conv, engram at layers 3 and 7, width 512), gated cross attention (x gets x plus sigma of g times softmax of q-hat K-hat transpose over root d, times V), beam search times 5 with keyword bias by Aho-Corasick automaton, and a transcript of 8,192 text pieces plus 7 language tokens, up to 320, with word times from the decoder's attention."
  caption="Whistle end to end. The two highlighted rows are the speech-specific parts: a cross memory computed once per clip, and a gated cross-attention in every decoder layer. Everything marked 'shared with Needle' runs Needle's engine code (Cactus blog, 'The model')."
/>

## What is in the 16.9 MB

The size claim holds, which is worth saying because [Needle 3's did not](/articles/needle-3): its 20-layer file measured 35.34 MB against a stated 29 MB ceiling. `whistle.cact` is 16,919,407 bytes. That is 16.92 MB, or 16.14 MiB (measured, from the Hub's file listing).

The checkpoint beside it, `checkpoints/whistle.safetensors`, is 220,618,620 bytes. Its header lists 116 tensors, all F32, totalling 55,151,433 parameters. 55,151,433 × 4 bytes is the payload to the byte (measured). Here is where those parameters live:

| Part | Parameters | Share |
|---|---|---|
| Convolutional stem | 691,584 | 1.3% |
| Encoder, 8 blocks | 13,665,760 | 24.8% |
| Decoder self-attention, mixers, norms (8 blocks) | 7,223,008 | 13.1% |
| Decoder cross-attention (8 blocks) | 9,446,152 | 17.1% |
| Token embedding, 8,199 × 512 | 4,197,888 | 7.6% |
| Engram tables, 2 sites | 19,927,040 | 36.1% |
| **Total** | **55,151,433** | |

All measured from the header. Three things in that table are not in the blog.

**The encoder is Conformer-shaped.** Each encoder block holds a pointwise projection from 512 to 1,024 channels, a depthwise convolution with kernel 9, and a projection back to 512. That is the convolution module of a [Conformer](https://arxiv.org/abs/2005.08100), and it is 792,064 of each block's 1,658,977 parameters (measured). Each block also has *two* Monarch Hadamard MLPs, 29,184 parameters together; two half-step MLPs is the Conformer's macaron layout (reasoned from the names). The blog calls these "Simple Attention blocks, shared with Needle" and puts "kernel 9" on the stem. In the checkpoint the stem's convolutions are 3×3 and the kernel-9 convolution sits in every encoder block (measured). A labelling slip, but the encoder is more than Needle's block with the mask removed.

**Cross-attention is the decoder's biggest part.** Of each decoder block's 2,034,402 parameters, 1,180,257 are cross-attention (measured). The decoder's self-attention uses grouped queries, 8 query heads sharing 2 key-value heads. Cross-attention does not: its keys are 8 heads × 48 and its values 8 heads × 64.

**Over a third of the model is lookup tables.** The two engram tables are each 4 × 18,432 × 128 (orders 2 and 3, two hashes each). Their 18.87M rows of numbers are read by gather, not matmul. So the parameter count overstates the compute: about 36M parameters ever touch a multiply (reasoned).

The header's `__metadata__` records the recipe: SpecAugment (2 frequency masks of width 17, 10 time masks), speed perturbation 0.9 to 1.1, an averaged checkpoint (`avg230-250.pkl`), then a quantization-aware stage `qapt` of 25,000 steps with `weight_bits` `embedding=4,stack/mhc=4,encoder/mhc=4,default=2` and 8-bit KV and activations (measured). The training *data* is not disclosed anywhere I found.

Does that recipe produce 16.9 MB? Cactus Quants cost $b + \tfrac{1}{8}$ bits per weight: $b$ bits of index plus a 16-bit norm shared by 128 weights ([derivation in the Needle 3 piece](/articles/needle-3)). Token embedding and mHC weights at 4 bits, every other matrix at 2, small vectors at 16: that predicts 16.29 MB, 0.63 MB under the real file, the rest being headers and tokenizer. The engram tables must be at 2 bits; at 4 the estimate passes 21 MB (both reasoned).

## Why a 55M model can sit next to Whisper base

Whisper base is 74M parameters, 6 layers wide 512 (Whisper paper, Table 1), trained on "680,000 hours of multilingual and multitask supervision". So how is a smaller model level with it on LibriSpeech?

**The 9x is mostly bits, not architecture.** 145.3 MB for a 74M model is about two bytes per parameter: Whisper base as distributed, at 16-bit precision. Whistle has 1.34x fewer parameters. The other 6.4x of the 8.6x size ratio is that Whistle stores about 2 bits per weight. At CQ2's 2.125 bits, Whisper base's 74M parameters would weigh about 19.7 MB (all reasoned). Whisper was never trained to survive 2 bits, which is the point: Whistle's real trick is quantization-aware training to 2 bits, as in Needle, not a radically smaller network.

**Whisper base spends its budget elsewhere.** Its GPT-2-sized vocabulary, around 50,000 tokens, makes an embedding of around 26M parameters at width 512, a third of the model (reasoned), and it also translates and covers dozens of languages. Whistle's vocabulary is 8,199 rows, 4.2M parameters, for seven languages and one task.

**Narrow wins on clean read speech.** Whistle leads on LibriSpeech audiobooks and loses on TED talks, meetings and the multilingual averages (below). That is what a smaller model with a narrower target looks like, and also what training data closer to LibriSpeech would look like. The data is undisclosed, so I can't tell which. Cactus states that no test audio is in its training or validation data, checked by audio checksums and speaker IDs (reported).

## Gated cross-attention, and eight gates that are wide open

Every decoder layer reads the encoder with

$$
x \leftarrow x + \sigma(g)\cdot \mathrm{softmax}\!\left(\frac{\hat q \hat K^\top}{\sqrt d}\right) V
$$

where $x$ is the text state, $\hat q$ and $\hat K$ are the normalised query and keys, $V$ the values from the audio, and $g$ a learned scalar per layer. The gate idea comes from [Flamingo](https://arxiv.org/abs/2204.14198), which bolted cross-attention onto a frozen language model with a tanh gate initialised at zero. At step one the new layers contribute nothing; training opens the gates as cross-attention learns something useful.

The checkpoint has the gate as `stack/layers/block/cross_gate`, eight floats. I fetched them:

```text
cross_gate logits:  22.806  22.427  22.022  22.011  22.314  22.843  22.760  22.842
sigmoid:             1.0000 to nine decimal places, in all eight
```

σ(22) differs from 1 by about 3 × 10⁻¹⁰ (measured values, reasoned arithmetic). At inference the scalar gate does nothing: Whistle's "gated cross attention" is plain cross-attention, the gate having earned its keep in training and saturated. The decoder's self-attention gates sit at 19.5 to 22.8; the only partly closed gate in the model is the encoder's layer-3 self-attention, at 0.909. Each cross-attention also has a 512 × 512 `gate_proj`, an elementwise output gate the formula leaves out (measured presence; role reasoned from Needle's self-attention, which has the same tensor).

The systems part is the **cross memory**. The keys and values come from the audio, which does not change while the decoder writes. So they are projected once per clip and held: 375 frames × (384 + 512) numbers × 8 layers = 2.69M numbers for a 30-second clip (reasoned from the measured shapes). Five beams share it, so beam search costs five short text caches, not five passes over the audio, which helps explain 1,319 tokens/s with 5 beams on an M4 Pro CPU (reported).

## The ladder is on the decoder

"Laddered like Needle's" means the same mechanism [the Needle 3 piece](/articles/needle-3) traced to its training report. Blocks are added in a fixed order: both ends first, then the midpoint of the widest remaining gap. Depth $d$ runs the first $d$ blocks of that order in their original sequence. Needle's `ladder_order()` on eight blocks gives `0, 7, 3, 5, 1, 2, 4, 6` (computed from `needle/model/architecture.py`), so `--audio-depth 2` keeps blocks 0 and 7. That Whistle uses this order is reasoned: its blocks "run Needle's code", but its training code is not published.

<DecoderLadder />

Two consequences follow from where the ladder sits.

1. **It doesn't touch time to first token.** The encoder always runs all eight blocks over all 375 frames. The first token waits for it whatever the depth, so a shallower decoder buys decode speed and memory, not first-token latency (reasoned).
2. **Nobody has published what a shallow Whistle can hear.** Blog, card and README give word error rates for the 8-layer decoder only. Needle 3's 2-layer slice scored 0.9% on Mobile Actions before fine-tuning; a depth-2 Whistle keeps one engram site and two of eight cross-attention reads. Whether it transcribes anything is unverified.

## Keyword biasing: a trie walked beside the beams

A speech decoder is a language model, confident about common words. "Siobhan" sounds like *shiv-AWN*, and a model that rarely saw the spelling prefers "shivon" to Si-ob-han, whose rare sub-word pieces each cost probability. The audio is fine; the prior is wrong.

Biasing fixes the prior at search time. You pass `keywords=["Siobhan", "Krzysztof"]`. The engine builds an [Aho-Corasick automaton](https://en.wikipedia.org/wiki/Aho%E2%80%93Corasick_algorithm) over those phrases, a trie with failure links that tracks all of them at once. Each beam carries a state in it, and when a beam's next token advances that state, its log probability is lifted. That is shallow fusion with a contextual bonus. It needs no retraining, the phrases can change per call, and the cost is one table lookup per beam per token.

The catch is that the bonus has no idea whether the name was said. Push it hard enough and the model writes "Siobhan" over "show on". The widget below has made-up log probabilities: three clips, four finished beams each, scored by length-normalised log prob plus a bonus per keyword token.

<KeywordBiasBeam />

With those illustrative numbers, "Siobhan" overtakes "shivon" once the bonus passes 0.36 nats per token. "Krzysztof" needs 0.80 to beat "Christoph", a common spelling with a better prior. Past 1.22 the bonus inserts "Siobhan" into a clip that said "put the show on now". The useful setting is a window, and it moves with how rare the name is and what it sounds like.

The Python API takes the phrase list but not the bias strength, so the engine's setting is fixed and unpublished. A reply under the announcement asked how biasing affects false insertions; Cactus answered "pass in your surname and see if Whistle picks it up". That tests recall, not false alarms. No false-insertion rate is published.

One more detail sits in the same search. The detected language is emitted as one of the 7 language tokens in the vocabulary, so detection is a beam decision like any other, and `language="de"` forces it.

## Word timestamps from the decoder's own attention

Cross-attention already says where the decoder looked to write each token: a distribution over 375 frames of 80 ms. Take those maps from heads that behave like alignments, find the monotonic path through them (openai-whisper uses dynamic time warping over chosen "alignment heads" for its `word_timestamps`), merge sub-word pieces into words, and each word gets a start and an end with no second model. The resolution floor is the 80 ms frame (reasoned from the frame rate). Which heads Whistle uses, and how it smooths them, is inside the closed engine.

## Silence returns nothing

Whisper hallucinates on non-speech, writing plausible sentences over silence ([the hallucination-projection piece](/articles/whisper-hallucination-projection) measured how often). Whistle sidesteps it before the model runs: the engine measures the clip's loudness range and, below a threshold, returns an empty transcript and empty language without entering beam search.

A *range* rather than a level is what makes "steady noise" work. A fan is loud but flat; speech swings. The repository's test feeds one second of zeros and asserts the result is `{"text": "", "language": "", "words": [], "ttft_ms": 0.0, "decode_tps": 0.0}` (measured, from `tests/test_whistle.py`). The threshold isn't published, nor how it treats quiet speech in a noisy room, where energy gates usually fail.

## The benchmarks, against the Whisper paper

<Figure
  src="https://ai.thesatyajit.com/articles/whistle-stt/fig2.png"
  alt="Grouped bar chart, word error rate percent, lower is better, for Whistle (orange), Whisper base and Moonshine tiny v2 (greys). Whistle's labelled values: LibriSpeech test-clean 4.31, test-other 10.49, SPGISpeech 7.65, Earnings-22 19.01, AMI 26.07, AMI cleaned 22.87, TED-LIUM 7.61, FLEURS 21.4, MLS 24.9. Whisper base has no bar for SPGISpeech, Earnings-22 or AMI cleaned. Below, three horizontal panels: size in MB (16.9, 145.3, 41.9), time to first token in ms on a 10 s clip (11.1, 73.2, 22.8), and decode tokens per second (1,319, 266, 262)."
  caption="Cactus's chart. Only Whistle's bars are labelled; the Whisper base bars below are read from the SVG's bar heights. Speed is 10 s of audio on an Apple M4 Pro, each model at its runtime's defaults (Cactus Whistle model card, benchmark figure)."
/>

The model card says Whisper's numbers are "as published (Whisper paper Tables 9, 10 and 13)". The chart labels only Whistle's bars, so I read Whisper's from the SVG geometry (6.87 px per WER point) and then from the tables themselves:

| Benchmark | Whistle (reported) | Whisper base, chart bar (measured) | Whisper base, paper (reported) | Lower |
|---|---|---|---|---|
| LibriSpeech test-clean | 4.31 | 4.9 | 4.9 (Table 9) | Whistle |
| LibriSpeech test-other | 10.49 | 11.0 | 11.0 (Table 9) | Whistle |
| TED-LIUM | 7.61 | 5.0 | 5.0 (Table 9) | Whisper |
| AMI | 26.07 | 21.5 | 21.5, AMI-IHM (Table 9) | Whisper |
| MLS, 6 languages | 24.9 | 23.1 | 23.2 mean (Table 10) | Whisper |
| FLEURS, 7 languages | 21.4 | **24.5** | **21.0 mean (Table 13)** | **Whisper** |

The first four rows match to the decimal; Table 9 is Whisper with 5-beam search and temperature fallback, a fair match for Whistle's 5 beams. The MLS bar is within rounding of the six-language mean (reasoned).

**FLEURS does not.** Whisper's Table 13 gives base 17.9 (German), 8.9 (English), 9.9 (Spanish), 28.5 (French), 17.9 (Italian), 33.0 (Dutch) and 30.8 (Polish). Those seven sum to 146.9, and the mean is **21.0** (reasoned). The chart draws 24.5. The tweet repeats it, "21.4 on the FLEURS average against 24.5". As published, Whisper base is 0.4 points *ahead* on FLEURS. I can't reconstruct 24.5 from any Whisper base table. The six-language mean without English is 23.0, and the MLS means are 21.6 and 23.2. Whatever produced it, it isn't the table the card cites.

That changes the tally. Against Whisper base, Whistle is ahead on two of the six shared benchmarks: both LibriSpeech splits, by 0.59 and 0.51 points. To be fair to Cactus, their blog lists TED-LIUM, AMI and MLS as Whisper wins. The "mostly beats" is the tweet's, and it rests on the FLEURS bar.

The rows with no Whisper bar:

- **SPGISpeech (7.65).** Whisper published none. Moonshine tiny v2's bar reads about 7.70 (measured from the SVG), so "ahead" is by 0.05.
- **Earnings-22 (19.01).** The card says Whisper reports no Earnings-22 figure. Table 16 reports 18.4 for base, but on *long-form* whole calls, not the segmented clips leaderboards score. Leaving the bar off is defensible; saying no figure exists is not.

**Speed** is reported, not re-run; the ratios are mine:

- **Size** 145.3 / 16.9 is 8.6x. **First token** 73.2 / 11.1 ms is 6.6x, the tweet's "6x speed". **Decode** 1,319 / 266 tokens/s is 5.0x.
- Whisper pads every input to 30 s, while Whistle's first token tracks the clip: 5.9 ms at 5 s, 11.1 at 10, 36.3 at 30. On a 30-second clip the gap is about 2x, if Whisper's stays near its flat 73.2 (reasoned).
- The published comparison command, `needle whistle compare`, calls openai-whisper's `transcribe()` with only `language` and `fp16=False`, which decodes greedily. So the speed panel runs Whisper greedy in fp32 PyTorch while the WER panel quotes Whisper with 5 beams. That flatters Whisper's speed, not Whistle's.

Whistle's own word error rates are reported over 86,174 utterances, scored with the Whisper normalizers.

## One engine, two models

<Figure
  src="https://ai.thesatyajit.com/articles/whistle-stt/fig3.png"
  alt="Diagram titled 'One engine, three ways to load it'. LOAD row: needle --model needle3.cact --model whistle.cact --tools tools.json. IN row: a waveform labelled clip.wav. OUT row: JSON with function_calls set_lights room kitchen on false, confidence 0.94, audio_text 'turn off the kitchen lights', audio_language en."
  caption="The middle of three loadouts: both models in one process. The clip goes in, the transcript stays inside the engine, and one JSON object comes out with Needle's tool call and Whistle's fields prefixed audio_ (Cactus Whistle model card, 'With Needle'; one frame of an animated figure)."
/>

This is the product argument, and a good one: a device already running Needle gets speech input without a second runtime, quantization format or transcript round-trip through app code. The tool-calling half is Needle 3, whose confidence head is covered in [the fine-tune piece](/articles/needle-3-finetune) and [the environments audit](/articles/needle-environments). The 17-platform claim checks out: the engine repository on the Hub has 17 platform folders, `android-arm64` to `wasm-component` (counted). For other speech stacks here, see [Nemotron 3 Diarization](/articles/nemotron-3-diarization), [Qwen-Audio-3.1](/articles/qwen-audio-3-1) and [Index-Translate](/articles/index-translate).

## What is open and what isn't

- **Open:** the weights (Apache-2.0, `.cact` and fp32 checkpoint), the Python package, and Needle's JAX training code with its ladder.
- **Closed:** the C++ engine, shipped prebuilt; the Linux x86-64 wheel's `libneedle3.so` (package 3.2.0) is 1,581,448 bytes (measured). Whistle's model and training code are not in the repository, so the keyword automaton, timestamp alignment, silence threshold and beam scoring can't be checked from source.
- **Telemetry:** Python `transcribe()` sends an anonymous usage event per call, without audio or text; `NEEDLE_TELEMETRY=0` turns it off (measured, from `needle/_telemetry.py`).

## The ledger

| Claim | Verdict |
|---|---|
| One 16.9 MB file | **Holds.** 16,919,407 bytes (measured) |
| 4.31 / 10.49 on LibriSpeech vs 4.9 / 11.0 for Whisper base | Whistle's figures reported; Whisper's match its paper's Table 9 |
| FLEURS 21.4 against 24.5 | **Does not hold.** Table 13's seven-language mean for base is 21.0 |
| "Mostly beats Whisper base" | **Does not hold.** Two of six comparable benchmarks |
| "9x less file size" | 8.6x, and mostly 2-bit storage vs 16-bit, not a smaller network (1.34x fewer parameters) |
| "6x speed" | 6.6x to first token on 10 s, 5.0x decode (reported); Whisper ran greedy in the speed test |
| Gated cross-attention at every layer | Present; all eight gates are saturated at σ = 1 (measured) |
| Every decoder depth from 2 up is deployable | Loads by design; no accuracy published below 8 layers |
| Keyword biasing rescues rare names | Mechanism is standard and sound; no false-insertion rate published |
| Silence and steady noise return empty | Silence verified by the repo's test; steady noise and quiet speech untested here |

Whistle is real engineering: a 2-bit quantization-aware encoder-decoder that matches Whisper base on clean English and fits beside a tool-calling model in one runtime. It is not ahead of Whisper base on most benchmarks, and the one bar that makes it look that way doesn't match the paper it cites.
