# Spark-X2.5-4B: fifteen minutes with the files behind a viral model

> Satyajit Ghana — Head of Engineering @ Inkers Technology
> canonical: https://ai.thesatyajit.com/articles/spark-x2-5-4b
> date: 2026-10-06
> tags: llm, open-weights, small-models, on-device, long-context, hybrid-attention, evaluation, benchmarks, security

On 2026-10-05 a post on X with about 12,300 views announced "a new open-source model that just hit
#1 on Hugging Face and it runs fully local on a 16GB laptop". The bullets: 4 billion parameters, a
"native 1 million token context window", "strong coding + agent capabilities", and support for
Ollama, LM Studio and llama.cpp. The link went to a GitHub repository. The video under the post
scrolls that repository's README and nothing else.

The model is [Spark-X2.5-4B](https://huggingface.co/XHToken/Spark-X2.5-4B). Its Hugging Face repo
carries the `custom_code` tag: `transformers` loads it with `trust_remote_code=True`, which runs
Python that the publisher wrote. That is exactly the situation where
a viral post and a checkpoint deserve fifteen minutes of reading before anyone runs anything.

This is that reading, done end to end, followed by the checklist so you can repeat it on the next one.
Short version: **the model is real, original and well built**. Most of the post's claims are true in
a narrower sense than they sound, and one of them I could not verify at all.

<Callout type="note">
**What this is and is not.** I did not run the model or its code, and I did not reproduce any
benchmark. Everything below comes from public metadata, files under 25 MB, and HTTP range requests
into the shards. I label each number: **measured** (I computed it from a file), **reported** (the
publisher's or press figure, not re-run) or **reasoned** (my arithmetic on the other two).
</Callout>

<ModelCard repo="XHToken/Spark-X2.5-4B" />

<Figure
  src="https://ai.thesatyajit.com/articles/spark-x2-5-4b/fig4.jpg"
  alt="A frame from the video attached to the X post: the GitHub page for XHToken/Spark-X2.5 in a dark theme, listing an assets folder, .gitignore, CODE_OF_CONDUCT.md, LICENSE, README.md and SECURITY.md."
  caption="The post's video opens on the GitHub repository it links. The file list is the whole repo: a README, a licence and policy files, no code and no weights. The weights live on Hugging Face (frame from the video attached to the X post by @Forhanvv, cropped to the browser window)."
/>

## Who published it

Start with the name, because "XHToken" reads like a crypto project and that matters for trust.
It is not one. Chinese launch coverage on 2026-09-01 ([Beijing News](https://m.bjnews.com.cn/detail/1788270778129337.html),
[10jqka](https://stock.10jqka.com.cn/20260901/c679482046.shtml)) names the publisher as **词元星火**,
a wholly owned subsidiary of iFlytek (**reported**). 星火 (Xinghuo) is the Chinese name of iFlytek's
Spark model family, and 词元 is the standard Chinese term for a *token* in the NLP sense. XH reads as
the initials of Xinghuo and Token as a translation of 词元 (**reasoned**). Nothing in the model's files, its
Hugging Face org page or the launch coverage mentions a coin; search hits for similar names belong
to unrelated projects. The card's quickstart even asks for "the capital of Anhui Province", where
iFlytek is headquartered, and the model was "trained on Huawei Ascend clusters" (**reported**).

The Hugging Face org is called SparkLLM, has 18 members and no verified badge (**measured**, from the
org API). It published ten repositories in the series between 2026-08-24 and 2026-09-03: 4B and 1.7B,
each as a base model, an instruct model, GGUF, FP8 and INT8.

The GitHub repository in the post contains six files: README, licence, code of conduct, security
policy, `.gitignore` and two images (**measured**, shallow clone at commit `a24ca4e`). That is
normal for a model release. The thing to read is on Hugging Face.

## What config.json says the model is

`config.json` declares `Spark2_5ForCausalLM`, `model_type: spark2_5`. Not a Qwen or Llama class
with a new name: an architecture of its own. The numbers (**measured**):

| field | Spark-X2.5-4B |
|---|---|
| layers · hidden | 36 · 2,560 |
| attention heads · KV heads · head_dim | 16 · 4 · 256 |
| MLP | gated GELU, intermediate 10,240 |
| layer pattern | 3 sliding-window, then 1 full attention, repeated 9 times |
| sliding window | 512 tokens |
| RoPE, sliding layers | all 256 dims, theta 10,000 |
| RoPE, full layers | 64 of 256 dims (`partial_rotary_factor` 0.25), theta 5,000,000 |
| attention output gate | one sigmoid gate per head (`g_proj`, 16 outputs) |
| vocabulary · tied embeddings | 131,072 · yes |
| `max_position_embeddings` | 1,048,576 |

The publisher's own diagram and spec table say the same thing, which is the first consistency check
passing.

<Figure
  src="https://ai.thesatyajit.com/articles/spark-x2-5-4b/fig1.png"
  alt="Left: a block diagram with an embedding at the bottom, a repeated unit of three sliding-window attention blocks and one global attention block, each with RMSNorm and a GeGLU feed-forward, and an LM head at the top. Right: a spec table for Spark-X2.5-1.7B and 4B listing total parameters 1,707,657,216 and 4,112,079,360, 27 SWA layers and 9 GA layers for the 4B, 16 heads, 4 KV heads, head dimension 256, RoPE dimension 64, context 1M, window 512 and vocabulary 131,072."
  caption="The hybrid layout: three sliding-window layers to every full-attention layer, and the spec table the safetensors headers can be checked against (Spark-X2.5-4B model card, architecture figure)."
/>

## Count the tensors without downloading them

A safetensors file begins with 8 bytes giving the length of a JSON header, then the header itself,
which lists every tensor's name, dtype, shape and byte offsets. Two range requests per shard read the
whole layout of an 8 GB checkpoint:

```python
# read_header.py: the layout of a safetensors shard, a few KB per file
import json, struct, subprocess
def header(url):
    n = struct.unpack("<Q", subprocess.run(["curl", "-sL", "-r", "0-7", url],
                      capture_output=True).stdout)[0]
    raw = subprocess.run(["curl", "-sL", "-r", f"8-{8 + n - 1}", url],
                         capture_output=True).stdout
    return json.loads(raw)
```

Across the five shards (**measured**): **290 tensors, all BF16, 4,112,079,360 parameters**. That is
the card's total to the last digit. Subtract the 131,072 x 2,560 embedding (335,544,320) and you get
3,776,535,040, also the card's non-embedding figure exactly. 290 is 36 layers x 8 tensors, plus the
embedding and the final norm; there is no `lm_head` tensor because the embeddings are tied. The
shards total 8.23 GB on disk.

Each layer is a fused `q_k_v_proj` of shape 6,144 x 2,560 (4,096 query plus 2 x 1,024 key and value
rows), the 16-row `g_proj`, an `out_proj`, three MLP matrices and two RMSNorm vectors. The headers
and the config describe the same network.

## Is it someone else's model?

The commonest way a viral small model disappoints is that it is a well-known base with new
paperwork. The cheap test is shapes. I pulled configs for every plausible 2-4B base (**measured**):

| model | layers | hidden | heads / KV | head_dim | MLP | vocab |
|---|---|---|---|---|---|---|
| **Spark-X2.5-4B** | 36 | 2,560 | 16 / 4 | 256 | 10,240 | 131,072 |
| Qwen3-4B | 36 | 2,560 | 32 / 8 | 128 | 9,728 | 151,936 |
| Qwen3.5-4B | 32 | 2,560 | 16 / 4 | 256 | 9,216 | 248,320 |
| Gemma-4-E4B | 42 | 2,560 | 8 / 2 | 256 | 10,240 | 262,144 |
| Gemma-3-4B | 34 | 2,560 | 8 / 4 | 256 | 10,240 | 262,208 |
| Llama-3.2-3B | 28 | 3,072 | 24 / 8 | 128 | 8,192 | 128,256 |
| Ministral-3-3B | 26 | 3,072 | 32 / 8 | 128 | 9,216 | 131,072 |

Nothing matches. The design clearly borrows ideas: the 3:1 sliding-to-full pattern, the 0.25 partial
rotary factor and 16/4 heads at head_dim 256 sit close to Qwen3.5-4B's
recipe; the window of 512, GELU MLP at 10,240, and theta 10,000 on sliding layers sit close to
[Gemma 4's](/articles/gemma-4). Borrowed recipes are how the field works. Borrowed weights would be a
different claim.

The tokenizer settles part of it. It is a byte-level BPE with **131,072 entries**, and its ids do
not line up with any candidate: of the tokens it shares with DeepSeek-V3, Mistral's Tekken and Qwen3,
only 8, 5 and 68 sit at the same id (**measured**), which is byte-table coincidence. Its
pre-tokenizer is DeepSeek-V3's (the same digit and CJK splits, a near-identical main regex) plus an
extra per-digit split, and its special tokens use DeepSeek's `<｜end▁of▁sentence｜>` spelling. A
DeepSeek-style tokenizer *recipe*, trained into its own vocabulary. An embedding table cannot be
copied across vocabularies, so at minimum the embedding is new.

For the transformer blocks I went one step further. Four of the candidates have MLP rows of the
same width (2,560), so I range-requested layer 0's `gate_proj` from Spark-X2.5-4B and from each,
about 50 MB per model, and for every Spark row found the best-matching row anywhere in the other
matrix by absolute cosine similarity. As a positive control I ran the same test against
Spark-X2.5-4B-Base, which the card names as its base.

| Spark-X2.5-4B layer 0 vs | median best-match cosine | column-norm correlation |
|---|---|---|
| Spark-X2.5-4B-Base (control) | **0.987** | 0.996 |
| Qwen3-4B | 0.078 | 0.006 |
| Qwen3.5-4B | 0.078 | 0.000 |
| Gemma-4-E4B | 0.079 | 0.037 |
| Gemma-3-4B | 0.079 | -0.016 |

All **measured**. 0.078-0.079 is the noise floor: Gemma-4-E4B against Gemma-3-4B, two models nobody
claims are related, also scores 0.079. The control scores 0.987, so the test can see a fine-tune when
there is one. The instruct model is a fine-tune of its own base, and neither is a fine-tune of these
four.

What this is not: proof against every possible base, or against a deliberate disguise. Someone
determined could rotate the hidden basis and defeat a row match. But the vocabulary, the layer
inventory and the weights all point the same way (**reasoned**): this is a model trained by its
publisher, not a rename. The press coverage's ~20 trillion pretraining tokens is **reported**.

## Read the code you would be trusting

`auto_map` points `transformers` at two files: `configuration_spark.py` (4,329 bytes) and
`modeling_spark.py` (19,418 bytes). I read both in full (**measured**):

```bash
grep -nE "import|exec|eval|subprocess|socket|requests|urllib|http|base64|os\.|open\(" \
  modeling_spark.py configuration_spark.py
```

The only imports are `math`, `torch` and `transformers`; the only `http` is the Apache licence URL
in a comment. No network calls, no file access, no `exec`, `eval`, `pickle` or base64 blobs. It is a
plain PyTorch decoder: RMSNorm, a fused QKV projection, partial RoPE computed per layer type, eager
softmax attention with sliding-window masks from `transformers`' own helpers, the per-head sigmoid
output gate, a GELU-gated MLP, and a residual stream kept in float32. The one surprise is cosmetic:
the MLP raises a Chinese-language error if `hidden_act` is anything but `gelu`.

Better still, you do not need it. llama.cpp's `src/llama-arch.cpp` on master registers
`LLM_ARCH_SPARK2_5` as `spark2_5` (**measured**, read from the source), and the card says Ollama
0.34.1 and LM Studio runtime 2.34.0 added it natively (**reported**). The GGUF path runs C++ from
those projects, not Python from the model repo. The post's "works with Ollama, LM Studio, llama.cpp"
holds. The `transformers` path, and the vLLM command in the card, do need `--trust-remote-code`;
having read the code, I would be comfortable with that here, but read it again if the repository
changes.

## Does "native 1M context" hold?

Two separate questions hide in that phrase: can the architecture *address* a million positions, and
was it *trained* on them.

The first is checkable from the config. There is no `rope_scaling` block, so the 1,048,576 is not a
YaRN stretch applied at inference. Sliding layers only ever see 512 tokens back, so their RoPE never
has to reach far. The nine full-attention layers rotate only 64 of their 256 dimensions; the other
192 carry no position at all. On the rotated 64, theta 5,000,000 puts the slowest pair's wavelength
at about 19.4 million positions (**reasoned**: $2\pi \cdot 5{,}000{,}000^{62/64}$), comfortably past
one million, so positions out to the claimed length remain distinguishable. The config is
**consistent with** a native 1M context.

The second is not checkable from files. The card says long context came from "a dedicated training
stage comprising hundreds of billions of tokens, with sequence lengths extending to 1M tokens"
(**reported**). One small inconsistency: `tokenizer_config.json` sets `model_max_length` to 131,072,
which will make some tooling warn or truncate at 128K (**measured**). And the card's one long-context
reasoning benchmark, AA-LCR, has Spark-X2.5-4B at 56.3 against Qwen3.5-4B's 57.0 (**reported**), so
on the long-context number the vendor chose to publish, it is level with the model it beats nearly
everywhere else.

## Does it run on a 16 GB laptop?

The weights do, easily. XHToken's own GGUF files are 2.60 GB at Q4_K_M, 4.38 GB at Q8_0 and 8.23 GB
at BF16 (**measured**, from the repository listing). A 4B model at 4 bits fits on a phone.

The million tokens do not. In a sliding-window layer the KV cache is capped at 512 positions; in a
full-attention layer it grows with the context. Per token, per full layer, the cache holds a key and
a value for 4 KV heads of 256 dimensions:

$$
\text{KV bytes per token} = 9 \times 2 \times 4 \times 256 \times 2 = 36{,}864
$$

(9 full layers; K and V; 4 heads; 256 dims; 2 bytes in bf16): 36,864 bytes per token. At 1,048,576 tokens that is **36 GiB**
(**reasoned**). The 27 sliding layers add 54 MiB in total, whatever the context. If all 36 layers
were full attention, the cache would be 144 GiB, so the hybrid layout is a real fourfold saving. It
is still more than twice a 16 GB laptop's entire memory.

<KvBudget />

With 4 GiB set aside for the operating system and Q4_K_M weights, about 9.6 GiB remains for the
cache: roughly 279,000 tokens in bf16, 525,000 with llama.cpp's 8-bit `q8_0` cache, and 992,000 with
the 4-bit `q4_0` cache (**reasoned**, ignoring compute buffers, so these are upper bounds). A
million tokens on 16 GB is just about reachable with a lossy 4-bit cache. Nothing on the card
measures what that does to quality. "Runs on a 16 GB laptop" and "1M context" are both true. They are
not true at the same time.

## The benchmarks

The card's table has 21 rows and eight columns: the two Spark models against Qwen3.5 at 9B, 4B and
2B, Gemma 4 at 12B, E4B and E2B. Spark-X2.5-4B beats Qwen3.5-4B on 19 of 21 and has the best score
in the whole table on 12 (**measured**, tallied from the card; the scores themselves are
**reported**). On agent benchmarks the margins are large: 30.4 on τ³-bench against 6.7 for
Qwen3.5-4B, 40.9 on BrowseComp against 14.3.

<Figure
  src="https://ai.thesatyajit.com/articles/spark-x2-5-4b/fig2.png"
  alt="Eight bar charts comparing seven models on τ³-bench, MCP-Atlas, BrowseComp, SciCode, AIME 2026, HMMT Feb 2026, HLE and IFBench. Spark-X2.5-4B's dark-blue bar is the tallest among the 4B-class models in every panel, for example 30.4 on τ³-bench, 54.6 on MCP-Atlas and 40.9 on BrowseComp; Qwen3.5-9B in grey is taller on HLE."
  caption="The headline chart. These are the vendor's numbers, not re-run, and it shows 8 of the 21 benchmarks in the card's table, without the Gemma4-12B column (Spark-X2.5-4B model card, benchmark figure; rendered from the card's SVG)."
/>

Four things to weigh before repeating any of it:

1. **Who ran the rivals.** The footnote says an asterisk marks "reported results from
   publicly‑released model cards / papers". Of 114 rival cells, 36 have one (**measured**). The
   other 78 were, by the card's own key, not copied from a published card, which in practice means
   the vendor ran them, under its harness and settings. That is normal, and it is also where most
   benchmark disagreement comes from.
2. **The chart is a selection.** It shows eight of the 21 rows, and in each of the eight
   Spark-X2.5-4B leads its size class. The rows where it trails Qwen3.5-4B (τ²-bench 75.1 vs 79.9,
   AA-LCR 56.3 vs 57.0) are in the table but not the picture, and so is the whole Gemma4-12B column,
   which beats it on SciCode (39.8 vs 34.7) and IFEval (94.8 vs 93.0). The chart also prints 12.2
   for HLE where the table says 12.3 (**measured**), a small slip, but a sign the two were made
   separately.
3. **One pattern breaks.** SWE-Bench Pro is the harder sibling of SWE-Bench Verified, and every
   other model in the table with both scores is lower on Pro: Qwen3.5-4B 29.4 vs 38.8, Gemma4-12B 21.9 vs 44.2.
   Spark-X2.5-4B is the only one that scores higher, 44.4 vs 41.6 (**measured** from the table). It
   may be a harness difference, a scaffold tuned for Pro, or a real strength. It is the row I would
   reproduce first.
4. **There is nothing to rerun.** No evaluation code, configs, prompts or logs ship with the release,
   and I found no independent evaluation in the five weeks since launch.

The post-training recipe is at least described. Supervised fine-tuning, then reinforcement learning
per domain to produce general, reasoning, agent and code teachers, then what the card calls MOPD,
multi-teacher on-policy distillation into one student with a reverse-KL objective. On-policy distillation is covered in
[Inkling-Small](/articles/inkling-small); nothing in Spark's files contradicts the description, and
nothing in them can confirm it either.

<Figure
  src="https://ai.thesatyajit.com/articles/spark-x2-5-4b/fig3.png"
  alt="A three-panel pipeline. Supervised fine-tuning on chat, instruction-following, code and agentic data produces an initial policy. Domain expert training uses reinforcement learning with environment interaction to produce general, reasoning, agent and code experts. Multi-teacher on-policy distillation routes prompts to the expert teachers and trains a unified student on its own trajectories with reverse KL."
  caption="Post-training as the publisher describes it: SFT, per-domain RL experts, then multi-teacher on-policy distillation into one model (Spark-X2.5-4B model card, post-training pipeline figure)."
/>

## "#1 on Hugging Face"

On 2026-10-06 the Hugging Face API gives Spark-X2.5-4B a `trendingScore` of 33, which places it
**#149** on the trending list; the top model scores 1,427 (**measured**). The repo has 1,380 likes,
37,787 downloads in the last 30 days and 43,329 all-time; the official GGUF repo has 478,743
downloads (**measured**). That is a popular small model. It is not, today, the top of anything.

The repository was created on 2026-08-24 and launched publicly on 2026-09-01. The post is from
2026-10-05, five weeks later, and says it "just hit #1". It may well have topped the trending list
in its launch week; the API does not keep history, and I could not confirm it either way
(**unverified**). One reply under the post, in Chinese, asks (my translation) whether the author
actually downloaded and tested it, and says that if not, this is an ad. The post does not say.

## The checklist

Here is the whole procedure as a list, with what Spark-X2.5-4B showed at each step. None of it needs
a GPU. My downloads came to about 350 MB, nearly all of it the optional weight comparison.

<VettingChecklist />

1. **Publisher (1 min).** Org API, then search in the publisher's own language. Here: an iFlytek
   subsidiary; the "token" is the NLP kind. Pass.
2. **config.json (2 min).** Architecture class, layers, context, scaling blocks, and fields that
   disagree with each other. Here: a custom `spark2_5` class matching the card, plus one stray
   `model_max_length`. Pass.
3. **Safetensors headers (2 min).** Range-request the 8-byte length and the JSON. Count parameters.
   Here: 4,112,079,360, exact. Pass.
4. **Base match (4 min).** Compare shapes and vocab with likely bases; if any match, compare one
   weight tensor, with the declared base as a positive control. Here: no match on any axis, noise
   floor on weights, 0.987 against its own base. Pass.
5. **Remote code (3 min).** Read every file `auto_map` names; grep for network, filesystem,
   subprocess and dynamic execution. Check whether llama.cpp supports the architecture natively, so
   you can avoid running it at all. Here: clean, and avoidable. Pass.
6. **Context claim (1 min).** RoPE settings, scaling, which layers are global, and the KV cache the
   claim implies. Here: consistent, 36 GiB at full length. True with a cost.
7. **Benchmarks (1 min).** Who ran the rival numbers, whether the chart is a subset, any eval code,
   any broken pattern. Here: vendor-run, selected, no code, one odd row. Unclear.
8. **Popularity (1 min).** Trending score and rank today, downloads, and the gap between release and
   post. Here: #149 today. Stale or unverified.

The same procedure has caught real problems elsewhere: a benchmark table that could not add up and a
LoRA nobody mentioned in [BTL-4](/articles/btl-4), and a "mixture of architectures" that was public
checkpoints and a tool loop in [Interfaze 1 Lite](/articles/interfaze-1-lite). Here it mostly found
a model that is what it says it is.

## What I would tell someone who saw the post

Spark-X2.5-4B is a genuine, original small model from a major Chinese speech-and-language company,
with Apache-2.0 weights, a readable and harmless loader, an efficient hybrid-attention design and
native support in the local runtimes the post names. That is more than many viral models can say.

The post overstates three things. The 1M context is real in the architecture and reported in the
training, but it needs about 36 GiB of cache in bf16, so the laptop and the million tokens are two
different deployments. The coding and agent strength is the vendor's measurement, with a selected
chart and no harness to rerun. And "#1 on Hugging Face" was, at best, true a month earlier.

If the agent numbers are even half right, a 4B model at 30.4 on τ³-bench is worth an afternoon. Run
the GGUF, at the context you actually need, on the tasks you actually have, and measure. For how the
sliding-window and global layers split the cache, see the [attention and KV cache explainer](/architectures/attention-kv)
and [RoPE](/architectures/rope).
