~/satyajit

Spark-X2.5-4B: fifteen minutes with the files behind a viral model

mdjsonmcp

2026-10-06 · 18 min · llm · open-weights · small-models · on-device · long-context · hybrid-attention · evaluation · benchmarks · security

On 2026-10-05 a post on X with about 12,300 views announced "a new open-source model that just hit #1 on Hugging Face and it runs fully local on a 16GB laptop". The bullets: 4 billion parameters, a "native 1 million token context window", "strong coding + agent capabilities", and support for Ollama, LM Studio and llama.cpp. The link went to a GitHub repository. The video under the post scrolls that repository's README and nothing else.

The model is Spark-X2.5-4B. Its Hugging Face repo carries the custom_code tag: transformers loads it with trust_remote_code=True, which runs Python that the publisher wrote. That is exactly the situation where a viral post and a checkpoint deserve fifteen minutes of reading before anyone runs anything.

This is that reading, done end to end, followed by the checklist so you can repeat it on the next one. Short version: the model is real, original and well built. Most of the post's claims are true in a narrower sense than they sound, and one of them I could not verify at all.

XHToken/Spark-X2.5-4B@0bcb356 · snapshot 2026-10-06
parameters
4.11B
repo size
8.23 GB
architecture
Spark2_5ForCausalLM
task
text-generation
library
transformers
license
apache-2.0
safetensors
5 shards
largest file
1.99 GB
files
23
downloads
37.8K
likes
1.4K
parameters by dtype
BF164.11B
llmsparkx2_5agent

repo last modified 2026-09-15

A frame from the video attached to the X post: the GitHub page for XHToken/Spark-X2.5 in a dark theme, listing an assets folder, .gitignore, CODE_OF_CONDUCT.md, LICENSE, README.md and SECURITY.md.
The post's video opens on the GitHub repository it links. The file list is the whole repo: a README, a licence and policy files, no code and no weights. The weights live on Hugging Face (frame from the video attached to the X post by @Forhanvv, cropped to the browser window).

Who published it

Start with the name, because "XHToken" reads like a crypto project and that matters for trust. It is not one. Chinese launch coverage on 2026-09-01 (Beijing News, 10jqka) names the publisher as 词元星火, a wholly owned subsidiary of iFlytek (reported). 星火 (Xinghuo) is the Chinese name of iFlytek's Spark model family, and 词元 is the standard Chinese term for a token in the NLP sense. XH reads as the initials of Xinghuo and Token as a translation of 词元 (reasoned). Nothing in the model's files, its Hugging Face org page or the launch coverage mentions a coin; search hits for similar names belong to unrelated projects. The card's quickstart even asks for "the capital of Anhui Province", where iFlytek is headquartered, and the model was "trained on Huawei Ascend clusters" (reported).

The Hugging Face org is called SparkLLM, has 18 members and no verified badge (measured, from the org API). It published ten repositories in the series between 2026-08-24 and 2026-09-03: 4B and 1.7B, each as a base model, an instruct model, GGUF, FP8 and INT8.

The GitHub repository in the post contains six files: README, licence, code of conduct, security policy, .gitignore and two images (measured, shallow clone at commit a24ca4e). That is normal for a model release. The thing to read is on Hugging Face.

What config.json says the model is

config.json declares Spark2_5ForCausalLM, model_type: spark2_5. Not a Qwen or Llama class with a new name: an architecture of its own. The numbers (measured):

fieldSpark-X2.5-4B
layers · hidden36 · 2,560
attention heads · KV heads · head_dim16 · 4 · 256
MLPgated GELU, intermediate 10,240
layer pattern3 sliding-window, then 1 full attention, repeated 9 times
sliding window512 tokens
RoPE, sliding layersall 256 dims, theta 10,000
RoPE, full layers64 of 256 dims (partial_rotary_factor 0.25), theta 5,000,000
attention output gateone sigmoid gate per head (g_proj, 16 outputs)
vocabulary · tied embeddings131,072 · yes
max_position_embeddings1,048,576

The publisher's own diagram and spec table say the same thing, which is the first consistency check passing.

Left: a block diagram with an embedding at the bottom, a repeated unit of three sliding-window attention blocks and one global attention block, each with RMSNorm and a GeGLU feed-forward, and an LM head at the top. Right: a spec table for Spark-X2.5-1.7B and 4B listing total parameters 1,707,657,216 and 4,112,079,360, 27 SWA layers and 9 GA layers for the 4B, 16 heads, 4 KV heads, head dimension 256, RoPE dimension 64, context 1M, window 512 and vocabulary 131,072.
The hybrid layout: three sliding-window layers to every full-attention layer, and the spec table the safetensors headers can be checked against (Spark-X2.5-4B model card, architecture figure).

Count the tensors without downloading them

A safetensors file begins with 8 bytes giving the length of a JSON header, then the header itself, which lists every tensor's name, dtype, shape and byte offsets. Two range requests per shard read the whole layout of an 8 GB checkpoint:

# read_header.py: the layout of a safetensors shard, a few KB per file
import json, struct, subprocess
def header(url):
    n = struct.unpack("<Q", subprocess.run(["curl", "-sL", "-r", "0-7", url],
                      capture_output=True).stdout)[0]
    raw = subprocess.run(["curl", "-sL", "-r", f"8-{8 + n - 1}", url],
                         capture_output=True).stdout
    return json.loads(raw)

Across the five shards (measured): 290 tensors, all BF16, 4,112,079,360 parameters. That is the card's total to the last digit. Subtract the 131,072 x 2,560 embedding (335,544,320) and you get 3,776,535,040, also the card's non-embedding figure exactly. 290 is 36 layers x 8 tensors, plus the embedding and the final norm; there is no lm_head tensor because the embeddings are tied. The shards total 8.23 GB on disk.

Each layer is a fused q_k_v_proj of shape 6,144 x 2,560 (4,096 query plus 2 x 1,024 key and value rows), the 16-row g_proj, an out_proj, three MLP matrices and two RMSNorm vectors. The headers and the config describe the same network.

Is it someone else's model?

The commonest way a viral small model disappoints is that it is a well-known base with new paperwork. The cheap test is shapes. I pulled configs for every plausible 2-4B base (measured):

modellayershiddenheads / KVhead_dimMLPvocab
Spark-X2.5-4B362,56016 / 425610,240131,072
Qwen3-4B362,56032 / 81289,728151,936
Qwen3.5-4B322,56016 / 42569,216248,320
Gemma-4-E4B422,5608 / 225610,240262,144
Gemma-3-4B342,5608 / 425610,240262,208
Llama-3.2-3B283,07224 / 81288,192128,256
Ministral-3-3B263,07232 / 81289,216131,072

Nothing matches. The design clearly borrows ideas: the 3:1 sliding-to-full pattern, the 0.25 partial rotary factor and 16/4 heads at head_dim 256 sit close to Qwen3.5-4B's recipe; the window of 512, GELU MLP at 10,240, and theta 10,000 on sliding layers sit close to Gemma 4's. Borrowed recipes are how the field works. Borrowed weights would be a different claim.

The tokenizer settles part of it. It is a byte-level BPE with 131,072 entries, and its ids do not line up with any candidate: of the tokens it shares with DeepSeek-V3, Mistral's Tekken and Qwen3, only 8, 5 and 68 sit at the same id (measured), which is byte-table coincidence. Its pre-tokenizer is DeepSeek-V3's (the same digit and CJK splits, a near-identical main regex) plus an extra per-digit split, and its special tokens use DeepSeek's <|end▁of▁sentence|> spelling. A DeepSeek-style tokenizer recipe, trained into its own vocabulary. An embedding table cannot be copied across vocabularies, so at minimum the embedding is new.

For the transformer blocks I went one step further. Four of the candidates have MLP rows of the same width (2,560), so I range-requested layer 0's gate_proj from Spark-X2.5-4B and from each, about 50 MB per model, and for every Spark row found the best-matching row anywhere in the other matrix by absolute cosine similarity. As a positive control I ran the same test against Spark-X2.5-4B-Base, which the card names as its base.

Spark-X2.5-4B layer 0 vsmedian best-match cosinecolumn-norm correlation
Spark-X2.5-4B-Base (control)0.9870.996
Qwen3-4B0.0780.006
Qwen3.5-4B0.0780.000
Gemma-4-E4B0.0790.037
Gemma-3-4B0.079-0.016

All measured. 0.078-0.079 is the noise floor: Gemma-4-E4B against Gemma-3-4B, two models nobody claims are related, also scores 0.079. The control scores 0.987, so the test can see a fine-tune when there is one. The instruct model is a fine-tune of its own base, and neither is a fine-tune of these four.

What this is not: proof against every possible base, or against a deliberate disguise. Someone determined could rotate the hidden basis and defeat a row match. But the vocabulary, the layer inventory and the weights all point the same way (reasoned): this is a model trained by its publisher, not a rename. The press coverage's ~20 trillion pretraining tokens is reported.

Read the code you would be trusting

auto_map points transformers at two files: configuration_spark.py (4,329 bytes) and modeling_spark.py (19,418 bytes). I read both in full (measured):

grep -nE "import|exec|eval|subprocess|socket|requests|urllib|http|base64|os\.|open\(" \
  modeling_spark.py configuration_spark.py

The only imports are math, torch and transformers; the only http is the Apache licence URL in a comment. No network calls, no file access, no exec, eval, pickle or base64 blobs. It is a plain PyTorch decoder: RMSNorm, a fused QKV projection, partial RoPE computed per layer type, eager softmax attention with sliding-window masks from transformers' own helpers, the per-head sigmoid output gate, a GELU-gated MLP, and a residual stream kept in float32. The one surprise is cosmetic: the MLP raises a Chinese-language error if hidden_act is anything but gelu.

Better still, you do not need it. llama.cpp's src/llama-arch.cpp on master registers LLM_ARCH_SPARK2_5 as spark2_5 (measured, read from the source), and the card says Ollama 0.34.1 and LM Studio runtime 2.34.0 added it natively (reported). The GGUF path runs C++ from those projects, not Python from the model repo. The post's "works with Ollama, LM Studio, llama.cpp" holds. The transformers path, and the vLLM command in the card, do need --trust-remote-code; having read the code, I would be comfortable with that here, but read it again if the repository changes.

Does "native 1M context" hold?

Two separate questions hide in that phrase: can the architecture address a million positions, and was it trained on them.

The first is checkable from the config. There is no rope_scaling block, so the 1,048,576 is not a YaRN stretch applied at inference. Sliding layers only ever see 512 tokens back, so their RoPE never has to reach far. The nine full-attention layers rotate only 64 of their 256 dimensions; the other 192 carry no position at all. On the rotated 64, theta 5,000,000 puts the slowest pair's wavelength at about 19.4 million positions (reasoned: 2π⋅5,000,00062/642\pi \cdot 5{,}000{,}000^{62/64}), comfortably past one million, so positions out to the claimed length remain distinguishable. The config is consistent with a native 1M context.

The second is not checkable from files. The card says long context came from "a dedicated training stage comprising hundreds of billions of tokens, with sequence lengths extending to 1M tokens" (reported). One small inconsistency: tokenizer_config.json sets model_max_length to 131,072, which will make some tooling warn or truncate at 128K (measured). And the card's one long-context reasoning benchmark, AA-LCR, has Spark-X2.5-4B at 56.3 against Qwen3.5-4B's 57.0 (reported), so on the long-context number the vendor chose to publish, it is level with the model it beats nearly everywhere else.

Does it run on a 16 GB laptop?

The weights do, easily. XHToken's own GGUF files are 2.60 GB at Q4_K_M, 4.38 GB at Q8_0 and 8.23 GB at BF16 (measured, from the repository listing). A 4B model at 4 bits fits on a phone.

The million tokens do not. In a sliding-window layer the KV cache is capped at 512 positions; in a full-attention layer it grows with the context. Per token, per full layer, the cache holds a key and a value for 4 KV heads of 256 dimensions:

KV bytes per token=9×2×4×256×2=36,864\text{KV bytes per token} = 9 \times 2 \times 4 \times 256 \times 2 = 36{,}864

(9 full layers; K and V; 4 heads; 256 dims; 2 bytes in bf16): 36,864 bytes per token. At 1,048,576 tokens that is 36 GiB (reasoned). The 27 sliding layers add 54 MiB in total, whatever the context. If all 36 layers were full attention, the cache would be 144 GiB, so the hybrid layout is a real fourfold saving. It is still more than twice a 16 GB laptop's entire memory.

Spark-X2.5-4B · memory at a given contextdoes not fit in 16 GiB
context length1M tokens (1,048,576)
weights
KV
0dashed line = 16 GiB · bar = 42.5 GiB
OS + apps (assumed)
4.00 GiB
weights
2.42 GiB
27 sliding layers
0.05 GiB
9 full layers
36.0 GiB

The sliding-window layers stop growing at 512 tokens, so past a few thousand tokens almost all of the cache is the nine full-attention layers. If all 36 layers were full attention, this cache would be 144.0 GiB instead of 36.1 GiB. The hybrid layout is a real 4x saving. It still does not make a million bf16 tokens fit in a laptop: only a 4-bit cache gets near, and a lossy cache is a quality trade the card does not measure.

With 4 GiB set aside for the operating system and Q4_K_M weights, about 9.6 GiB remains for the cache: roughly 279,000 tokens in bf16, 525,000 with llama.cpp's 8-bit q8_0 cache, and 992,000 with the 4-bit q4_0 cache (reasoned, ignoring compute buffers, so these are upper bounds). A million tokens on 16 GB is just about reachable with a lossy 4-bit cache. Nothing on the card measures what that does to quality. "Runs on a 16 GB laptop" and "1M context" are both true. They are not true at the same time.

The benchmarks

The card's table has 21 rows and eight columns: the two Spark models against Qwen3.5 at 9B, 4B and 2B, Gemma 4 at 12B, E4B and E2B. Spark-X2.5-4B beats Qwen3.5-4B on 19 of 21 and has the best score in the whole table on 12 (measured, tallied from the card; the scores themselves are reported). On agent benchmarks the margins are large: 30.4 on τ³-bench against 6.7 for Qwen3.5-4B, 40.9 on BrowseComp against 14.3.

Eight bar charts comparing seven models on τ³-bench, MCP-Atlas, BrowseComp, SciCode, AIME 2026, HMMT Feb 2026, HLE and IFBench. Spark-X2.5-4B's dark-blue bar is the tallest among the 4B-class models in every panel, for example 30.4 on τ³-bench, 54.6 on MCP-Atlas and 40.9 on BrowseComp; Qwen3.5-9B in grey is taller on HLE.
The headline chart. These are the vendor's numbers, not re-run, and it shows 8 of the 21 benchmarks in the card's table, without the Gemma4-12B column (Spark-X2.5-4B model card, benchmark figure; rendered from the card's SVG).

Four things to weigh before repeating any of it:

  1. Who ran the rivals. The footnote says an asterisk marks "reported results from publicly‑released model cards / papers". Of 114 rival cells, 36 have one (measured). The other 78 were, by the card's own key, not copied from a published card, which in practice means the vendor ran them, under its harness and settings. That is normal, and it is also where most benchmark disagreement comes from.
  2. The chart is a selection. It shows eight of the 21 rows, and in each of the eight Spark-X2.5-4B leads its size class. The rows where it trails Qwen3.5-4B (τ²-bench 75.1 vs 79.9, AA-LCR 56.3 vs 57.0) are in the table but not the picture, and so is the whole Gemma4-12B column, which beats it on SciCode (39.8 vs 34.7) and IFEval (94.8 vs 93.0). The chart also prints 12.2 for HLE where the table says 12.3 (measured), a small slip, but a sign the two were made separately.
  3. One pattern breaks. SWE-Bench Pro is the harder sibling of SWE-Bench Verified, and every other model in the table with both scores is lower on Pro: Qwen3.5-4B 29.4 vs 38.8, Gemma4-12B 21.9 vs 44.2. Spark-X2.5-4B is the only one that scores higher, 44.4 vs 41.6 (measured from the table). It may be a harness difference, a scaffold tuned for Pro, or a real strength. It is the row I would reproduce first.
  4. There is nothing to rerun. No evaluation code, configs, prompts or logs ship with the release, and I found no independent evaluation in the five weeks since launch.

The post-training recipe is at least described. Supervised fine-tuning, then reinforcement learning per domain to produce general, reasoning, agent and code teachers, then what the card calls MOPD, multi-teacher on-policy distillation into one student with a reverse-KL objective. On-policy distillation is covered in Inkling-Small; nothing in Spark's files contradicts the description, and nothing in them can confirm it either.

A three-panel pipeline. Supervised fine-tuning on chat, instruction-following, code and agentic data produces an initial policy. Domain expert training uses reinforcement learning with environment interaction to produce general, reasoning, agent and code experts. Multi-teacher on-policy distillation routes prompts to the expert teachers and trains a unified student on its own trajectories with reverse KL.
Post-training as the publisher describes it: SFT, per-domain RL experts, then multi-teacher on-policy distillation into one model (Spark-X2.5-4B model card, post-training pipeline figure).

"#1 on Hugging Face"

On 2026-10-06 the Hugging Face API gives Spark-X2.5-4B a trendingScore of 33, which places it #149 on the trending list; the top model scores 1,427 (measured). The repo has 1,380 likes, 37,787 downloads in the last 30 days and 43,329 all-time; the official GGUF repo has 478,743 downloads (measured). That is a popular small model. It is not, today, the top of anything.

The repository was created on 2026-08-24 and launched publicly on 2026-09-01. The post is from 2026-10-05, five weeks later, and says it "just hit #1". It may well have topped the trending list in its launch week; the API does not keep history, and I could not confirm it either way (unverified). One reply under the post, in Chinese, asks (my translation) whether the author actually downloaded and tested it, and says that if not, this is an ad. The post does not say.

The checklist

Here is the whole procedure as a list, with what Spark-X2.5-4B showed at each step. None of it needs a GPU. My downloads came to about 350 MB, nearly all of it the optional weight comparison.

vetting a viral model · Spark-X2.5-4B5 pass2 unclear1 fail
  1. fetch
    GET /api/organizations/<org>/overview, then a search in the publisher's own language
    look for
    A known lab, a history of releases, press that names a company rather than a handle.
    found
    XHToken is 词元星火, which Chinese press describes as a wholly owned iFlytek subsidiary. 词元 is the Chinese word for a token in the NLP sense. No coin is mentioned anywhere in the model's files, org page or launch coverage. The org has no verified badge.
  1. Publisher (1 min). Org API, then search in the publisher's own language. Here: an iFlytek subsidiary; the "token" is the NLP kind. Pass.
  2. config.json (2 min). Architecture class, layers, context, scaling blocks, and fields that disagree with each other. Here: a custom spark2_5 class matching the card, plus one stray model_max_length. Pass.
  3. Safetensors headers (2 min). Range-request the 8-byte length and the JSON. Count parameters. Here: 4,112,079,360, exact. Pass.
  4. Base match (4 min). Compare shapes and vocab with likely bases; if any match, compare one weight tensor, with the declared base as a positive control. Here: no match on any axis, noise floor on weights, 0.987 against its own base. Pass.
  5. Remote code (3 min). Read every file auto_map names; grep for network, filesystem, subprocess and dynamic execution. Check whether llama.cpp supports the architecture natively, so you can avoid running it at all. Here: clean, and avoidable. Pass.
  6. Context claim (1 min). RoPE settings, scaling, which layers are global, and the KV cache the claim implies. Here: consistent, 36 GiB at full length. True with a cost.
  7. Benchmarks (1 min). Who ran the rival numbers, whether the chart is a subset, any eval code, any broken pattern. Here: vendor-run, selected, no code, one odd row. Unclear.
  8. Popularity (1 min). Trending score and rank today, downloads, and the gap between release and post. Here: #149 today. Stale or unverified.

The same procedure has caught real problems elsewhere: a benchmark table that could not add up and a LoRA nobody mentioned in BTL-4, and a "mixture of architectures" that was public checkpoints and a tool loop in Interfaze 1 Lite. Here it mostly found a model that is what it says it is.

What I would tell someone who saw the post

Spark-X2.5-4B is a genuine, original small model from a major Chinese speech-and-language company, with Apache-2.0 weights, a readable and harmless loader, an efficient hybrid-attention design and native support in the local runtimes the post names. That is more than many viral models can say.

The post overstates three things. The 1M context is real in the architecture and reported in the training, but it needs about 36 GiB of cache in bf16, so the laptop and the million tokens are two different deployments. The coding and agent strength is the vendor's measurement, with a selected chart and no harness to rerun. And "#1 on Hugging Face" was, at best, true a month earlier.

If the agent numbers are even half right, a 4B model at 30.4 on τ³-bench is worth an afternoon. Run the GGUF, at the context you actually need, on the tasks you actually have, and measure. For how the sliding-window and global layers split the cache, see the attention and KV cache explainer and RoPE.

Cite this article

For attribution, please use the following reference or BibTeX:

Satyajit Ghana, "Spark-X2.5-4B: fifteen minutes with the files behind a viral model", ai.thesatyajit.com, October 2026.

bibtex
@misc{ghana2026sparkx254b,
  author = {Satyajit Ghana},
  title  = {Spark-X2.5-4B: fifteen minutes with the files behind a viral model},
  url    = {https://ai.thesatyajit.com/articles/spark-x2-5-4b},
  year   = {2026}
}
share