2026-10-06 · 18 min · llm · open-weights · small-models · on-device · long-context · hybrid-attention · evaluation · benchmarks · security
On 2026-10-05 a post on X with about 12,300 views announced "a new open-source model that just hit #1 on Hugging Face and it runs fully local on a 16GB laptop". The bullets: 4 billion parameters, a "native 1 million token context window", "strong coding + agent capabilities", and support for Ollama, LM Studio and llama.cpp. The link went to a GitHub repository. The video under the post scrolls that repository's README and nothing else.
The model is Spark-X2.5-4B. Its Hugging Face repo
carries the custom_code tag: transformers loads it with trust_remote_code=True, which runs
Python that the publisher wrote. That is exactly the situation where
a viral post and a checkpoint deserve fifteen minutes of reading before anyone runs anything.
This is that reading, done end to end, followed by the checklist so you can repeat it on the next one. Short version: the model is real, original and well built. Most of the post's claims are true in a narrower sense than they sound, and one of them I could not verify at all.
- architecture
- Spark2_5ForCausalLM
- task
- text-generation
- library
- transformers
- license
- apache-2.0
- safetensors
- 5 shards
- largest file
- 1.99 GB
- files
- 23
- downloads
- 37.8K
- likes
- 1.4K
repo last modified 2026-09-15

Who published it
Start with the name, because "XHToken" reads like a crypto project and that matters for trust. It is not one. Chinese launch coverage on 2026-09-01 (Beijing News, 10jqka) names the publisher as 词元星火, a wholly owned subsidiary of iFlytek (reported). 星火 (Xinghuo) is the Chinese name of iFlytek's Spark model family, and 词元 is the standard Chinese term for a token in the NLP sense. XH reads as the initials of Xinghuo and Token as a translation of 词元 (reasoned). Nothing in the model's files, its Hugging Face org page or the launch coverage mentions a coin; search hits for similar names belong to unrelated projects. The card's quickstart even asks for "the capital of Anhui Province", where iFlytek is headquartered, and the model was "trained on Huawei Ascend clusters" (reported).
The Hugging Face org is called SparkLLM, has 18 members and no verified badge (measured, from the org API). It published ten repositories in the series between 2026-08-24 and 2026-09-03: 4B and 1.7B, each as a base model, an instruct model, GGUF, FP8 and INT8.
The GitHub repository in the post contains six files: README, licence, code of conduct, security
policy, .gitignore and two images (measured, shallow clone at commit a24ca4e). That is
normal for a model release. The thing to read is on Hugging Face.
What config.json says the model is
config.json declares Spark2_5ForCausalLM, model_type: spark2_5. Not a Qwen or Llama class
with a new name: an architecture of its own. The numbers (measured):
| field | Spark-X2.5-4B |
|---|---|
| layers · hidden | 36 · 2,560 |
| attention heads · KV heads · head_dim | 16 · 4 · 256 |
| MLP | gated GELU, intermediate 10,240 |
| layer pattern | 3 sliding-window, then 1 full attention, repeated 9 times |
| sliding window | 512 tokens |
| RoPE, sliding layers | all 256 dims, theta 10,000 |
| RoPE, full layers | 64 of 256 dims (partial_rotary_factor 0.25), theta 5,000,000 |
| attention output gate | one sigmoid gate per head (g_proj, 16 outputs) |
| vocabulary · tied embeddings | 131,072 · yes |
max_position_embeddings | 1,048,576 |
The publisher's own diagram and spec table say the same thing, which is the first consistency check passing.

Count the tensors without downloading them
A safetensors file begins with 8 bytes giving the length of a JSON header, then the header itself, which lists every tensor's name, dtype, shape and byte offsets. Two range requests per shard read the whole layout of an 8 GB checkpoint:
# read_header.py: the layout of a safetensors shard, a few KB per file
import json, struct, subprocess
def header(url):
n = struct.unpack("<Q", subprocess.run(["curl", "-sL", "-r", "0-7", url],
capture_output=True).stdout)[0]
raw = subprocess.run(["curl", "-sL", "-r", f"8-{8 + n - 1}", url],
capture_output=True).stdout
return json.loads(raw)Across the five shards (measured): 290 tensors, all BF16, 4,112,079,360 parameters. That is
the card's total to the last digit. Subtract the 131,072 x 2,560 embedding (335,544,320) and you get
3,776,535,040, also the card's non-embedding figure exactly. 290 is 36 layers x 8 tensors, plus the
embedding and the final norm; there is no lm_head tensor because the embeddings are tied. The
shards total 8.23 GB on disk.
Each layer is a fused q_k_v_proj of shape 6,144 x 2,560 (4,096 query plus 2 x 1,024 key and value
rows), the 16-row g_proj, an out_proj, three MLP matrices and two RMSNorm vectors. The headers
and the config describe the same network.
Is it someone else's model?
The commonest way a viral small model disappoints is that it is a well-known base with new paperwork. The cheap test is shapes. I pulled configs for every plausible 2-4B base (measured):
| model | layers | hidden | heads / KV | head_dim | MLP | vocab |
|---|---|---|---|---|---|---|
| Spark-X2.5-4B | 36 | 2,560 | 16 / 4 | 256 | 10,240 | 131,072 |
| Qwen3-4B | 36 | 2,560 | 32 / 8 | 128 | 9,728 | 151,936 |
| Qwen3.5-4B | 32 | 2,560 | 16 / 4 | 256 | 9,216 | 248,320 |
| Gemma-4-E4B | 42 | 2,560 | 8 / 2 | 256 | 10,240 | 262,144 |
| Gemma-3-4B | 34 | 2,560 | 8 / 4 | 256 | 10,240 | 262,208 |
| Llama-3.2-3B | 28 | 3,072 | 24 / 8 | 128 | 8,192 | 128,256 |
| Ministral-3-3B | 26 | 3,072 | 32 / 8 | 128 | 9,216 | 131,072 |
Nothing matches. The design clearly borrows ideas: the 3:1 sliding-to-full pattern, the 0.25 partial rotary factor and 16/4 heads at head_dim 256 sit close to Qwen3.5-4B's recipe; the window of 512, GELU MLP at 10,240, and theta 10,000 on sliding layers sit close to Gemma 4's. Borrowed recipes are how the field works. Borrowed weights would be a different claim.
The tokenizer settles part of it. It is a byte-level BPE with 131,072 entries, and its ids do
not line up with any candidate: of the tokens it shares with DeepSeek-V3, Mistral's Tekken and Qwen3,
only 8, 5 and 68 sit at the same id (measured), which is byte-table coincidence. Its
pre-tokenizer is DeepSeek-V3's (the same digit and CJK splits, a near-identical main regex) plus an
extra per-digit split, and its special tokens use DeepSeek's <|end▁of▁sentence|> spelling. A
DeepSeek-style tokenizer recipe, trained into its own vocabulary. An embedding table cannot be
copied across vocabularies, so at minimum the embedding is new.
For the transformer blocks I went one step further. Four of the candidates have MLP rows of the
same width (2,560), so I range-requested layer 0's gate_proj from Spark-X2.5-4B and from each,
about 50 MB per model, and for every Spark row found the best-matching row anywhere in the other
matrix by absolute cosine similarity. As a positive control I ran the same test against
Spark-X2.5-4B-Base, which the card names as its base.
| Spark-X2.5-4B layer 0 vs | median best-match cosine | column-norm correlation |
|---|---|---|
| Spark-X2.5-4B-Base (control) | 0.987 | 0.996 |
| Qwen3-4B | 0.078 | 0.006 |
| Qwen3.5-4B | 0.078 | 0.000 |
| Gemma-4-E4B | 0.079 | 0.037 |
| Gemma-3-4B | 0.079 | -0.016 |
All measured. 0.078-0.079 is the noise floor: Gemma-4-E4B against Gemma-3-4B, two models nobody claims are related, also scores 0.079. The control scores 0.987, so the test can see a fine-tune when there is one. The instruct model is a fine-tune of its own base, and neither is a fine-tune of these four.
What this is not: proof against every possible base, or against a deliberate disguise. Someone determined could rotate the hidden basis and defeat a row match. But the vocabulary, the layer inventory and the weights all point the same way (reasoned): this is a model trained by its publisher, not a rename. The press coverage's ~20 trillion pretraining tokens is reported.
Read the code you would be trusting
auto_map points transformers at two files: configuration_spark.py (4,329 bytes) and
modeling_spark.py (19,418 bytes). I read both in full (measured):
grep -nE "import|exec|eval|subprocess|socket|requests|urllib|http|base64|os\.|open\(" \
modeling_spark.py configuration_spark.pyThe only imports are math, torch and transformers; the only http is the Apache licence URL
in a comment. No network calls, no file access, no exec, eval, pickle or base64 blobs. It is a
plain PyTorch decoder: RMSNorm, a fused QKV projection, partial RoPE computed per layer type, eager
softmax attention with sliding-window masks from transformers' own helpers, the per-head sigmoid
output gate, a GELU-gated MLP, and a residual stream kept in float32. The one surprise is cosmetic:
the MLP raises a Chinese-language error if hidden_act is anything but gelu.
Better still, you do not need it. llama.cpp's src/llama-arch.cpp on master registers
LLM_ARCH_SPARK2_5 as spark2_5 (measured, read from the source), and the card says Ollama
0.34.1 and LM Studio runtime 2.34.0 added it natively (reported). The GGUF path runs C++ from
those projects, not Python from the model repo. The post's "works with Ollama, LM Studio, llama.cpp"
holds. The transformers path, and the vLLM command in the card, do need --trust-remote-code;
having read the code, I would be comfortable with that here, but read it again if the repository
changes.
Does "native 1M context" hold?
Two separate questions hide in that phrase: can the architecture address a million positions, and was it trained on them.
The first is checkable from the config. There is no rope_scaling block, so the 1,048,576 is not a
YaRN stretch applied at inference. Sliding layers only ever see 512 tokens back, so their RoPE never
has to reach far. The nine full-attention layers rotate only 64 of their 256 dimensions; the other
192 carry no position at all. On the rotated 64, theta 5,000,000 puts the slowest pair's wavelength
at about 19.4 million positions (reasoned: ), comfortably past
one million, so positions out to the claimed length remain distinguishable. The config is
consistent with a native 1M context.
The second is not checkable from files. The card says long context came from "a dedicated training
stage comprising hundreds of billions of tokens, with sequence lengths extending to 1M tokens"
(reported). One small inconsistency: tokenizer_config.json sets model_max_length to 131,072,
which will make some tooling warn or truncate at 128K (measured). And the card's one long-context
reasoning benchmark, AA-LCR, has Spark-X2.5-4B at 56.3 against Qwen3.5-4B's 57.0 (reported), so
on the long-context number the vendor chose to publish, it is level with the model it beats nearly
everywhere else.
Does it run on a 16 GB laptop?
The weights do, easily. XHToken's own GGUF files are 2.60 GB at Q4_K_M, 4.38 GB at Q8_0 and 8.23 GB at BF16 (measured, from the repository listing). A 4B model at 4 bits fits on a phone.
The million tokens do not. In a sliding-window layer the KV cache is capped at 512 positions; in a full-attention layer it grows with the context. Per token, per full layer, the cache holds a key and a value for 4 KV heads of 256 dimensions:
(9 full layers; K and V; 4 heads; 256 dims; 2 bytes in bf16): 36,864 bytes per token. At 1,048,576 tokens that is 36 GiB (reasoned). The 27 sliding layers add 54 MiB in total, whatever the context. If all 36 layers were full attention, the cache would be 144 GiB, so the hybrid layout is a real fourfold saving. It is still more than twice a 16 GB laptop's entire memory.
The sliding-window layers stop growing at 512 tokens, so past a few thousand tokens almost all of the cache is the nine full-attention layers. If all 36 layers were full attention, this cache would be 144.0 GiB instead of 36.1 GiB. The hybrid layout is a real 4x saving. It still does not make a million bf16 tokens fit in a laptop: only a 4-bit cache gets near, and a lossy cache is a quality trade the card does not measure.
With 4 GiB set aside for the operating system and Q4_K_M weights, about 9.6 GiB remains for the
cache: roughly 279,000 tokens in bf16, 525,000 with llama.cpp's 8-bit q8_0 cache, and 992,000 with
the 4-bit q4_0 cache (reasoned, ignoring compute buffers, so these are upper bounds). A
million tokens on 16 GB is just about reachable with a lossy 4-bit cache. Nothing on the card
measures what that does to quality. "Runs on a 16 GB laptop" and "1M context" are both true. They are
not true at the same time.
The benchmarks
The card's table has 21 rows and eight columns: the two Spark models against Qwen3.5 at 9B, 4B and 2B, Gemma 4 at 12B, E4B and E2B. Spark-X2.5-4B beats Qwen3.5-4B on 19 of 21 and has the best score in the whole table on 12 (measured, tallied from the card; the scores themselves are reported). On agent benchmarks the margins are large: 30.4 on τ³-bench against 6.7 for Qwen3.5-4B, 40.9 on BrowseComp against 14.3.

Four things to weigh before repeating any of it:
- Who ran the rivals. The footnote says an asterisk marks "reported results from publicly‑released model cards / papers". Of 114 rival cells, 36 have one (measured). The other 78 were, by the card's own key, not copied from a published card, which in practice means the vendor ran them, under its harness and settings. That is normal, and it is also where most benchmark disagreement comes from.
- The chart is a selection. It shows eight of the 21 rows, and in each of the eight Spark-X2.5-4B leads its size class. The rows where it trails Qwen3.5-4B (τ²-bench 75.1 vs 79.9, AA-LCR 56.3 vs 57.0) are in the table but not the picture, and so is the whole Gemma4-12B column, which beats it on SciCode (39.8 vs 34.7) and IFEval (94.8 vs 93.0). The chart also prints 12.2 for HLE where the table says 12.3 (measured), a small slip, but a sign the two were made separately.
- One pattern breaks. SWE-Bench Pro is the harder sibling of SWE-Bench Verified, and every other model in the table with both scores is lower on Pro: Qwen3.5-4B 29.4 vs 38.8, Gemma4-12B 21.9 vs 44.2. Spark-X2.5-4B is the only one that scores higher, 44.4 vs 41.6 (measured from the table). It may be a harness difference, a scaffold tuned for Pro, or a real strength. It is the row I would reproduce first.
- There is nothing to rerun. No evaluation code, configs, prompts or logs ship with the release, and I found no independent evaluation in the five weeks since launch.
The post-training recipe is at least described. Supervised fine-tuning, then reinforcement learning per domain to produce general, reasoning, agent and code teachers, then what the card calls MOPD, multi-teacher on-policy distillation into one student with a reverse-KL objective. On-policy distillation is covered in Inkling-Small; nothing in Spark's files contradicts the description, and nothing in them can confirm it either.

"#1 on Hugging Face"
On 2026-10-06 the Hugging Face API gives Spark-X2.5-4B a trendingScore of 33, which places it
#149 on the trending list; the top model scores 1,427 (measured). The repo has 1,380 likes,
37,787 downloads in the last 30 days and 43,329 all-time; the official GGUF repo has 478,743
downloads (measured). That is a popular small model. It is not, today, the top of anything.
The repository was created on 2026-08-24 and launched publicly on 2026-09-01. The post is from 2026-10-05, five weeks later, and says it "just hit #1". It may well have topped the trending list in its launch week; the API does not keep history, and I could not confirm it either way (unverified). One reply under the post, in Chinese, asks (my translation) whether the author actually downloaded and tested it, and says that if not, this is an ad. The post does not say.
The checklist
Here is the whole procedure as a list, with what Spark-X2.5-4B showed at each step. None of it needs a GPU. My downloads came to about 350 MB, nearly all of it the optional weight comparison.
- fetchGET /api/organizations/<org>/overview, then a search in the publisher's own languagelook forA known lab, a history of releases, press that names a company rather than a handle.foundXHToken is 词元星火, which Chinese press describes as a wholly owned iFlytek subsidiary. 词元 is the Chinese word for a token in the NLP sense. No coin is mentioned anywhere in the model's files, org page or launch coverage. The org has no verified badge.
- Publisher (1 min). Org API, then search in the publisher's own language. Here: an iFlytek subsidiary; the "token" is the NLP kind. Pass.
- config.json (2 min). Architecture class, layers, context, scaling blocks, and fields that
disagree with each other. Here: a custom
spark2_5class matching the card, plus one straymodel_max_length. Pass. - Safetensors headers (2 min). Range-request the 8-byte length and the JSON. Count parameters. Here: 4,112,079,360, exact. Pass.
- Base match (4 min). Compare shapes and vocab with likely bases; if any match, compare one weight tensor, with the declared base as a positive control. Here: no match on any axis, noise floor on weights, 0.987 against its own base. Pass.
- Remote code (3 min). Read every file
auto_mapnames; grep for network, filesystem, subprocess and dynamic execution. Check whether llama.cpp supports the architecture natively, so you can avoid running it at all. Here: clean, and avoidable. Pass. - Context claim (1 min). RoPE settings, scaling, which layers are global, and the KV cache the claim implies. Here: consistent, 36 GiB at full length. True with a cost.
- Benchmarks (1 min). Who ran the rival numbers, whether the chart is a subset, any eval code, any broken pattern. Here: vendor-run, selected, no code, one odd row. Unclear.
- Popularity (1 min). Trending score and rank today, downloads, and the gap between release and post. Here: #149 today. Stale or unverified.
The same procedure has caught real problems elsewhere: a benchmark table that could not add up and a LoRA nobody mentioned in BTL-4, and a "mixture of architectures" that was public checkpoints and a tool loop in Interfaze 1 Lite. Here it mostly found a model that is what it says it is.
What I would tell someone who saw the post
Spark-X2.5-4B is a genuine, original small model from a major Chinese speech-and-language company, with Apache-2.0 weights, a readable and harmless loader, an efficient hybrid-attention design and native support in the local runtimes the post names. That is more than many viral models can say.
The post overstates three things. The 1M context is real in the architecture and reported in the training, but it needs about 36 GiB of cache in bf16, so the laptop and the million tokens are two different deployments. The coding and agent strength is the vendor's measurement, with a selected chart and no harness to rerun. And "#1 on Hugging Face" was, at best, true a month earlier.
If the agent numbers are even half right, a 4B model at 30.4 on τ³-bench is worth an afternoon. Run the GGUF, at the context you actually need, on the tasks you actually have, and measure. For how the sliding-window and global layers split the cache, see the attention and KV cache explainer and RoPE.