~/satyajit

Interfaze 1 Lite: the mixture of architectures is ten public checkpoints and a tool loop

mdjsonmcp

2026-10-06 · 17 min · ocr · document-parsing · open-weights · multimodal · tool-calling · calibration · asr

Interfaze announced interfaze-1-lite as "the first general open-weight model for deterministic work": OCR, speech-to-text, extraction and classification, with "confidence scores and metadata like bounding boxes on every run". The second post in the thread names the design: "Mixture of Architecture (MoA)", where "each specialist model is a dedicated architecture built to solely process a single task", and the whole thing runs on one H100.

That is a claim about structure, and structure is checkable. A Hugging Face repo is a list of files with sizes and SHA-256 hashes, every safetensors file starts with a JSON header naming each tensor's shape, and the routing code ships alongside because the model needs trust_remote_code=True. I read all of it without downloading the 44 GB of weights, then read the GitHub server that wraps it. No third-party code was executed.

interfaze-ai/interfaze-1-lite@1bfa94c · snapshot 2026-10-06
repo size
44.27 GB
architecture
InterfazeLiteModel
task
image-text-to-image
library
transformers
license
apache-2.0
safetensors
68 shards
largest file
10.59 GB
files
130
downloads
6
likes
32
languages
en, zh, es, fr, de, it
vlmmultimodalmultilingualmixture-of-architecturesocrdocument-understandingspeech-recognitionspeaker-diarization

A bundle, not one network: eight upstream checkpoints in named folders plus routing code. 27.78B of its parameters are the Qwen3.8-27B-FP8 core, counted from the safetensors headers.

repo last modified 2026-10-05

Weightsinterfaze-ai/interfaze-1-lite, Apache-2.0, commit 1bfa94c (2026-10-05)
ServerInterfazeAI/interfaze-1-lite, commit b7c642c, an OpenAI-compatible service over vLLM
Launch postThe First Open Weight Model for Deterministic Work
Bundle size44,289,111,287 bytes across all files (measured, Hub API)
Hardwareone 80 GB GPU, compute capability 8.9 or newer (reported)

What "mixture of architectures" means here

Start with what it is not. A mixture of experts is one network whose feed-forward layers are split into experts, with a learned router choosing a few per token, trained end to end. Nothing in interfaze-1-lite is trained jointly, and there is no learned router. The routing is a 27B vision-language model deciding which tools to call, the way any tool-calling agent does.

The model card says as much, plainly: "Interfaze 1 Lite is not one network. It is a reasoning core plus specialists, each chosen for the task it is best at, connected by tool calls." The launch blog draws it the same way.

Architecture diagram. Four inputs (text, images, PDFs and Word, audio) feed a yellow 'Reasoning core' box, described as a hybrid-attention decoder with a vision encoder in FP8, which reads the request, picks the specialists, grounds boxes on a 0 to 1000 grid and writes the answer. Arrows labelled 'runs' and 'results' connect it to a 'Specialist models' panel of eight boxes: document reader, line geometry, layout, speech, diarization, segmentation, forecasting and guardrails. The core outputs an Answer; the specialists output Precontext, each specialist's raw result.
The reasoning core calls specialists and writes the answer; each specialist's raw output is returned beside it as precontext (Interfaze launch blog, architecture figure).

What the diagram does not say is where the specialists come from. The card's architecture table describes each by type only ("Vision-language model trained for page reading", "Encoder-decoder speech recognizer"). The configuration file goes further and says it on purpose: the capability-to-component mapping "is supplied by the deployment (COMPONENT_<CAPABILITY>), so this file declares which capabilities exist without naming what provides them."

The names are elsewhere, in three places that agree:

"Inspiration" undersells it. The Hub API returns the LFS SHA-256 of every large file. I pulled the same listing for the ten upstream repos and compared hashes. 76 of the bundle's 79 LFS files are byte-identical to an upstream file (measured): 44.23 GB of the 44.27 GB of large files. The three that differ are the banner image and two tokenizer.json files; the document reader's is one byte longer than Chandra's. Every weight file is someone else's checkpoint, unmodified.

That is not a criticism of the engineering. It is a correction to the framing, and it has a licensing edge worth knowing. The repo is tagged Apache-2.0 and ships one LICENSE file, while Chandra OCR 2's card declares license: openrail and pyannote's diarization pipeline is gated on the Hub (measured, from the two file listings). The configuration docstring itself says bundled mode comes with "the obligation to carry every upstream licence"; read each upstream licence before you build on the bundle. "Each specialist model is a dedicated architecture built to solely process a single task" is true. It is also true that none of them was built by Interfaze.

Folder by folder

The bundle has one folder per component. Counting parameters from the safetensors headers takes two range requests per file: 8 bytes for the header length, then the header itself.

interfaze-1-lite, folder by folder · pick a call to see which checkpoints it loads

ocr() loads 10.8 GB of the 44.3 GB bundle

RoleUpstream checkpointParamsGBHands back
Reasoning coreQwen/Qwen3.8-27B-FP827.78Bm30.89boxes on a 0-1000 grid (from generated text)
Document readerdatalab-to/chandra-ocr-25.30Bm10.61text and reading order; no confidence
Speechopenai/whisper-large-v3-turbo0.81Bm1.62timestamps; no confidence
Segmentationfacebook/sam2.1-hiera-large224Mh0.90outlines; its mask scores are dropped
LayoutPaddlePaddle/PP-DocLayout_plus-L~32Mr0.13typed blocks with a score
Line geometryPaddlePaddle/PP-OCRv5_server_det~22Mr0.09a quad per line
Line geometryPaddlePaddle/en_PP-OCRv5_mobile_rec~1.9Mr0.01a score per line: the confidence
Diarizationpyannote/speaker-diarization-community-1~8Mr0.03speaker turns
Guardrailsmeta-llama/Llama-Guard-3-1B1.50Bh—unsafe probability
Forecastinggoogle/timesfm-2.5-200m-pytorch231Mh—point forecast; quantiles dropped

Three readers run concurrently on every page: the document reader for text, the Paddle pair for a box and a score per line, the layout detector for typed blocks. The brain is never loaded.

m = measured from the safetensors headers · h = the upstream repo’s own Hub metadata · r = reasoned, file bytes / 4 for float32 Paddle and PyTorch files. Every weight file in the bundle has the same SHA-256 as the upstream file it came from.

The numbers, with how each was obtained:

The brain is 70% of the bundle by bytes (reasoned). Its text stack is Qwen3.8-27B's hybrid design, the one Cinference's speed claim turns on: 64 layers, of which layer_types marks 48 as linear attention and 16 as full attention, one full layer in every four (measured from brain/config.json). That is what the card means by "hybrid-attention decoder", and why it warns that without flash-linear-attention, "generation slows to minutes per page".

Two specialists are not in the bundle

config.json sets "bundle_weights": true, and the loader's docstring says what that means: "component weights live in this repo under the prefixes below." The configuration code lists eleven capabilities, guard and forecaster among them. The file listing has no guard/ folder and no forecaster/ folder (measured). moderate() needs Llama Guard 3 1B and forecast() needs TimesFM 2.5; neither ships.

Reading _repo(), a bundled load of either asks snapshot_download for guard/* or forecaster/*, gets an empty pattern match, and hands the path to from_pretrained. I expect that to fail, but I did not run it, so treat it as unverified. The Docker path in the GitHub repo avoids the question: it pulls each component from its upstream repo and requires an HF_TOKEN because, in its own words, "the diarization and guard models are gated". The card's "Self-contained. One repository, one GPU, runs offline" holds for OCR, speech, detection and chat. It does not hold for the two capabilities whose weights are gated or absent.

One more thing is optional. The config keeps a list of "the 25 languages the fast recogniser TDT v3 covers", and components.env.example names nvidia/parakeet-tdt-0.6b-v3 behind ENABLE_FAST_ASR=1. Off by default, so the default speech path is Whisper alone.

The orchestration is the product

If the weights are borrowed, the work is the Python around them: 4,552 lines in the bundle's nine .py files and 9,329 in the GitHub server's package (measured, wc -l). Some of it is careful.

The tool loop. chat() runs agent.run: the brain sees the user's files as ref-0, ref-1, … and four tools, ocr, stt, object_detection and gui_detection. It gets up to six tool-calling turns. If a file is attached and the brain answers without calling anything, it is reminded once ("answered by eye, the text is invented"). The images themselves are shown to the brain only after a tool has read them: "shown an image first, the model transcribes it by eye and invents text." Every tool's full result goes back to the caller as precontext, next to the answer.

Direct calls skip the brain. ocr(), transcribe(), moderate() and forecast() never load the 27B core; components load lazily on first use, so "a caller who only wants OCR never pulls the 28 GB brain." For the deterministic jobs the announcement leads with, the reasoning core is not in the loop at all.

Detection is the brain, segmentation is SAM. detect() prompts the brain to write boxes as text on a 0-1000 grid, parses them, and passes each box to SAM 2.1 for an outline. ground() is the brain alone, tiling large screenshots. RefCOCO's 83.8 (reported) is therefore a Qwen3.8 score with a prompt and a parser around it (reasoned: ground() and detect() boxes come from the brain alone).

Speech is Whisper with its own guards. Audio past 120 s is cut at quiet points into windows of at most 30 s and decoded 16 at a time; the card reports a 95-minute recording in about 90 seconds. Speaker labels come from pyannote, with each word assigned to the turn it overlaps most. The comments are worth reading: greedy decoding alone "looped '5-5-5' for 60 s of a card number", so the code restores Whisper's reference-decoder fallbacks. The site's piece on Whisper's hallucinations covers why those guards exist, and Nemotron 3 Diarization covers the speaker side.

OCR: two readers, one page

OCR is where the design earns its keep, and it is the one place the confidence claim is fully backed. Three jobs run on every page at once: Chandra OCR 2 reads the page, the PP-OCRv5 detector and recognizer find and read every line, and PP-DocLayout labels the blocks.

Three panels. Left, 'Document reader': three lines of a bank statement as text, captioned 'Complete text in reading order, no positions, no confidence'. Middle, 'Line geometry': three dashed boxes with confidences 0.86, 0.99 and 0.92, captioned 'A box and confidence for every line, its own text can be incomplete'. Right, 'Stitched line': each line's text with its box coordinates and confidence, captioned 'Each line takes the reader's words: exact box, complete text, confidence'.
OCR stitching: the document reader supplies the text, the line detector supplies each line's box and confidence, and each stitched line carries all three (Interfaze launch blog, OCR stitching figure).

The division of labour, from the stitch() docstring: "the reader owns the page's text -- reading order, tables, structure -- and the detector owns every box and confidence." Chandra emits HTML with data-bbox attributes, but the code does not trust those boxes: on portrait pages they "came back squeezed toward the left", and matching by position "dropped 62 of a memo's 98 lines". So lines are matched by what they say. For each detector line, PageText.read searches the reader's flattened text with rapidfuzz.partial_ratio_alignment, widens the hit to whole words, and accepts it at a fuzzy score of 88 or more (100 for lines of four characters or fewer). Best matches claim their span first, so three near-misses cannot all take the same "Dependent 1". A line with no match keeps the recognizer's own text.

There are pragmatic guards on top. A page whose mean line confidence falls below 0.75 is re-read at a denser pixel budget, up to about 8.4 MP; images under about 2.1 MP are read at twice their size, after a phone photo of a receipt came back with "EARBUDDS" and "DIL" for OIL. These are numbers measured by the authors on their own documents and written into comments (reported).

Where the confidence comes from

The confidence on each OCR line is the PP-OCRv5 recognizer's rec_score for its own reading of that line crop (measured from line_runtime.py and _detect_text_lines). That is a legitimate, roughly calibrated signal; it is what made classic OCR pipelines gateable. It comes from a recognizer with about 1.9M parameters.

Two details follow from the code, and both matter if you gate on it.

First, the score is for the recognizer's text, and the line shows the reader's text. When the stitcher replaces a line's words, it keeps average_confidence=line.average_confidence. A long line is accepted at a fuzzy-match score of 88 out of 100, so the two readings can disagree on several characters and the line still carries the recognizer's score for its own version. A reply under the launch post made this point; it is correct as written.

Second, in the released code, word scores are not per-word. split_words() divides the line's box among its words by character count and gives each word confidence=confidence, the line's own value. That applies to stitched lines and to the detector's raw lines alike. The launch blog's worked example shows something richer, so try the two views:

The launch blog’s bank statement · gate each row on its OCR line’s confidence
9 auto · 1 to a person
04/01/24balance1125780.310.99
04/15/24contribution100000.98
04/22/24dividend185.670.95
05/03/24withdrawal-50000.86
05/15/24contribution50000.99
05/24/24interest23.140.92
06/03/24dividend132.980.95
06/17/24contribution100000.98
06/28/24capital_gain286.350.98
06/30/24balance1209509.650.99
05/03/240.99
WITHDRAWAL0.99
-0.92
ACH0.99
-0.42
line: 0.86

The blog’s five word scores average 0.862, which rounds to the line’s 0.86. The trailing dash at 0.42 is the word that sends this row to a person.

In the blog's output (reported, from the hosted API), the withdrawal line's words score 0.99, 0.99, 0.92, 0.99 and 0.42. Their mean is 0.862 (reasoned), which rounds to the line's 0.86, and the gate at 0.9 sends that one row of ten to a person. The row-level gate works the same with the open code. The per-word explanation does not: run locally, all five words would read 0.86 (measured from split_words, not run). The hosted service may compute word scores differently; the GitHub server's package has the same split_words and I found no other source of word scores in it.

Confidence elsewhere

"Confidence scores … on every run" is an OCR claim. Elsewhere in the code:

How "deterministic" is meant

The blog's definition is about the task, not the arithmetic: work "where there's one right answer", with outputs read by an if statement rather than a person. Its table puts "Same input, same output" under consistency.

The code does decode greedily. Every generate() call on the brain and the guard passes do_sample=False, and the vLLM path sends "temperature": 0.0 to the document reader. That is despite brain/generation_config.json shipping Qwen's sampling defaults (do_sample: true, temperature 1.0, top-k 20), which the code overrides. The brain also runs with enable_thinking=False.

Greedy is not the same as bitwise reproducible. Floating-point reductions depend on batch composition and kernel choice, so the same input in a different vLLM batch can flip a near-tie token; a reply under the launch post asked exactly this and got no answer. And Whisper's guards are a temperature ladder, (0.0, 0.2, 0.4, 0.6, 0.8, 1.0): a window that looks like a loop or scores below a log-probability of -1.0 is re-decoded with sampling, which varies run to run unless seeded (reasoned). "Deterministic" here means typed outputs, greedy decoding and a schema. It is not a reproducibility guarantee, and nothing in the repo claims one.

The benchmarks

Table of nine benchmarks comparing interfaze-1-lite, interfaze-1 and the best of four other models. MMMU-Pro 73.2% vs 71.1% vs Grok-4.3 68.7%; RefCOCO 83.8% vs 82.1% vs Gemini-3.7-Flash 80.9%; SOB value accuracy 81.5% vs 80.5% vs Gemini-3.7-Flash 80.2%; olmOCR 83.8% vs 85.7% vs Claude-Sonnet-5 83.5%; VoxPopuli WER 3.0% vs 2.4% vs Gemini-3.7-Flash 4.0%; Spider 2.0-Lite 48.9% vs 52.9% vs Claude-Sonnet-5 50.6%; GPQA Diamond 85.9% vs 92.4% vs Gemini-3.7-Flash 91.4%; OCRBench V2 60.9% vs 70.7% vs Gemini-3.7-Flash 63.9%; MMMLU 87.8% vs 90.9% vs Grok-4.3 89.7%.
interfaze-1-lite's nine headline benchmarks against the hosted interfaze-1 and the best of four frontier models. Every number is the publisher's (Interfaze launch thread and blog, benchmark table).

All of these are reported: Interfaze scored Lite itself "with each benchmark's official scorer" and took every other model's score from its own leaderboard. I re-ran none of them. Two things are worth knowing before you quote them.

The comparison set changes between documents. The model card compares against Claude-Sonnet-4.6 and Gemini-3-Flash; the blog compares against Claude-Sonnet-5 and Gemini-3.7-Flash. On olmOCR-Bench, the card's best rival is Grok-4.3 at 81.9, and the blog's is Claude-Sonnet-5 at 83.5. Lite's own numbers are the same in both. The context window also varies: 131k tokens in the card, 128k in the blog, and max_context: 262144 in config.json.

Since the weights are public, the fair baseline is each component alone. On the olmOCR leaderboard, Interfaze lists Chandra OCR 2 by itself at 84.3% and Lite at 83.8% (both reported, same page). Lite's OCR is Chandra OCR 2 plus the Paddle stitcher, and the stitched pipeline scores half a point below its own document reader. By category, Lite gains on tables (88.8 against 87.9) and old scans (51.3 against 49.2), and loses on headers and footers (90.0 against 92.5), long tiny text (91.9 against 93.9) and multi-column pages (80.1 against 81.3). Datalab's own card reports Chandra 2 at 85.8 ± 0.8 in its harness (reported). The stitcher is not there to raise the text score; it is there to attach boxes and scores, and on this benchmark that costs a little.

The same reading applies to the reasoning rows. Lite's GPQA Diamond of 85.9 comes from byte-identical Qwen3.8-27B-FP8 weights; Qwen's card reports 89.2 for Qwen3.8-27B (reported). The gap is plausibly the enable_thinking=False setting I found in the transformers path, but the evaluation harness is in a separate repository I did not read, so that is a guess (reasoned). And VoxPopuli's 3.01% WER is Whisper large-v3 turbo's number with long-form windowing.

Jina-OCR-v1 is the other OCR release I have taken apart at the header level; the difference is that its new piece was a trained speculative-decoding head, and here there are no new weights at all.

What it costs to run

The card's hardware line is one 80 GB GPU with compute capability 8.9 or newer. The loader's comments put numbers on that. Through transformers, the FP8 brain is dequantized to BF16 because the FP8 kernels "computed nonsense" on a receipt question, "for about 54 GB of memory" (reported). 27.78B parameters at 2 bytes is 55.6 GB (reasoned), so that checks out. Add the 10.6 GB document reader and a chat that touches both holds about 66 GB of weights before activations or KV cache (reasoned). It fits; it is not roomy, and the card itself warns that large PDFs "can spike in significant use of CUDA memory". The vLLM path in the GitHub repo keeps the brain in FP8 and is the one to use for throughput.

The API alternative is priced at $0.85 per million input tokens and $1.50 per million output tokens (reported).

What I'd take from it

The useful idea here is old and sound: let a general model plan and write, and let small specialist models do the perception, because the specialists already produce the metadata a pipeline needs. A 1.9M-parameter line recognizer gives you a per-line score that a 27B VLM does not. Interfaze's stitcher, which matches text by content rather than trusting a VLM's boxes, is a good piece of engineering with honest comments about what broke.

What it is not is a new model. "Mixture of Architecture" names a deployment: ten public checkpoints, eight of them in one repo, wired together by a tool loop and a few thousand lines of careful glue. If you want the confidence scores, call ocr() directly, gate on the line score, and remember whose reading that score belongs to.

Sources, read on 2026-10-06: interfaze-ai/interfaze-1-lite (model card, config.json, every component config, modeling_interfaze_lite.py, contracts.py, agent.py, guard.py, and the safetensors headers of brain/, ocr_vlm/ and asr_fallback/ via HTTP range requests); InterfazeAI/interfaze-1-lite (components.env.example, docker-compose.yml, README.md); the Hub file listings of Qwen/Qwen3.8-27B-FP8, datalab-to/chandra-ocr-2, openai/whisper-large-v3-turbo, facebook/sam2.1-hiera-large, PaddlePaddle/PP-OCRv5_server_det, PaddlePaddle/en_PP-OCRv5_mobile_rec, PaddlePaddle/PP-DocLayout_plus-L, pyannote/speaker-diarization-community-1, meta-llama/Llama-Guard-3-1B and google/timesfm-2.5-200m-pytorch; the launch blog, the leaderboards and the olmOCR board. Figures are reproduced for commentary from Interfaze's launch blog and thread; the two interactives are original.

Cite this article

For attribution, please use the following reference or BibTeX:

Satyajit Ghana, "Interfaze 1 Lite: the mixture of architectures is ten public checkpoints and a tool loop", ai.thesatyajit.com, October 2026.

bibtex
@misc{ghana2026interfaze1lite,
  author = {Satyajit Ghana},
  title  = {Interfaze 1 Lite: the mixture of architectures is ten public checkpoints and a tool loop},
  url    = {https://ai.thesatyajit.com/articles/interfaze-1-lite},
  year   = {2026}
}
share