~/satyajit

pplx-embed-v2-late: the shared space is between two model sizes, not two modalities

mdjsonmcp

2026-10-08 · 24 min · retrieval · multimodal · distillation · small-models · benchmarks

Why read this

Hightop 30%

Recounts both checkpoints from their headers, derives vectors and bytes per page from the configs, and shows cross-model means index big, query small.

  • Runs on a consumer GPU
  • Original analysis
  • Widely used

Vision & multimodalMITPractitioner model

How this was scored
Is it new?
2 of 3: A real new idea, method or capability
Can I trust it?
2 of 3: Measures key facts from files, code or configs
Can I run it?
3 of 3: Open, permissive, runs on reader hardware with instructions
Will I understand it?
2 of 3: Mechanism from first principles with figures
Can I act on it?
2 of 3: A concrete recipe, numbers or comparison
Will it last?
2 of 3: A reference for a year or more
Does it affect many?
2 of 3: A widely used model, tool or lab release
Only here?
2 of 3: A teardown or measurement few others did

Score 73 of 100, ranked 61 of 476 rated articles. Each question is answered 0–3 by hand, and a 3 is rare. How articles are scored

Perplexity's announcement said the two new models "retrieve text, images, and pages with a shared embedding space for cross-model querying." I read that as cross-modal the first time, and so did the note that sent me to it. Every multimodal retriever has a shared space between text and images; that is what makes a text query find a picture. It would not be news.

The word is model. The 0.6B and the 9B produce vectors that live in the same space, so a corpus indexed with the big one can be searched with queries from the small one. In a late-interaction retriever, where the index is huge and the query path is latency-bound, that is a more useful property than it sounds, and it is the part of this release I would actually build around.

So I went through the files: both model cards, both safetensors headers (two HTTP range requests each, no weights downloaded), every config in the repos, the Sentence Transformers 6 code that loads them, the transformers code for the Qwen3.5 backbone and its image processor, and the launch post with its ten charts. I did not run the models.

Weightspplx-embed-v2-late-0.6b, pplx-embed-v2-late-9b, float32 safetensors
LicenceMIT, both sizes
Outputone 128-dimensional, unit-length vector per token or image patch
ScoringMaxSim, similarity_fn_name: "maxsim"
Inputstext, or images (photos, screenshots, rendered pages); not both in one input
Runtimesentence-transformers >= 6.0.0, transformers >= 5.4.0, no custom code
Papernone yet; the post promises "a technical report later this year"
perplexity-ai/pplx-embed-v2-late-9b@77e936a · snapshot 2026-10-08
announced
9B
measured
8,392,695,024
parameters
8.39B
repo size
33.59 GB
architecture
Qwen3_5Model
task
feature-extraction
library
sentence-transformers
license
mit
safetensors
2 shards
largest file
33.57 GB
files
17
downloads
240
likes
15
languages
multilingual
parameters by dtype
F328.39B
feature-extractionsentence-similaritymtebsentence-transformers

The header holds 8,392,695,024 parameters: 6.92B text layers, 456M vision, 1.02B token embeddings, and no output head.

repo last modified 2026-10-05

perplexity-ai/pplx-embed-v2-late-0.6b@dd4e95b · snapshot 2026-10-08
announced
0.6B
measured
594,321,600
parameters
594.3M
repo size
2.40 GB
architecture
Qwen3_5Model
task
feature-extraction
library
sentence-transformers
license
mit
safetensors
2 shards
largest file
2.38 GB
files
17
downloads
251
likes
21
languages
multilingual
parameters by dtype
F32594.3M
feature-extractionsentence-similaritymtebsentence-transformers

The header holds 594,321,600 parameters: 239M text layers (12 of Qwen3.5-0.8B's 24), 101M vision, 254M token embeddings.

repo last modified 2026-10-05

One vector is a lossy summary

A dense embedding model turns a document into one vector and a query into one vector, and the score is their dot product. This is why it scales: a billion documents is a billion points in an approximate-nearest-neighbour index, and a search is one probe. The price is that a 4,000-token contract and a slide full of numbers each get squeezed into the same few thousand floats, before anyone knows what the query will ask. Perplexity's post points at the LIMIT benchmark, which makes the stronger version of that argument: for a fixed dimension and a large enough corpus, some patterns of relevance cannot be produced by any single-vector model, however it is trained.

The opposite end is a cross-encoder, which reads query and document together and can match anything to anything, but needs a forward pass for every candidate on every search. Nobody runs that over a corpus. It reranks a shortlist.

Late interaction, the idea from ColBERT, sits between them. Documents are still encoded alone and offline, like a dense model. But the encoder keeps one vector per token instead of pooling them, and the score is computed from all of them at query time:

s(q,d)=∑i=1∣q∣max⁡1≤j≤∣d∣  qi⊤djs(q, d) = \sum_{i=1}^{|q|} \max_{1 \le j \le |d|} \; \mathbf{q}_i^\top \mathbf{d}_j

where qi\mathbf{q}_i is the vector for query token ii and dj\mathbf{d}_j the vector for document token (or image patch) jj. Every query token looks across the whole document for its single best partner, and the partners' similarities are added up. The vectors are unit length (the last module in the pipeline is Normalize), so each term is a cosine between -1 and 1, and the score grows with the number of query tokens.

Perplexity's own figure is the clearest picture of it I have seen, so here it is, and below it the same numbers in a grid you can poke.

A 3 by 5 grid of similarities. Query tokens quick, pasta and recipe down the side; document tokens easy, homemade, noodle, dish and minutes across the top. The highlighted maxima are quick to easy at 0.78, pasta to noodle at 0.88 and recipe to dish at 0.70. Arrows carry those three values to a box reading MaxSim score 2.36.
MaxSim on a toy pair: each query token keeps its best-matching document token and the maxima add up to 2.36. The similarities are illustrative, drawn by Perplexity to explain the scoring (launch post, late-interaction figure).
MaxSim, one query against one documentsimilarities from Perplexity's figure; click tokens to drop them
query \ docbest
0.780.310.120.180.520.78
0.150.340.880.610.090.88
0.290.570.430.700.220.70
MaxSim score (sum)
2.36
the figure's 2.36
divided by query tokens
0.787
MeanMaxSim, a length-free scale
dot products for this pair
15
3 x 5, each over 128 numbers; a single vector needs 1

Drop "noodle" and "pasta" falls back to "dish" at 0.61, so the score loses 0.27, not 0.88. Each query token only ever needs one good partner, which is why removing redundant document vectors costs less than the vector count suggests, and why removing the one vector that carries a rare fact costs a lot.

Two things fall out of the sum that matter later. A query token only needs one good partner, so a document can lose many redundant vectors and barely notice: drop "noodle" and "pasta" falls back to "dish" at 0.61, costing 0.27 rather than 0.88. And a document's cost is its vector count. Scoring one candidate is ∣q∣×∣d∣|q| \times |d| dot products of 128 numbers, where a dense model does one.

What is in the two checkpoints

The model card says "built on Qwen3.5 with bidirectional attention." The configs say which Qwen3.5. The 9B's text tower is Qwen3.5-9B's exactly (hidden size 4,096, 32 layers, 12,288-wide MLP) with its 27-block vision tower. The 0.6B's widths are Qwen3.5-0.8B's (hidden 1,024, MLP 3,584, a 12-block vision tower), but num_hidden_layers is 12 where the base has 24. The post confirms it: they "prune its 24-layer text tower to 12 layers." It doesn't say which twelve survived.

Summing every tensor in the headers:

Group0.6B9B
Text layers239,449,0246,919,565,824
Vision tower100,592,896456,010,480
Token-embedding table254,279,6801,017,118,720
Total in the file594,321,6008,392,695,024
Card's "active"340M7.4B
128-wide projection (1_Dense)131,072524,288

The card's "active parameters" is everything except the embedding table: 594.3M minus 254.3M is 340M, and 8.39B minus 1.02B is 7.38B. It's a fair way to count an encoder, since a table lookup costs nothing at inference, and the post says so plainly for the small model. For text-only input the 0.6B's vision tower sits idle too, so a text query runs through about 240M parameters.

The 9B has no output head at all, which is why its file holds 8.39B rather than Qwen3.5-9B's nine-and-a-bit: an embedding model never predicts a next token, so lm_head was dropped. The 9B config still carries mtp_num_hidden_layers: 1 from the base, but no multi-token-prediction weights are in the header. Everything is stored in float32, which makes the 9B's single model.safetensors 33.57 GB on disk. Cast it to bfloat16 when you load it; nothing in the configs asks for float32 at runtime.

After the backbone, modules.json lists the rest of the pipeline, and it is short:

0  Transformer         Qwen3.5, last_hidden_state for every position
1  Dense               4096 -> 128 (1024 -> 128 on the 0.6B), no bias, no activation
2  MultiVectorMask     drop bare punctuation tokens, documents only
3  Normalize           unit length, so MaxSim terms are cosines

Queries are prefixed with [Q] and documents with [D] , both added to the tokenizer as special tokens; the card warns that PyLate puts these markers second where this model expects them first. Queries are cut at 1,024 tokens and documents at 4,096. query_expansion is null, so there is no ColBERT-style padding of short queries with [MASK] tokens; a five-word query is scored with roughly five vectors plus the marker. The punctuation skiplist is the classic ColBERT one (all 32 ASCII punctuation characters) and applies only on the document side.

"Bidirectional" covers a quarter of the layers

This is the detail I didn't expect. Qwen3.5 is a hybrid: three of every four layers are Gated DeltaNet linear-attention layers, and every fourth is ordinary softmax attention. The configs keep that pattern (full_attention_interval: 4), so the 9B has 8 softmax layers out of 32 and the 0.6B has 3 out of 12.

The config sets is_causal: false, and in transformers that flag does what it says for the softmax layers: create_causal_mask returns a bidirectional mask, and the flag is also passed to the attention kernel. But it does not reach the DeltaNet layers. Qwen3_5GatedDeltaNet runs a causal 1-D convolution and then a left-to-right recurrent scan, and nothing in it reads is_causal. In those layers each token still sees only what came before it.

So "bidirectional attention" is true of a quarter of the depth. Information from the end of a passage reaches its first tokens only through the softmax layers, and since the last layer in both models is a softmax layer, every output vector does see the whole input at least once. Whether that costs anything I can't tell without running it; ColBERT-style vectors are mostly about the local token in context, and the benchmark numbers below say the hybrid works. It's a design trade the card's one line hides. If the gated delta rule is new to you, the article on it explains why it is a recurrence and not an attention matrix.

How a page becomes 1,692 vectors

The headline modality is pages: PDFs, slides and scans rendered to images and embedded directly, with no OCR in between. How many vectors an image produces is not on the card, but it is fully determined by processor_config.json.

The image processor is Qwen2-VL's. Its vision tower cuts the image into 16-pixel patches and then merges each 2×2 block of patches into one token, so each output vector covers a 32×32-pixel block. Before that, smart_resize rounds the image to multiples of 32 and, if it has more than max_pixels (1,800,964, which is 1,342 squared), scales it down to fit. The cap works out to at most 1,758 vectors per image. (The same file also carries a size.longest_edge of 16,777,216; in the current transformers code the legacy max_pixels key overrides it at load time, so the 1.8-megapixel cap is the one that applies.)

What that means for the things you would actually feed it:

InputResized toVectors
Letter page at 72 dpi (612×792)608×800475
Letter page at 100 dpi864×1,088918
Letter page at 150 dpi1,152×1,5041,692
A4 page at 150 dpi1,120×1,5681,715
1080p screenshot1,760×9921,705
12 MP phone photo1,536×1,1521,728

Add three or four marker tokens ([D] , the vision start and end tokens) to each. The mask module in this export does not strip them: keep_only_token_ids is null, so nothing restricts the index to image-patch vectors the way colpali-engine's mask_non_image_embeddings does. With so few text tokens around an image it hardly matters here.

Anything rendered at 150 dpi or more hits the cap, so a realistic page is about 1,700 vectors. The render resolution is the main dial on your index size, and it is set by your PDF rasteriser, not by the model. Perplexity doesn't say what resolution ViDoRe's pages were embedded at, so I can't tell you what quality you give up at 72 dpi.

One limitation is stated plainly on the card and is easy to miss: "Mixed text+image inputs are not supported." A document is either text or an image. A page with a useful OCR layer can't be embedded as both at once; you would index it twice and merge.

What the index costs

The launch post gives storage one paragraph and leaves the sums to the reader. Every one of those vectors is stored. encode_document returns float32, so a 150-dpi page is 1,692 × 128 × 4 bytes, 866 KB. In float16 it is 433 KB. A dense model with a 4,096-wide float32 vector stores 16 KB for the same page, so late interaction costs about 26 times as much at half precision. A million pages is 433 GB of vectors before any index structure.

Nothing in the release reduces that. There is no token pooling in modules.json and no quantised variant. Sentence Transformers 6 does ship the tools: a HierarchicalTokenPooling module that clusters a document's vectors and keeps about one in pool_factor, which you can pass per call to encode_document, and ColBERTv2-style engines store each 128-d vector as a 4-byte centroid id plus 2 or 1 bits per dimension (36 or 20 bytes instead of 256). Perplexity reports no results with either, so how much quality survives them on these models is unmeasured.

one vector per token vs one vector per documentvector counts from the shipped configs; no graph or metadata
what one document is
multi-vector storage, per 128-d vector
single-vector baseline: dimensions
single-vector baseline: precision
documents in the corpus
index size, 1M documents
late interaction
433 GB
4096-d single
16.4 GB
vectors per document
1,692
resized to 1152 x 1504, one per 32 x 32 block
bytes per document
433 KB
1,692 x 256 B; half of float32
single-vector document
16.4 KB
4096 x 4 B
late interaction costs
26x
the bytes of the single-vector index
multiply-adds to score one candidate
3.5M
16 x 1,692 x 128, against 4,096 for one dot product
same index from the 0.6B or the 9B
same bytes
both emit 128-d vectors from the same tokenizer and image processor

A Letter page rendered at 150 dpi is past the 1,800,964-pixel cap, so it is shrunk to 1,152 x 1,504 and gives 1,692 vectors: 433 KB in float16, about 26 times a 4,096-wide float32 vector. Rendering at 72 dpi gives 475 vectors instead; how sharp you rasterise a page sets the index bill, and that is a choice the model card leaves to you.

Two things in that calculator are worth a second look. The single-vector baseline's width matters less than you'd think: going from 1,024 to 4,096 dimensions changes the dense index fourfold, while the late-interaction index is ten to a hundred times bigger in every setting. And the last cell is the quiet advantage of the shared space. Both models emit 128-d vectors from the same tokenizer and image processor (I diffed the configs: only the backbone widths and depths differ), so a 9B index costs exactly the same bytes as a 0.6B one. The bigger model costs you indexing compute, once.

The post makes a fair point on the other side. The vision-only multi-vector models it compares against emit wider vectors: 2,048 dimensions for EVIE-4.5B, 2,560 for Nemotron ColEmbed V2 4B and 4,096 for EVIE-8B and Nemotron ColEmbed V2 8B. Per vector that is 16 to 32 times the bytes of a 128-d vector, which is why Perplexity says indexing its 100,000-image internal benchmark at those widths "was impractical" and left those models off that chart. Their patch counts per page depend on their own processors, which I didn't check, so I'd treat the per-vector ratio as the fair comparison rather than a per-page one.

The shared space is between sizes

Now the property in the headline. Perplexity trained an 18B ColBERT teacher contrastively (built, the post says, by the same layer pruning from a 27B backbone), then distilled it separately into the 9B and the 0.6B. The distillation is not the usual "match the teacher's ranking." It is LEAF-style representation distillation, done per token: for every retained token, the student's output vector is pulled toward the teacher's vector for that token. LEAF did this for single-vector models; the extension here is that the target is a whole matrix of token vectors.

If both students reproduce the teacher's vectors, they reproduce each other's, and the two models become interchangeable at the vector level, which is all "cross-model querying" means. Encode the corpus once with the 9B, offline, and encode live queries with the 0.6B, which runs about 240M parameters for text and could sit on a laptop or a phone.

Dot plot of nDCG@10 per domain with three markers each: 0.6B queries on 0.6B docs (hollow), 0.6B queries on 9B docs (light), 9B on 9B (dark). Average 78.0, 79.6, 81.3. Finance 79.9, 81.5, 82.5. Legal 78.2, 79.4, 81.6. Health 78.0, 80.2, 82.5. Conversation 79.7, 80.7, 81.5. Tech 66.6, 68.6, 70.4. Multilingual 85.8, 87.2, 89.5. ViDoRe v3 (8 tasks) 62.3, 63.5, 65.2.
Indexing with the 9B and querying with the 0.6B lands between the two symmetric setups in every domain: 79.6 against 78.0 and 81.3 on the six-domain average, 63.5 against 62.3 and 65.2 on ViDoRe v3 images (launch post, cross-model retrieval figure).

The arithmetic on that chart is better than "a 9B index helps." On the text average the asymmetric setup gains 1.6 points of the 3.3 between the two models, so about half, which is what the post claims. On ViDoRe v3 images it gains 1.2 of 2.9, about 41%. Query cost is identical to running the 0.6B alone, because it is the 0.6B encoding the query, and the index is the same size.

Two cautions. Only one direction is reported: 0.6B queries against a 9B corpus. The post also suggests encoding private documents locally with the 0.6B and merging them with results from a cloud 9B index, which is the reverse pairing mixed into one ranking, and there is no number for it. Scores from two different encoders sharing a space are not guaranteed to share a scale. And "shared space" is a property of how well each student matched the teacher; a fine-tune of one size on your own data will break it unless you fine-tune the other to match.

Does it hold up

The evaluation is broad: 72 text tasks, Perplexity's own web benchmark, two visual-document benchmarks, a natural-image one and two agentic ones. I checked every claim in the post's prose against the numbers on its own charts. They agree, with one wording issue I come to at the end.

The training mix is worth knowing first, because it frames the text results. Both models saw 186 million query-document pairs from 594 datasets in 46 languages: 88.3% text-to-text, 8.3% text-to-image and 3.4% text-to-page, rebalanced by sampling weights to 56.5%, 30.9% and 12.6%. Perplexity says it excluded every dataset tied to a benchmark it reports. I have no way to check that, but it is the right policy and they paid for it in lost training data.

Text

On 72 MTEB-style tasks across finance, legal, health, conversation, tech and a multilingual group, the 9B averages 81.3, 1.6 ahead of nemotron-embed-8b at 79.7, and the 0.6B averages 78.0, 0.3 behind gemini-embedding-2. The 9B is not top everywhere: in health, nemotron-embed-8b scores 83.9 to its 82.5.

Table of nDCG@10 on domain-specific benchmarks. Columns: pplx-embed-v2-late-9b, nemotron-embed-8b, gemini-embedding-2, voyage-large-4, qwen3-embed-8b, then under 1B pplx-embed-v2-late-0.6b, voyage-4-nano, qwen3-embed-0.6b. Average row: 81.3, 79.7, 78.3, 79.0, 74.8, 78.0, 74.6, 68.4. Health row: 82.5, 83.9 (bold), 82.4, 83.0, 80.2, 78.0, 76.0, 72.4.
Domain-specific text retrieval, nDCG@10; the average weights six domains equally and the parentheses count tasks (launch post, domain-specific benchmarks figure).

The more striking text result is ViDoRe v3 with the pages converted to Markdown by OCR, a pure text task on visually rich documents. The 9B scores 64.7 and the 0.6B 61.2; the best model from anyone else is nemotron-embed-8b at 60.6, and the best other sub-1B model, topk-embed-v1-0.8b, is 4.3 points behind the 0.6B.

Table of nDCG@10 on the public ViDoRe v3 subset after OCR to Markdown, with embedding width and vector type per model. pplx 9B 128 multi 64.7; nemotron-embed-8b 4096 single 60.6; topk-embed-v1-2b 128 multi 59.2; gemini-embedding-2 3072 single 58.9; voyage-large-4 1024 single 58.9; qwen3-embed-8b 4096 single 51.5; pplx 0.6B 128 multi 61.2; topk-embed-v1-0.8b 128 multi 56.9; voyage-4-nano 2048 single 53.3; qwen3-embed-0.6b 1024 single 44.2.
ViDoRe v3 as a text task over OCR-derived Markdown; the second row of the header gives each model's vector width and whether it is single- or multi-vector (launch post, ViDoRe v3 Markdown figure).

Q2D-Web is Perplexity's own benchmark: about 70,000 agent-reformulated queries from production traffic over 190 million web documents, scored by Recall@1000. The 9B and 0.6B reach 74.8 and 73.6 on the combined judgments against 69.3 for nemotron-embed-8b. Read it with the caveat the post itself gives: two of the three judgment sets come from what Perplexity's own systems already cited or ranked, which favours a model trained to rank like them, so they lead with the combined set, which adds LLM judgments of documents neither system surfaced. It is still their benchmark, their queries and their judge.

Three bar panels of Recall@1000 on Q2D-Web: Citation, Web and Combined. pplx 9B 88.2, 68.8, 74.8; nemotron-embed-8b 83.1, 61.7, 69.3; qwen3-embed-8b 79.3, 58.3, 65.3; topk-embed-v1-2b 79.2, 57.7, 65.3; pplx 0.6B 86.7, 67.3, 73.6; voyage-4-nano 69.5, 50.1, 57.4; qwen3-embed-0.6b 71.2, 51.8, 58.7.
Q2D-Web, Perplexity's own web-scale benchmark, Recall@1000 under three judgment sets (launch post, Q2D-Web figure).

Pages and images

On the ViDoRe v3 public subset with real page images, the 9B averages 65.2 and the 0.6B 62.3. Both beat every other vision-language model on the chart, including topk-embed-v1-2b at 61.9 and the single-vector qwen3-vl-embed-8b at 58.3 and gemini-embedding-2 at 46.0. The vision-only multi-vector models are separated out, and two of them beat the 9B: EVIE-8B at 66.4 and EVIE-4.5B at 65.9. Nemotron ColEmbed V2 8B scores 63.5, 1.2 above the 0.6B.

Table of nDCG@10 on ViDoRe v3 images, public subset, eight domains. Vision plus language: pplx 0.6B 62.3, topk-embed-v1-0.8b 59.0, pplx 9B 65.2, topk-embed-v1-2b 61.9, qwen3-vl-embed-8b 58.3, gemini-embedding-2 46.0. Vision-only: nemotron-colembed-v2-8b 63.5, nemotron-colembed-v2-4b 61.4, evie-8b 66.4, evie-4.5b 65.9. Embedding widths 128 for the four pplx and topk models, 4096 and 3072 for the single-vector ones, 4096, 2560, 4096 and 2048 for the vision-only ones.
ViDoRe v3 page-image retrieval, nDCG@10, with each model's vector width; the vision-only multi-vector models are grouped on the right (launch post, ViDoRe v3 images figure).

The post's "our 0.6B model matches models with five times as many active parameters" comes from this chart plotted against size. On the vision-language panel the 0.6B at 340M active sits at 62.3 and the next frontier point is topk-embed-v1-2b, at roughly 1.7B on the log axis, with 61.9. The footer adds that nine models without a published active count are left off the plot.

Two scatter plots of ViDoRe v3 image nDCG@10 against active parameters on a log axis from 100M to 10B. Left, vision plus language: a frontier from colmodernvbert near 26 to pplx-embed-v2-late-0.6b at 62.3, topk-embed-v1-2b near 63 and pplx-embed-v2-late-9b near 65. Right, including vision-only: the frontier continues to EVIE-4.5B and EVIE-8B near 66, with the 9B just below.
ViDoRe v3 image retrieval against active parameters; lines join each panel's Pareto frontier. Scores are from the MTEB leaderboard on 28 September 2026 and topk-embed-v1's model cards (launch post, Pareto figure).

Natural images are the weaker side. On MIRACL-Vision (Wikipedia screenshots) the 9B scores 75.7, behind gemini-embedding-2 at 77.5, and the 0.6B's 67.5 is a point under qwen3-vl-embed-8b. On PPLX-Q2I, an internal benchmark from Perplexity's image-search logs (10,000 queries, 100,000 images), the 9B's 65.0 trails gemini's 67.0 and both sizes beat qwen3-vl-embed-8b by more than 9 points. The second benchmark is internal and unreleased, so treat it as a claim.

Two bar panels of nDCG@10. MIRACL-Vision: pplx 9B 75.7, pplx 0.6B 67.5, qwen3-vl-embed-8b 68.5, gemini-embedding-2 77.5, nemotron-colembed-v2-8b 68.6, nemotron-colembed-v2-4b 62.7. pplx-q2i: 65.0, 60.4, 50.6, 67.0, and the two Nemotron models marked not evaluated due to indexing cost.
MIRACL-Vision and Perplexity's internal PPLX-Q2I image benchmark; the Nemotron models were not run on the second because of index size (launch post, additional image benchmarks figure).

Agents

The agentic results are the ones I find most interesting, because they measure the thing late interaction is supposed to help with: an agent issuing many short, specific queries. On BrowseComp+, with GPT-OSS-120B doing the searching, the 9B gets 64.0% of questions right, 4.9 points ahead of the next ColBERT model (Reason-ModernColBERT) and 8.7 ahead of the best dense one, and the two pplx models need the fewest searches on the chart. A better retriever ending an agent's loop sooner is a real cost saving that a recall table doesn't show.

Two scatter plots against average number of searches from 16 to 21. Answer accuracy: pplx 9B 64.0 at about 16.7 searches, pplx 0.6B about 63 at 16.5, Reason-ModernColBERT about 59 at 18.9, voyage-4-large about 55 at 18.3, gemini-embedding-2 about 52 at 19.3, voyage-4-nano about 47 at 19.4, qwen3-embedding-8b about 44 at 20.5. Retrieval recall shows the same order with the pplx models near 76 and 73.
BrowseComp+ with a GPT-OSS-120B agent at high effort: answer accuracy and retrieval recall against searches per question; up and to the left is better (launch post, BrowseComp+ figure).

MADQA is 500 human-written questions over 800 PDFs (more than 18,000 pages), answered by a Gemini 3.5 Flash agent using the retriever under test. The 9B reaches 92.4% and the 0.6B 90.1%, against 88.9% with Mixedbread's retriever and 93.4% for Mixedbread's agentic search, which runs its own planning sub-agent. Perplexity calls the 9B "a new state of the art among retrievers." On the point estimate it is. But the chart prints the error bars: 92.4 ± 2.6 against 88.9 ± 1.7, so the ranges overlap, and on page F1, the measure of whether the agent found the right evidence pages, Mixedbread's retriever is slightly ahead, 83.4 to 83.0. I'd call that a tie on retrieval quality with a lead in answers that 500 questions can't confirm.

Bar chart of MADQA accuracy and page F1. Human plus oracle retriever 99.4 ± 0.4, F1 n/a. Gemini 3.5 Flash plus Mixedbread Agentic Search 93.4 ± 1.3, F1 84.3. Plus pplx 9B 92.4 ± 2.6, F1 83.0. Plus pplx 0.6B 90.1 ± 2.9, F1 82.5. Plus Mixedbread 88.9 ± 1.7, F1 83.4. Human plus BM25 82.2 ± 2.0, F1 79.3.
MADQA answer accuracy with confidence intervals, and page-level F1 against the annotated evidence pages (launch post, MADQA figure).

What I'd do with it

If you are indexing PDFs, slide decks or scans and today you run OCR into a text embedder, this is the most practical open option I know of for skipping the OCR step. It is MIT-licensed, it loads in stock Sentence Transformers, and the 0.6B is small enough to try on a laptop:

# pip install 'sentence-transformers>=6.0.0' 'transformers>=5.4.0'
from PIL import Image
from sentence_transformers import MultiVectorEncoder
 
model = MultiVectorEncoder("perplexity-ai/pplx-embed-v2-late-0.6b",
                           model_kwargs={"torch_dtype": "bfloat16"})
pages = model.encode_document([Image.open("page-001.png").convert("RGB")])
q = model.encode_query(["which clause sets the limitation period?"])
print(pages[0].shape)               # (vectors, 128): about 1,700 for a 150-dpi page
print(model.similarity(q, pages))   # MaxSim

The deployment I'd pick is the one Perplexity recommends: index with the 9B on a GPU box, query with the 0.6B wherever the query arrives. You pay for the 9B once per document and get about half its quality lead for free, with no change in index size.

Then budget the index before you commit. At about 1,700 vectors a page, storage is the real cost of this design, and the release gives you no compressed variant and no numbers on how pooling or 2-bit residuals affect it. Plan to measure that yourself on your own pages, and to render them at the lowest resolution your retrieval quality tolerates. For general web-text search, where a page of prose is a few hundred tokens and you need billions of documents, a single-vector model with an ANN index is still the economical choice; the post itself positions late interaction as a quality-sensitive first stage or a later stage in a larger pipeline. For more on the single-vector side of that trade, see EmbeddingGemma 2, where Qdrant squeezes 768 dimensions to 40 bytes, and TurboVec for what quantised ANN search costs. If you are still deciding whether you need embeddings at all, BM25 is a fair baseline to beat first.

One wording note to close on, the one I mentioned earlier. The post says these are "the first to combine" multi-vector representations, text and image input, and a shared space across sizes. The first two together are not new: the topk-embed-v1 models on Perplexity's own charts are 128-d multi-vector retrievers scored on both text and page images. The claim only holds with the third, and the third is the one that is genuinely new.

How I checked

Cite this article

For attribution, please use the following reference or BibTeX:

Satyajit Ghana, "pplx-embed-v2-late: the shared space is between two model sizes, not two modalities", ai.thesatyajit.com, October 2026.

bibtex
@misc{ghana2026pplxembedv2late,
  author = {Satyajit Ghana},
  title  = {pplx-embed-v2-late: the shared space is between two model sizes, not two modalities},
  url    = {https://ai.thesatyajit.com/articles/pplx-embed-v2-late},
  year   = {2026}
}
share