perplexity logo bg
Perplexity AI releases pplx-embed-v2-late models for multimodal search
Two open-weight late-interaction embedding models on Hugging Face can search text, images, PDFs and slides without OCR
Perplexity AI has released two new embedding models that read pages more like people do: words, pictures, tables and all. The models, called pplx-embed-v2-late, come in 0.6B and 9B parameter versions and are now publicly available on Hugging Face.
Announced on October 7, 2026, the pair handles retrieval across text, images and visual documents such as PDFs and slide decks.
What Perplexity actually shipped
Perplexity’s new models use a ColBERT-style late-interaction architecture that keeps a separate 128-dimensional vector for each token. Rather than one summary per document, the system keeps a running set of notes on every piece.
Matching is handled with MaxSim scoring. Each part of a query looks for its best match inside a document, and those best matches drive the final score.
That shared space enables the release’s most practical feature. A corpus indexed with the larger 9B model can be searched using the lightweight 0.6B model, with testing showing this cross-model querying worked efficiently at no additional cost.
Translated for the people paying cloud bills: you can index once with the heavyweight model, then answer live queries with the smaller, faster one.
The numbers behind the release
Perplexity trained both models through distillation from an 18B teacher model.
The training data covered 186 million query-document pairs across 46 languages. Text-to-text pairs made up the dominant share, with text-to-image and visual document pairs filling out the rest of the multilingual multimodal mix.
AI, tech, and the markets they move—in one daily briefing.
Daily. Free. Join 34,000+ readers across crypto, finance, and policy.
On benchmarks, the research findings point to a noted increase on ViDoRe v3, a test focused on visual document retrieval, from 62.3% to 63.5%.
The 9B model achieved an average nDCG@10 score of up to 81.3% on specific retrieval tasks, matching or outperforming larger alternatives on those tasks.
Because the models work directly on visually rich content, there is no need for OCR or text parsing as a preprocessing step. Complex layouts, tables and figures stay intact rather than being mangled into a wall of extracted characters.
How this fits Perplexity’s 2026 roadmap
This is not Perplexity’s first embedding release this year. Earlier in 2026, the company introduced the dense pplx-embed-v1 model, followed by contextual v2 models.
The late-interaction pair extends that lineup into multimodal territory, with Perplexity steadily building out open-weight, production-oriented embeddings aimed at web-scale applications.
What this means for developers and the AI search market
The most immediate beneficiaries are teams sitting on piles of messy documents ā legal archives, research libraries, internal slide decks and scanned reports ā where meaning often lives in charts and tables rather than paragraphs.
The cross-model querying feature also changes the economics conversation. Late-interaction models have historically carried a reputation for being storage and compute hungry, since they keep many vectors per document instead of one. Letting a 0.6B model query a 9B-indexed corpus gives operators a way to keep query-time costs in check while retaining a higher-quality index.
Open weights mean developers can inspect, self-host and fine-tune without negotiating a contract first.
A few things are worth watching from here: whether independent evaluations replicate the reported ViDoRe v3 and nDCG@10 results across different document types and languages, and how storage costs play out for large corpora given the per-token vector approach.