Perplexity’s new contextual embedding model claims top spot in document retrieval

Perplexity’s new contextual embedding model claims top spot in document retrieval

The pplx-embed-v2-context-9b-preview encodes whole documents in a single pass, building on a v1 family that already beat rivals on the ConTEB benchmark

Perplexity AI has a new contextual embedding model, pplx-embed-v2-context-9b-preview, and the company says it sets a new state-of-the-art for document retrieval.

The trick is simple to describe and hard to pull off. The model encodes an entire document once, so every piece of it carries the context of the whole.

What the new model actually does

The problem is chunking. Long documents get sliced into smaller pieces before they are embedded, and those pieces often lose the plot.

A chunk that says “the company raised its forecast” is useless if the model has forgotten which company the document was about. Traditional approaches embed each chunk in isolation, like reading one page torn out of a novel.

Perplexity’s contextual approach tries to fix that. By processing the full document in one pass, the model lets each chunk inherit information from the surrounding text.

The model carries a “preview” tag, which signals it is an early release rather than a finished product. Perplexity is positioning it as the next step after its first generation of contextual models.

The v1 benchmark bar it has to clear

On February 26, 2026, the company released two embedding families: pplx-embed-v1 and pplx-embed-context-v1.

Advertisement

Both came in 0.6B and 4B parameter sizes.

On the ConTEB benchmark, pplx-embed-context-v1-4B hit an average nDCG@10 of 81.96%. nDCG@10 scores whether the most relevant results show up in the top 10, and whether the best ones land near the top rather than at slot nine.

For comparison, voyage-context-3 scored 79.45% on the same measure. A model from Anthropic came in at 72.4%.

Perplexity also reported strong recall gains over Qwen3 embeddings in internal web-scale benchmarks. Those tests analyzed over 1 billion production pages, though they were run by Perplexity itself rather than an independent party.

How Perplexity built it, and what it costs

The v1 training process started with diffusion-based techniques on approximately 250 billion multilingual tokens.

After that, the company used a multi-stage contrastive learning approach. The setup enables bidirectional attention, meaning the model looks at text on both sides of a word at once.

The models support a context window of 32K tokens.

The models support native INT8 and binary quantization. Perplexity says this delivers up to 32x storage savings.

The models also do not require instruction prefixes.

API pricing for the v1 lineup runs from $0.004 to $0.05 per million tokens, depending on the variant and size.

Why Perplexity is building its own plumbing

Perplexity is best known as an AI-powered answer engine. Building its own embedding models gives the company direct control over retrieval and turns internal infrastructure into something it can sell to outside developers.

Those developers are mostly working on semantic search and retrieval-augmented generation, or RAG. RAG is the technique of fetching relevant documents first, then handing them to a language model so it answers from evidence rather than memory.

What this means for developers and rivals

For developers, the practical questions are cost, storage, and switching effort. Quantization and the lack of instruction prefixes lower the operational burden, while the pricing range keeps experiments cheap.

There are caveats worth watching. Some of the most striking comparisons, such as the Qwen3 recall gains, come from Perplexity’s internal benchmarks, and a preview model may change before any general release.

Disclosure: This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our Editorial Policy.
Perplexity’s new contextual embedding model claims top spot in document retrieval
Perplexity’s new contextual embedding model claims top spot in document retrieval

The pplx-embed-v2-context-9b-preview encodes whole documents in a single pass, building on a v1 family that already beat rivals on the ConTEB benchmark

Perplexity AI has a new contextual embedding model, pplx-embed-v2-context-9b-preview, and the company says it sets a new state-of-the-art for document retrieval.

The trick is simple to describe and hard to pull off. The model encodes an entire document once, so every piece of it carries the context of the whole.

What the new model actually does

The problem is chunking. Long documents get sliced into smaller pieces before they are embedded, and those pieces often lose the plot.

A chunk that says “the company raised its forecast” is useless if the model has forgotten which company the document was about. Traditional approaches embed each chunk in isolation, like reading one page torn out of a novel.

Perplexity’s contextual approach tries to fix that. By processing the full document in one pass, the model lets each chunk inherit information from the surrounding text.

The model carries a “preview” tag, which signals it is an early release rather than a finished product. Perplexity is positioning it as the next step after its first generation of contextual models.

The v1 benchmark bar it has to clear

On February 26, 2026, the company released two embedding families: pplx-embed-v1 and pplx-embed-context-v1.

Advertisement

Both came in 0.6B and 4B parameter sizes.

On the ConTEB benchmark, pplx-embed-context-v1-4B hit an average nDCG@10 of 81.96%. nDCG@10 scores whether the most relevant results show up in the top 10, and whether the best ones land near the top rather than at slot nine.

For comparison, voyage-context-3 scored 79.45% on the same measure. A model from Anthropic came in at 72.4%.

Perplexity also reported strong recall gains over Qwen3 embeddings in internal web-scale benchmarks. Those tests analyzed over 1 billion production pages, though they were run by Perplexity itself rather than an independent party.

How Perplexity built it, and what it costs

The v1 training process started with diffusion-based techniques on approximately 250 billion multilingual tokens.

After that, the company used a multi-stage contrastive learning approach. The setup enables bidirectional attention, meaning the model looks at text on both sides of a word at once.

The models support a context window of 32K tokens.

The models support native INT8 and binary quantization. Perplexity says this delivers up to 32x storage savings.

The models also do not require instruction prefixes.

API pricing for the v1 lineup runs from $0.004 to $0.05 per million tokens, depending on the variant and size.

Why Perplexity is building its own plumbing

Perplexity is best known as an AI-powered answer engine. Building its own embedding models gives the company direct control over retrieval and turns internal infrastructure into something it can sell to outside developers.

Those developers are mostly working on semantic search and retrieval-augmented generation, or RAG. RAG is the technique of fetching relevant documents first, then handing them to a language model so it answers from evidence rather than memory.

What this means for developers and rivals

For developers, the practical questions are cost, storage, and switching effort. Quantization and the lack of instruction prefixes lower the operational burden, while the pricing range keeps experiments cheap.

There are caveats worth watching. Some of the most striking comparisons, such as the Qwen3 recall gains, come from Perplexity’s internal benchmarks, and a preview model may change before any general release.

Disclosure: This article was edited by Diego Almada Lopez. For more information on how we create and review content, see our Editorial Policy.