Skip to main content

Perplexity Releases pplx-embed-v2-context-9b-preview

Perplexity launches pplx-embed-v2-context-9b-preview, a 9B contextual embedding model on Hugging Face featuring late chunking for evidence retrieval.

AI-written
Inewgen
01 Oct 2026Source: MarkTechPost4 min read (0 views)
Share
Perplexity Releases pplx-embed-v2-context-9b-preview

Stock photo for illustration only, not from the actual event

Font size
  • Perplexity releases the pplx-embed-v2-context-9b-preview 9B contextual embedding model as a self-hosted preview
  • Weights are published on Hugging Face under the MIT license, requiring transformers>=5.4.0
  • Addresses traditional RAG limitations using late chunking and a teacher compression model for soft targets
  • Supports 2048 and 1024 dimensions via Matryoshka training with native int8 embedding quantization

Perplexity has introduced a brand-new model named pplx-embed-v2-context-9b-preview, a contextual embedding model engineered to accurately retrieve answers alongside their supporting evidence. The model is currently available as a self-hosted preview, with its weights officially published on Hugging Face under the MIT license.

Developers interested in testing the model must use the transformers library version 5.4.0 or higher, along with setting trust_remote_code=True. However, the model card notes that both the weights and the interface may change without backward compatibility guarantees, and the model is not yet accessible via the Perplexity API.

machine learning code screen workspace

Stock photo for illustration only, not from the actual event

In traditional Retrieval-Augmented Generation (RAG) systems, long documents are split into smaller chunks. Often, a chunk heavily depends on entities, definitions, or headings mentioned elsewhere in the document. Contextual models tackle this challenge through a technique known as Late Chunking, where the document is encoded in a single pass before being pooled per chunk.

Late Chunking represents a major leap forward for RAG architectures. Traditional token splitting often severs crucial contextual ties, causing pronouns and cross-references to lose their intended meaning. By preserving these relationships within vector representations, Perplexity's approach significantly enhances performance on complex documents like legal contracts or technical research papers.

Standard training pipelines typically assign one gold chunk per query, treating all other chunks as negatives—including sentences essential for verifying the answer. Perplexity outlines three primary drawbacks to this conventional approach:

Never miss the latest news?

Subscribe to get news summaries by email - not often enough to be annoying.

โฆษณา

  • Binary labels provide only a coarse signal for training.
  • LLM annotation costs scale linearly with the overall dataset size.
  • Labels remain strictly tied to a single chunking strategy.

To overcome these hurdles, Perplexity employs a query-aware context compression model acting as a teacher. It reads the query and document simultaneously to score every token. During each training batch, a random chunking strategy is sampled, separating chunks with a learned token before mean-pooling. This teacher model operates solely during training, adding zero latency or storage overhead during inference.

9BBase ColBERT parameter size
2048Max dimensions (supports 1024)
430+Datasets across 50+ languages

The model originates from an in-house 9B ColBERT retrieval model. A linear projection yields 2048 dimensions, while Matryoshka training also accommodates 1024 dimensions alongside quantization-aware training for native int8 embeddings. Released as a model soup combining several checkpoints, the training utilized approximately 430 datasets spanning over 50 languages without any ConTEB data.

neural network graph analytics dashboard

Stock photo for illustration only, not from the actual event

For evaluation benchmarks, Perplexity cites results from context-bench, built and privately held by turbopuffer to minimize training contamination. The benchmark encompasses 2,099 queries, 38,894 documents across 21 domains, and 2,458,072 sentence chunks, testing 12 distinct contextual capabilities ranging from pronoun resolution to table structure parsing.

Source: MarkTechPost

Comments

Leave a Comment
0/2000

Found something wrong in this article? Report an issue with this article