ToolNavs Find Useful AI Tools
Submit Sign in
Back to AI information
Perplexity Open-Sources Multimodal Embedding Models: Text and Images Share One Retrieval Space

Perplexity Open-Sources Multimodal Embedding Models: Text and Images Share One Retrieval Space

AI information • Admin • • 6 views

Perplexity's new embedding models put text and images into a single retrieval space. On October 7, 2026, Perplexity CEO Aravind Srinivas announced the open-source release of pplx-embed-v2-late, with weights published on Hugging Face under an MIT license. The family uses multi-vector embeddings and ships in two sizes, 9B and 0.6B: the larger model indexes multimodal material, the smaller one can run on-device to issue queries, and because both share one embedding space, indexing and querying do not have to use the same model.

How it differs from ordinary embedding models

Most embedding models compress a whole passage into a single vector — fast, but fine detail is lost. A late-interaction design instead keeps one 128-dimensional vector per token and, at query time, matches each query token against its best counterpart, a scoring method known as MaxSim. The cost is larger indexes and more computation; the payoff is matching precision on long, technical queries. According to the model card, both models are built on Qwen3.5 with bidirectional attention and were distilled from an internal 18B teacher model.

Searching PDF pages without OCR is the headline use case

Perplexity says the 9B model can index multimodal data directly, so PDF pages can be retrieved as they are instead of being flattened into plain text by OCR first; the scores Srinivas cited in his launch post include 92.4% on MADQA and 64% on BrowseComp+. On the public ViDoRe v3 benchmark for visual document retrieval, the model card reports nDCG@10 of 65.2% on image documents and 64.7% on Markdown for the 9B model, and 62.3% and 61.2% respectively for the 0.6B model. These are the vendor's own reported figures; independent replication by the community is still to come.

What it means for teams building RAG systems

Enterprise document stores — contracts, manuals, policy files — are the most direct fit: layout information such as tables no longer has to survive an OCR pass before retrieval even starts, and splitting the work between a small on-device query model and a large indexing model lowers online serving costs. The trade-off to budget for is index size: token-level vectors take far more storage than a single vector per document, and recall latency has to be re-measured. Whether the switch pays off depends on how much of a collection's value sits in images and layout rather than plain text.

Recommended Tools

More