⚡ Uncle Cat AI Radar
ModelsResearchOpen Source

Perplexity Releases Multimodal Late-Interaction Embeddings

Perplexity’s new 0.6B and 9B embedding models match text and images in one space, targeting more precise retrieval across rendered documents.

Perplexity has released pplx-embed-v2-late, a pair of late-interaction embedding models designed to retrieve text, images and rendered document pages through a shared representation space. The release follows the company’s earlier contextual embedding model and broadens its push beyond answer generation into the retrieval layer that feeds search and agent systems.

The models use multiple vectors per input instead of compressing an entire document or query into one embedding. That allows different parts of a query to match different parts of a page, preserving local detail that can disappear in single-vector retrieval. Perplexity says the system can also perform text-to-image retrieval, including searches over rendered PDF pages without first converting every page into machine-readable text.

The release includes a 0.6-billion-parameter query encoder and a 9-billion-parameter index model. Perplexity’s published material positions the smaller model for query-side use and the larger model for building the searchable index. That division is important for deployment economics: query infrastructure can remain relatively light while the heavier model is used offline or at indexing time.

This is not a headline model release, but it targets a bottleneck that increasingly determines whether AI search and agents are trustworthy: finding the right evidence before generation begins. Multimodal retrieval could improve access to charts, layouts and image-heavy reports, although the practical gains will depend on indexing cost, latency, licensing and evaluation outside Perplexity’s own tests. The strategic signal is that competition is moving deeper into the retrieval stack, where better representations can influence every downstream model call.

Sources