
voyage-multimodal-3.5 is a high-accuracy multimodal embedding model for retrieval across text, images, PDFs, screenshots, tables, figures, slides, and videos. It is designed for production search and RAG workflows that need to retrieve visually rich and mixed-format content.
The model embeds text, visual documents, and video frames into a shared vector space, helping teams build retrieval systems where similarity reflects semantic meaning across modalities. It supports 2048, 1024, 512, and 256 output dimensions, along with multiple quantization options for efficient storage and retrieval.
On-demand DeploymentDocs | On-demand deployments allow you to use voyage-multimodal-3.5 on dedicated GPUs with Fireworks' high-performance serving stack with high reliability and no rate limits. |