Instruction-tuned image reasoning model with 90B parameters from Meta. Optimized for visual recognition, image reasoning, captioning, and answering general questions about an image. The model can understand visual data, such as charts and graphs and also bridge the gap between vision and language by generating text to describe images details Note: This mode is served experimentally as a serverless model. If you're deploying in production, be aware that Fireworks may undeploy the model with short notice.
On-demand deployments give you dedicated GPUs for Llama 3.2 90B Vision Instruct using Fireworks' reliable, high-performance system with no rate limits.
Learn MoreMeta
131072
$0.9