Instruction-tuned image reasoning model from Meta with 11B parameters. Optimized for visual recognition, image reasoning, captioning, and answering general questions about an image. The model can understand visual data, such as charts and graphs and also bridge the gap between vision and language by generating text to describe images details
On-demand deployments give you dedicated GPUs for Llama 3.2 11B Vision Instruct using Fireworks' reliable, high-performance system with no rate limits.
Learn MoreMeta
131072
$0.2