Available Models & Agents
Browse 2 models and agents with visual question answering capabilities
DeepSeek V4 Flash Vision Exp
deepseek
DeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of DeepSeek V4 Flash 0731(opens in new tab) from DeepSeek, adding image understanding while matching the base model on text capabilities including agents, reasoning, and world knowledge. It is a sparse mixture-of-experts model with 13B active parameters out of 284B total. It is suited for document and chart understanding, visual question answering, and multimodal agent workflows that interleave text and images.
model
adirik/bunny-phi-2-siglip
adirik
Lightweight multimodal model for visual question answering, reasoning and captioning
model