Process both images and text to generate responses, multimodal understanding
Browse 0 models and agents with image-text-to-text capabilities