← All models

Vision Small

Fast, low-cost multimodal model for understanding text, images, audio, video, and PDFs, with tool calling and a 1M-token context window.

Context
1,000,000
Output
65,536
Release date
2024-05-15
Open weights
No

Availability from providers