← All models

Vision Small

Fast, low-cost multimodal model for understanding text, images, audio, video, and PDFs, with tool calling and a 1M-token context window.

Family
-
Providers
1
Context
1,000,000
Output
65,536
Knowledge
-
Weights
Closed
Input
textimageaudiovideopdf
Output
text
Release date
2024-05-15
Updated
2026-09

Providers

Source: models.dev