← All models
Vision Small
Fast, low-cost multimodal model for understanding text, images, audio, video, and PDFs, with tool calling and a 1M-token context window.
- Family
- -
- Providers
- 1
- Context
- 1,000,000
- Output
- 65,536
- Knowledge
- -
- Weights
- Closed
- Input
- textimageaudiovideopdf
- Output
- text
- Release date
- 2024-05-15
- Updated
- 2026-09
Providers
Source: models.dev