← All models
Vision Small
Fast, low-cost multimodal model for understanding text, images, audio, video, and PDFs, with tool calling and a 1M-token context window.
- Context
- 1,000,000
- Output
- 65,536
- Release date
- 2024-05-15
- Open weights
- No