← All models

Vision Medium

Balanced multimodal model pairing a 1M-token context window with deeper reasoning for analysis, content creation, and tool use across text, image, audio, video, and PDF inputs.

Context
1,000,000
Output
65,536
Release date
2024-05-15
Open weights
No

Availability from providers