← All models

Vision Medium

Balanced multimodal model pairing a 1M-token context window with deeper reasoning for analysis, content creation, and tool use across text, image, audio, video, and PDF inputs.

Family
-
Providers
1
Context
1,000,000
Output
65,536
Knowledge
-
Weights
Closed
Input
textimageaudiovideopdf
Output
text
Release date
2024-05-15
Updated
2026-09

Providers

Source: models.dev