
Evaluating Video In-Context Learning
for Multimodal Agents in Interactive Environments






Human demonstrations and recorded agent trajectories across every benchmark environment.
No matching games.
| Model | Overall | Navigation | Control | Strategy |
|---|---|---|---|---|
| Human | 94.2 | 96.8 | 93.4 | 92.5 |
| GPT-4o | 68.3 | 72.1 | 65.8 | 67.0 |
| Claude Sonnet 3.5 | 65.7 | 70.3 | 62.4 | 64.5 |
| Gemini 1.5 Pro | 63.2 | 68.9 | 59.8 | 61.0 |