Model comparisons
Compare the capability signal.
Public benchmark data helps you understand capability. Your private task telemetry tells you which model actually works best for you.
ModelRelative scoreScore
GPT-5.6 Sol
OpenAI
96
2
Claude Opus 4.8
Anthropic
94
3
Grok 4.5
xAI
89
4Google
Gemini 3.6 Flash
84
This is a clearly labeled demo composite used to exercise the multi-source benchmark architecture. Production benchmark feeds are the next integration step; no score is presented as an official provider claim.
Best for coding
GPT-5.6 Sol
96 / 100
Best instruction following
Claude Opus 4.8
97 / 100
Fastest
Gemini 3.6 Flash
96 / 100