I looked through some AI benchmarks and it seems to me that Open AI, Google, and Anthropic are dominating. The Chinese models seem to be the best open source models and they are not far behind the other models. I think that X's Grok is not far behind also.
IMHO, the "Best" benchmark which has images as part of its input is MMMU (https://mmmu-benchmark.github.io).
Top 5 on MMMU: GPT4o, Claude 3.5 Sonnet, Gemini 1.5 Pro, Claude 3 Opus, and Gemini 1.0 Ultra.
Best Chinese Model is Qwen-2.5-VL-72B which is about 10th.
IMHO, the "Best" text based benchmarks are MMLU and ARC
Top 5 MMLU: GPT-5, GPT-4.1, Claude Opus 4.1, Gemini 2.5 Pro, and Llama 3.1 405B.
Best Chinese Model: DeepSeek-R1
Top 5 ARC: Llama 3.1 405B, Claude 3 Opus, Claude 3.5 / 4.5, AI21 Labs, Meta / Mixed models.
Chinese Model: No data found.