Observations
Sampled Data
Sampled measurements of Earth's large language models. Search, filter by organization, or sort by score.
Demo data for layout preview only — replace with real evaluation results (src/data/benchmarks.json).
6 entries
| Model | Org | Params | MMLU | MATH | Code |
|---|---|---|---|---|---|
| Qwen3-235B | Alibaba | 235B / 激活 22B | 87.8 | 85.0 | 70.7 |
| DeepSeek-V3.1 | DeepSeek | 671B / 激活 37B | 88.5 | 89.3 | 74.8 |
| Kimi-K2 | Moonshot AI | 1T / 激活 32B | 89.5 | 87.5 | 76.1 |
| GLM-4.6 | Zhipu AI | 355B / 激活 32B | 86.2 | 84.1 | 72.5 |
| Llama-4-Maverick | Meta | 400B / 激活 17B | 85.5 | 80.2 | 68.9 |
| Mistral-Medium-3 | Mistral | 未公开 | 84.1 | 78.6 | 66.3 |