There is no single winner across this benchmark table. The ai-sage Hugging Face model card reports GigaChat 3.5 Ultra Reasoning slightly ahead on AIME 2025 and AIME 2026, while DeepSeek V4 Flash Preview Reasoning scores higher on several other listed tasks and on the card’s reported average. Read each row as a result for a specific benchmark and setup—not as a universal ranking. The figures below are reported by the model card and were not independently reproduced here.
How to read a benchmark table
- Start with the benchmark and protocol. A row measures performance on a particular task under a particular evaluation setup. Similar-looking scores from different rows need not measure the same capability.
- Compare models within the same row. Check the score direction, aggregation rule, sample count, judges, prompts, and any tools or time limits stated for that benchmark.
- Keep the model labels exact. The comparison names DeepSeek V4 Flash Preview Reasoning; these results do not automatically apply to another DeepSeek release or service configuration.
- Treat missing entries as missing. A dash is not a zero and should not be included as one when interpreting an average.
- Do not read small gaps as proven wins. The card does not provide uncertainty intervals or sufficient repeated-run detail to establish whether small differences are robust.
What does mean@32 mean?
In a benchmark label such as AIME 2025 (mean@32), “mean@32” identifies the aggregation reported for 32 attempts or samples. It is not directly interchangeable with a differently aggregated result, such as HMMT 2025 (mean@8). Read the aggregation label with the score rather than comparing bare numbers.
Can scores from different benchmarks be compared?
Not as though they were points on one common scale. A score of 90 on one benchmark does not necessarily indicate more overall capability than 80 on another: the tasks, scoring rules, and evaluation conditions may differ. Use each benchmark to answer a task-specific question.
Which AI model scored higher in this comparison?
The results are mixed. The following values are those reported in the ai-sage Hugging Face model card, not independently verified test results. Compare across models within each row only.
#1 Best Overall
STEM results
| Benchmark | GigaChat 3.5 Ultra Reasoning | DeepSeek V4 Flash Preview Reasoning | Higher reported score |
|---|---|---|---|
| AIME 2025 (mean@32) | 89 | 88.95 | GigaChat |
| AIME 2026 (mean@32) | 92 | 90.4 | GigaChat |
| HMMT 2025 (mean@8) | 83.13 | 95.21 | DeepSeek |
| IMOAnswerBench | 73 | 85.75 | DeepSeek |
| GPQA-Diamond | 82.32 | 87.4 | DeepSeek |
General-task results
| Benchmark | GigaChat 3.5 Ultra Reasoning | DeepSeek V4 Flash Preview Reasoning | Higher reported score |
|---|---|---|---|
| IFBench | 77 | 73.33 | GigaChat |
| StructEval | 85 | 80.19 | GigaChat |
| MERA-2.0 | 42.3 | not stated (ai-sage Hugging Face model card) | Not comparable from the reported entries |
| Function Calling V4 | 58.59 | 68.06 | DeepSeek |
| TAU3-bench | 47.8 | 67.7 | DeepSeek |
| Natural Plan | 80.19 | 88 | DeepSeek |
Code results
| Benchmark | GigaChat 3.5 Ultra Reasoning | DeepSeek V4 Flash Preview Reasoning | Higher reported score |
|---|---|---|---|
| Live Code Bench v6 | 85.4 | 87.87 | DeepSeek |
| SWE-bench Verified | 64.7 | 78.6 | DeepSeek |
| Terminal-Bench 2 | 30.3 | 56.6 | DeepSeek |
Arena results
| Benchmark | GigaChat 3.5 Ultra Reasoning | DeepSeek V4 Flash Preview Reasoning | Other model-card entry |
|---|---|---|---|
| Arena Hard Logs V3 | 56.5 | 53.7 | — |
| Arena Hard Ru | 60.7 | 36.8 | — |
| Ru LLM Arena | 64 | 48.5 | — |
| Pollux | 49 | 67.9 | GigaChat Ultra Instruct: 71.6 |
Across the listed rows, GigaChat leads on both AIME entries and several Arena measures; DeepSeek leads on HMMT, IMOAnswerBench, GPQA-Diamond, multiple general and code tasks, and the reported average. Pollux also includes a distinct GigaChat Ultra Instruct result, which should not be mistaken for Ultra Reasoning’s score.
What evaluation details change how to interpret the scores?
The model card documents several protocol details that matter when reading specific rows:
- IMOAnswerBench uses Qwen-3-235B-Instruct-2507 as judge.
- TAU3-bench averages Airline, Retail, Telecom, and Banking.
- Natural Plan uses a corrected scorer that normalizes UTF-8 characters to ASCII.
- SWE-bench Verified and Terminal-Bench 2 use mini-swe-agent with a three-hour timeout.
- Arena evaluations use MiniMax-M2.7 as judge and GPT-5.2 as baseline.
- Benchmarks without a methodology-defined system prompt were run with an empty system prompt.
These conditions help define what was measured; they do not make unlike benchmarks directly comparable. A separate harness-benchmark project also cautions that a single run does not establish repeatability or significance for small score differences, and that changing harnesses or profiles can change the measured task. That general caution is not independent validation of the model-card results.
How much weight should the reported average get?
The model card reports an average of 68.88 for GigaChat Ultra Reasoning and 72.71 for DeepSeek. It does not explain cross-task normalization sufficiently to interpret that average as a general-purpose quality score. The missing DeepSeek MERA-2.0 value is also not zero; without the averaging method and treatment of missing entries, the average is not a substitute for examining the underlying rows.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitchesRank #3
What do the reasoning-token figures say?
The ai-sage model card states that GigaChat 3.5 Reasoning used 37% fewer reasoning tokens overall than DeepSeek V4 Flash Preview across four reported evaluation samples. Its table gives the following sample counts, mean token counts, and reductions:
| Evaluation sample | Samples | GigaChat mean tokens | DeepSeek mean tokens | Reduction reported by the card |
|---|---|---|---|---|
| AIME 2025 | 240 | 13,980 | 19,129 | 27% |
| AIME 2026 | 240 | 13,635 | 17,697 | 23% |
| HMMT | 480 | 13,311 | 19,553 | 32% |
| IMOAnswerBench | 1,096 | 17,074 | 29,041 | 41% |
These are figures published by the model card, not independently reproduced measurements. The reviewed page does not establish when it published them; the years in AIME benchmark names identify task versions, not the publication date of the measurements. Token use alone does not establish lower cost, faster responses, or equivalent performance in a matched deployment.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.What model and deployment context is stated?
The model card describes GigaChat 3.5 Reasoning as a 432B Mixture-of-Experts model with 28B active parameters and a maximum supported context length of 262K tokens. It provides software inference instructions for frameworks and local tooling. Those details describe the card’s account of GigaChat; the reviewed evidence does not establish a matched cost, latency, or hardware comparison with DeepSeek.
The card also describes on-policy distillation with the sentence: “The student generates its own trajectory, while the expert for the corresponding domain provides token-level supervision on that trajectory.” The model-card text does not identify an individual speaker for this sentence.
Best Value
What the table can—and cannot—establish
The table supports a task-by-task reading: GigaChat is narrowly ahead on the two AIME rows, while DeepSeek is ahead on a number of other listed benchmarks and on the reported average. The figures alone do not establish a universal winner. Their provenance is an ai-sage Hugging Face repository; the reviewed evidence does not confirm it as an official publisher for either model developer or independently verify every result. Treat them as reported comparison data, not official vendor claims or a firsthand test.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

