Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

There is no single winner across this benchmark table. The ai-sage Hugging Face model card reports GigaChat 3.5 Ultra Reasoning slightly ahead on AIME 2025 and AIME 2026, while DeepSeek V4 Flash Preview Reasoning scores higher on several other listed tasks and on the card’s reported average. Read each row as a result for a specific benchmark and setup—not as a universal ranking. The figures below are reported by the model card and were not independently reproduced here.

How to read a benchmark table

  1. Start with the benchmark and protocol. A row measures performance on a particular task under a particular evaluation setup. Similar-looking scores from different rows need not measure the same capability.
  2. Compare models within the same row. Check the score direction, aggregation rule, sample count, judges, prompts, and any tools or time limits stated for that benchmark.
  3. Keep the model labels exact. The comparison names DeepSeek V4 Flash Preview Reasoning; these results do not automatically apply to another DeepSeek release or service configuration.
  4. Treat missing entries as missing. A dash is not a zero and should not be included as one when interpreting an average.
  5. Do not read small gaps as proven wins. The card does not provide uncertainty intervals or sufficient repeated-run detail to establish whether small differences are robust.

What does mean@32 mean?

In a benchmark label such as AIME 2025 (mean@32), “mean@32” identifies the aggregation reported for 32 attempts or samples. It is not directly interchangeable with a differently aggregated result, such as HMMT 2025 (mean@8). Read the aggregation label with the score rather than comparing bare numbers.

Can scores from different benchmarks be compared?

Not as though they were points on one common scale. A score of 90 on one benchmark does not necessarily indicate more overall capability than 80 on another: the tasks, scoring rules, and evaluation conditions may differ. Use each benchmark to answer a task-specific question.

Which AI model scored higher in this comparison?

The results are mixed. The following values are those reported in the ai-sage Hugging Face model card, not independently verified test results. Compare across models within each row only.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

STEM results

Benchmark GigaChat 3.5 Ultra Reasoning DeepSeek V4 Flash Preview Reasoning Higher reported score
AIME 2025 (mean@32) 89 88.95 GigaChat
AIME 2026 (mean@32) 92 90.4 GigaChat
HMMT 2025 (mean@8) 83.13 95.21 DeepSeek
IMOAnswerBench 73 85.75 DeepSeek
GPQA-Diamond 82.32 87.4 DeepSeek

General-task results

Benchmark GigaChat 3.5 Ultra Reasoning DeepSeek V4 Flash Preview Reasoning Higher reported score
IFBench 77 73.33 GigaChat
StructEval 85 80.19 GigaChat
MERA-2.0 42.3 not stated (ai-sage Hugging Face model card) Not comparable from the reported entries
Function Calling V4 58.59 68.06 DeepSeek
TAU3-bench 47.8 67.7 DeepSeek
Natural Plan 80.19 88 DeepSeek

Code results

Benchmark GigaChat 3.5 Ultra Reasoning DeepSeek V4 Flash Preview Reasoning Higher reported score
Live Code Bench v6 85.4 87.87 DeepSeek
SWE-bench Verified 64.7 78.6 DeepSeek
Terminal-Bench 2 30.3 56.6 DeepSeek

Arena results

Benchmark GigaChat 3.5 Ultra Reasoning DeepSeek V4 Flash Preview Reasoning Other model-card entry
Arena Hard Logs V3 56.5 53.7 —
Arena Hard Ru 60.7 36.8 —
Ru LLM Arena 64 48.5 —
Pollux 49 67.9 GigaChat Ultra Instruct: 71.6

Across the listed rows, GigaChat leads on both AIME entries and several Arena measures; DeepSeek leads on HMMT, IMOAnswerBench, GPQA-Diamond, multiple general and code tasks, and the reported average. Pollux also includes a distinct GigaChat Ultra Instruct result, which should not be mistaken for Ultra Reasoning’s score.

What evaluation details change how to interpret the scores?

The model card documents several protocol details that matter when reading specific rows:

  • IMOAnswerBench uses Qwen-3-235B-Instruct-2507 as judge.
  • TAU3-bench averages Airline, Retail, Telecom, and Banking.
  • Natural Plan uses a corrected scorer that normalizes UTF-8 characters to ASCII.
  • SWE-bench Verified and Terminal-Bench 2 use mini-swe-agent with a three-hour timeout.
  • Arena evaluations use MiniMax-M2.7 as judge and GPT-5.2 as baseline.
  • Benchmarks without a methodology-defined system prompt were run with an empty system prompt.

These conditions help define what was measured; they do not make unlike benchmarks directly comparable. A separate harness-benchmark project also cautions that a single run does not establish repeatability or significance for small score differences, and that changing harnesses or profiles can change the measured task. That general caution is not independent validation of the model-card results.

How much weight should the reported average get?

The model card reports an average of 68.88 for GigaChat Ultra Reasoning and 72.71 for DeepSeek. It does not explain cross-task normalization sufficiently to interpret that average as a general-purpose quality score. The missing DeepSeek MERA-2.0 value is also not zero; without the averaging method and treatment of missing entries, the average is not a substitute for examining the underlying rows.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What do the reasoning-token figures say?

The ai-sage model card states that GigaChat 3.5 Reasoning used 37% fewer reasoning tokens overall than DeepSeek V4 Flash Preview across four reported evaluation samples. Its table gives the following sample counts, mean token counts, and reductions:

Evaluation sample Samples GigaChat mean tokens DeepSeek mean tokens Reduction reported by the card
AIME 2025 240 13,980 19,129 27%
AIME 2026 240 13,635 17,697 23%
HMMT 480 13,311 19,553 32%
IMOAnswerBench 1,096 17,074 29,041 41%

These are figures published by the model card, not independently reproduced measurements. The reviewed page does not establish when it published them; the years in AIME benchmark names identify task versions, not the publication date of the measurements. Token use alone does not establish lower cost, faster responses, or equivalent performance in a matched deployment.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What model and deployment context is stated?

The model card describes GigaChat 3.5 Reasoning as a 432B Mixture-of-Experts model with 28B active parameters and a maximum supported context length of 262K tokens. It provides software inference instructions for frameworks and local tooling. Those details describe the card’s account of GigaChat; the reviewed evidence does not establish a matched cost, latency, or hardware comparison with DeepSeek.

The card also describes on-policy distillation with the sentence: “The student generates its own trajectory, while the expert for the corresponding domain provides token-level supervision on that trajectory.” The model-card text does not identify an individual speaker for this sentence.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What the table can—and cannot—establish

The table supports a task-by-task reading: GigaChat is narrowly ahead on the two AIME rows, while DeepSeek is ahead on a number of other listed benchmarks and on the reported average. The figures alone do not establish a universal winner. Their provenance is an ai-sage Hugging Face repository; the reviewed evidence does not confirm it as an official publisher for either model developer or independently verify every result. Treat them as reported comparison data, not official vendor claims or a firsthand test.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.