Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsIn Debashish Ghosal’s v0.1.0 field test of AdversarialDebate, DeepSeek + Mistral led on average score and verdict rate—but Ghosal also reported a 65% capitulation rate for that pair. The result shows why a high convergence or verdict score does not, by itself, tell you whether two models reached a sound conclusion: one model may simply have conceded without rebuttal.
These are results from one author-run test, not proof of a universally best or worst model pair. Ghosal’s later versions changed the ranking and added uncertainty that further narrows what the numbers can support.
Why the top score did not settle the question
AdversarialDebate is a multi-agent debate and review system. In Ghosal’s v0.1.0 report, DeepSeek + Mistral had the highest average score and verdict rate among the listed pairs. But those outcome measures do not reveal how a debate reached its endpoint. Agreement after evidence and rebuttal is different from one side yielding immediately, even if both interactions are recorded as resolved.
Ghosal’s framing captures the measurement problem: “Strong-looking multi-agent metrics are not automatically trustworthy multi-agent metrics.” He also poses the consequential question: “Should a verdict reached through capitulation count as a verdict at all?” A verdict can be useful as an outcome label, but it should not be treated as evidence of a robust review unless the process behind it is measured too.
#1 Best Overall
What the v0.1.0 results report
The figures below are attributed to Debashish Ghosal’s 2026 v0.1.0 field test. The article’s figures have not been independently audited or replicated here. “Average score” and verdict rate are the reported outcome measures; the concession count is included as a process clue, not as a rate.
| Model pair | Average score | Verdict rate | Concessions |
|---|---|---|---|
| DeepSeek + Mistral | 0.982 | 97% | 2,352 |
| GPT + Mistral | 0.754 | 48% | 1,728 |
| GPT + GPT | 0.688 | 57% | 1,444 |
| Gemini + DeepSeek | 0.622 | 10% | 1,470 |
| Gemini + Mistral | 0.512 | 4% | 1,073 |
| GPT + Gemini | 0.357 | 4% | 727 |
Across 411 debates in v0.1.0, Ghosal classified 80—19%—as capitulation cascades. The reported rate was 65% for DeepSeek + Mistral and 0% for GPT + Gemini. The cascade rule was at least 80% round-one concessions and zero rebuttals. That rule identifies a particular pattern; it does not establish that every concession is unjustified or that every debate outside the rule involved meaningful scrutiny.
Why convergence and verdict rate need process measures
A convergence measure can combine two very different paths: agents may exchange evidence and move toward agreement, or one agent may fold before a substantive challenge. As Ghosal puts it, “A 1.0 score can mean: 1. both sides genuinely converged after evidence exchange 2. one side folded immediately”. If a dashboard reports only the endpoint, those paths can look equivalent.
The reverse problem also matters. A pair that rarely reaches a verdict may be failing to converge, while a pair that reaches one almost every time may be converging too easily. Neither a low nor a high outcome score explains the interaction by itself. For a useful comparison, read outcome measures alongside process measures and the scope of the test.
Recommended Free Tools
Rank #3
- Convergence or average score: How often or how strongly the system reports agreement, according to its scoring method.
- Verdict rate: How often a debate produces a verdict, distinct from whether the verdict was well supported.
- Capitulation: Whether one side concedes early, especially without a rebuttal.
- Rebuttal activity: Whether claims were challenged before agreement.
- Corpus and subset: Which artifacts were tested, and whether results cover the full corpus or a smaller validation subset.
- Uncertainty: How much noise or statistical uncertainty could affect pair rankings, particularly for small samples.
How the pair rankings changed across versions
The v0.1.0 ranking did not remain the project’s operational recommendation. In later versions, the reported comparison changed in both scope and interpretation. The figures in this table are Ghosal’s 2026 reports for each version; they are not a single like-for-like ranking because v0.2.0 reports a full-corpus result alongside validation-subset results.
| Version and context | Pair | Reported result |
|---|---|---|
| v0.2.0, full corpus | GPT + Mistral | Average convergence 0.536; 2/150 verdicts; 2,927 concessions |
| v0.2.0, validation subset | DeepSeek + Mistral | Average convergence 0.572; 1/36 verdicts; 936 concessions |
| v0.2.0 | GPT + Gemini | Average convergence 0.033; 0/24 verdicts |
| v0.2.1, same 150-artifact corpus | DeepSeek + GPT-4o-mini | Average convergence 0.246 |
| v0.2.1, comparison values reported by the author | GPT + GPT | Average convergence 0.273 |
| v0.2.1, comparison values reported by the author | GPT + Mistral | Average convergence 0.536 |
| v0.2.1, comparison values reported by the author | DeepSeek + Mistral | Average convergence 0.572 |
| v0.2.1, comparison values reported by the author | GPT + Gemini | Average convergence 0.033 |
v0.2.0: a different operational default
Ghosal described GPT + Mistral as the full-corpus default in v0.2.0, while treating DeepSeek + Mistral as a validation pair. The reported 0.572 for DeepSeek + Mistral came from a 36-item validation subset, not the 150-artifact full-corpus run reported for GPT + Mistral. Comparing those two values as if they were measured on identical samples would overstate what the table establishes.
Rank #4
v0.2.1: another pair, not a final answer
In v0.2.1, Ghosal added DeepSeek + GPT-4o-mini on the same 150-artifact corpus and reported an average convergence score of 0.246, below the reported 0.273 for GPT + GPT and 0.536 for GPT + Mistral. The article interpreted this comparison as evidence that Mistral’s participation mattered more than simply mixing labs. But a comparison of these pairs does not isolate model identity from all other factors in the setup, and the reported scores alone do not establish why they differ.
The v0.2.1 release note reproduced in the article reports a 1.7–3.4% missed-issue rate, described as first recall data, and 55 new unit tests. Those release details do not independently validate the pair comparison or establish that a high convergence score corresponds to a reliable verdict.
v0.2.2: the lead is uncertain
In v0.2.2, Ghosal characterized the 0.572 versus 0.536 gap as 1.8 sigma, making it a narrow and uncertain difference rather than a decisive separation. He also raised shared RLHF conversational defaults as an alternative explanation: pair behavior might reflect similar conversational priors among models, rather than a uniquely beneficial Mistral effect. The update further described noise floors for pairs with n<30 as very wide. Taken together, those qualifications do not support a categorical rule that Mistral should always be included.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.What this field test can—and cannot—show
The version history is useful because it shows how adding process measures, comparison pairs, and uncertainty can change the interpretation of an apparent winner. It does not establish a universal best model pair. The reported results belong to Ghosal’s author-run AdversarialDebate tests, and the article is not an independent benchmark or controlled study.
Ghosal’s article does not establish exact prompts, provider settings, precise model snapshots, costs, or the full experimental protocol. Those details matter for reproducing the results and judging whether a ranking would transfer to another deployment. The pair labels should therefore be read as labels in this particular test, not as timeless descriptions of stable model products.
The article is Debashish Ghosal’s report on DEV Community, published August 29, 2026. Its results support a practical evaluation lesson: record how a verdict was produced, not just whether the system produced one.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallQuick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

