Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Sometimes an AI should stand by a well-supported answer when a user challenges it; sometimes it should correct an earlier answer when stronger evidence appears. The Source-of-Belief Asymmetry Benchmark (SoBA) asks whether a language model handles that evidence differently depending on whether the original claim came from itself, a user, a document, or no named source. Rajan Mishra’s September 25, 2026 DEV Community article describes the benchmark and reports initial results, but those figures are the author’s claims—not independently verified evidence about commercial AI systems.

What SoBA is designed to test

SoBA focuses on belief attribution: who or what was said to support an earlier claim. The benchmark then presents contradictory evidence and assesses the model’s final answer. Its central question is not simply whether a model changes its mind, but whether the source attached to its prior claim affects how it evaluates new evidence.

Mishra describes four attribution conditions: the claim is attributed to the model’s own earlier response, to the user, to a document, or to no source. A model that treats comparable counter-evidence differently across these conditions may show an attribution-related bias. That alone would not establish why the difference occurs, or whether the same pattern appears in ordinary real-world use.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How the benchmark is structured

In Mishra’s description, each multi-turn trial establishes an initial claim, may add neutral conversational filler, introduces an attribution and counter-evidence, and then evaluates a final structured answer. The benchmark uses invented facts in fictional domains—including distributed systems, deep-space exploration, biotechnology, and fictional geopolitics and history—to reduce reliance on familiar facts learned during pretraining.

Alongside the source of the prior claim, the author says the design varies four factors:

  • Source reliability: whether the counter-evidence is reliable or unreliable.
  • Recency: whether the evidence is recent or outdated.
  • Repetition: whether it appears once or five times.
  • Conversation depth: whether the follow-up is immediate or delayed.

The article also describes adversarial cases intended to expose simplistic behavior: unsupported user pushback, an outdated document presented as official, and a fresh but unverified rumor. These cases matter because a model that always accepts a challenge could appear flexible while making errors whenever the challenge is weak.

How to read the benchmark’s metrics

The article names five measures. They are intended to capture both whether the final answer is correct and how the model responds to claims and counter-evidence:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Final Accuracy: the share of trials with a correct final answer.
  • Persistence Error Rate (PER): how often the model persists in an incorrect belief when it should revise.
  • False Revision Rate (FRR): how often it changes a correct belief when it should not.
  • Source Sensitivity Index (SSI): a measure the article uses to report response sensitivity to source conditions.
  • Self-Authority Bias (SAB): PER under self-attribution minus PER under document attribution.

Under the author’s definition, a positive SAB means persistence errors are more frequent when the model’s own earlier answer is the attributed source than when a document is. A negative value points in the opposite direction. SAB is a comparison between two conditions, not a general score for trustworthiness, truthfulness, or accuracy. It also does not directly compare self-attribution with every other attribution condition.

What the article reports

Mishra reports a dataset of 760 balanced multi-turn trials, with adversarial traps accounting for 16% of the dataset. The article says the metrics use 95% bootstrap confidence intervals. The following are the author’s reported profile results, not independently replicated measurements:

Profile label in the article Final accuracy PER FRR SSI SAB
Calibrated Reasoner 89.2% 8.9% 11.4% +68.2% +6.6% (reported interval: −2.1% to +17.4%)
Self-Protective Stubborn Sloth 90.4% 23.4% under document attribution; 65.2% under self-attribution Not stated in the article Not stated in the article +41.8% (reported interval: +22.4% to +59.2%)
Sycophantic Agent (FlipFlop) 45.1% 13.4% 67.6% +27.6% −4.0% (reported interval: −16.9% to +11.0%)
Repetition-Biased Reasoner 49.9% 8.4% 63.0% +23.5% +0.2% (reported interval: −10.5% to +11.0%)

The figures illustrate why a single measure can be misleading. The article’s “Self-Protective Stubborn Sloth” profile has the highest reported final accuracy in this set, but its PER is much higher under self-attribution than document attribution. The “Sycophantic Agent” and “Repetition-Biased Reasoner” profiles have high reported false revision rates, illustrating the opposite risk: accepting challenges or repeated claims too readily can undermine accuracy. The article also reports an 18.4% increase in false revision when an unreliable rumor was repeated five times.

The reported SAB intervals for the Calibrated Reasoner, Sycophantic Agent, and Repetition-Biased Reasoner include zero. For those results, the article’s reported intervals do not establish a nonzero difference at the interval level. The Self-Protective Stubborn Sloth interval is entirely above zero in the reported results. These are interpretations of the intervals as presented in the article, not independent statistical checks.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What the results do—and do not—show

The benchmark’s useful distinction is between appropriate persistence and stubbornness on one side, and appropriate revision and credulity on the other. A system should not accept a weak contradiction merely because it is repeated or voiced by a user; nor should it defend an incorrect answer just because the answer was its own. Looking at Final Accuracy, PER, FRR, and attribution-specific differences together gives a more informative picture than treating willingness to revise as inherently good or bad.

However, the profile names and reported scores do not establish that named commercial AI models behave this way. The article’s figures should be understood as results reported by its author for the described benchmark. They do not, on their own, support a broad claim that AI generally trusts itself more than users, or that a particular deployed assistant has a specific bias.

Limitations and what would strengthen the test

Mishra identifies the use of synthetic knowledge as a limitation: fictional facts help isolate belief revision from familiar-world knowledge, but they do not reproduce the complexity of judging real sources and evidence. Version 1, as described, tests direct contradictions within a session. Cross-session vector-memory behavior and more semantically subtle contradictions are future directions identified by the author, not capabilities established by the reported results.

The article points to a Kaggle notebook, dataset, and code path. Their availability, licensing, execution, and compatibility with current Kaggle benchmarking SDK APIs are not established here, so the benchmark should not be assumed to be ready to reproduce from those references alone. Independent evaluation would also need to check how trials are constructed and scored, whether the reported differences hold across models and prompts, and whether patterns transfer beyond synthetic, within-session contradictions.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why the benchmark’s question matters

For users, the practical issue is not whether an assistant should always trust itself or always defer to people. It is whether the assistant evaluates evidence consistently: a claim’s origin should not substitute for its reliability, freshness, or support. SoBA offers a way to frame that question and to measure two opposing failure modes. Its reported results are an early benchmark account, not a settled verdict about AI behavior in general.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.