Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Often, yes. Recent studies of large language models (LLMs) find that the same moral dilemma or the same evaluation case can receive a different verdict when the wording is altered, the point of view shifts, the answer options are reordered, or the evaluation instructions change. The underlying facts stay the same; the presentation does not. That gap between facts and presentation is the core of the question, and it is also where careless conclusions tend to start.

The “boundary” in the title is best read as the frame around a question rather than a line drawn through the facts. The strongest evidence concerns AI systems, specifically how they judge disputes and how they grade other models’ outputs. Evidence about human juries or courts is much thinner, and it is covered separately below.

What counts as moving the boundary

Changing a question’s frame can mean very different things. Some changes alter only the surface of the text. Others change who is telling the story, what the answer options look like, or how the judge is instructed. The studies discussed here separate these effects, and a reader who lumps them together will misread the results.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Change to the question What is meant to stay fixed Reported effect in the studies
Surface wording (rephrasing sentences, reordering non-factual text) Facts and moral conflict 7.5% of verdicts flipped under surface perturbations, within the authors’ self-consistency noise floor of 4–13% (van Nuenen and Sachdeva, 2026 preprint)
Point of view (retelling the same events from another participant’s perspective) Events and outcomes 24.3% instability under point-of-view shifts (van Nuenen and Sachdeva, 2026 preprint)
Response mode (free-form graded rating versus forced yes/no) The case being judged Cross-form incoherence of 0.12–0.21 on a ±1 axis for graded ratings by the tested frontier models (Huang, 2026 arXiv paper; a study-specific measure)
Answer labels and option order (A/B labels, which option appears first) The same two options Order and lexical-label effects that could remain even when verdict-attached logical bias was approximately zero (Huang, 2026 arXiv paper)
Evaluation protocol (how instructions are structured and where they are placed) The same model response 67.6% agreement between structured evaluation protocols (κ=0.55); 35.7% of model-scenario units matched across all three protocols tested (van Nuenen and Sachdeva, 2026 preprint)

Keep these axes apart when you compare results. A flip caused by label order says something different from a flip caused by a change of narrator, and neither is measured by a rating-scale experiment.

A verdict flip is not the same as a changed stance

When a binary answer changes, it is tempting to conclude that the model’s view changed. The 2026 arXiv paper by Haonan Huang argues that this conclusion needs a check. The study separated answer-order and word effects from a effect attached to the verdict itself. For the tested frontier models using arbitrary A/B labels, the verdict-attached logical bias was approximately zero, while surface label and order effects could still appear.

In the author’s words: “the models are not drawn toward rejecting – the pull follows the printed surface, not the verdict it carries.” This is the author’s own statement about the paper’s findings, not an official consensus position, and it describes the models and setup that paper tested.

What the studies measured

Huang, 2026 arXiv paper: separating order, labels and verdict

The paper’s central methodological point is that a question should be crossed through equivalent frames rather than asked once. In the author’s words: “Measuring what an AI values requires crossing the frames of the question, not asking once.”

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Its numbers describe apparent binary yes/no bias for the tested models:

  • Claude Sonnet: a story-averaged result of −0.32, decomposed into order bias of −0.18 and lexical pull of −0.14.
  • Claude Haiku: a result of −0.86, decomposed into −0.33 for order and −0.53 for lexical pull.
  • GPT-5.5 and the tested Gemini models: approximately zero on that measure.

These figures describe specific models under the paper’s setup. They are not a ranking of which system judges better, and a near-zero score for one model does not show that it is free of framing sensitivity on other tasks.

Van Nuenen and Sachdeva, 2026 preprint: moral dilemmas from a public forum

This preprint evaluated 2,939 dilemmas drawn from r/AmItheAsshole, posted between January and March 2025. Four models produced 129,156 judgments. The authors generated perturbations of the stories and measured how often verdicts changed. Surface perturbations flipped 7.5% of verdicts, a figure the authors place within their own self-consistency noise floor of 4–13%. Point-of-view shifts produced 24.3% instability, which is well above that noise floor.

The authors draw a broad conclusion from these results: “These results show that LLM moral judgments are co-produced by narrative form and task scaffolding, raising reproducibility and equity concerns when outcomes depend on presentation skill rather than moral substance.” This is the authors’ interpretation, and it rests on their generated perturbations and tested models.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

JudgeSense, 2026 benchmark: when the judge is an AI

The JudgeSense benchmark, described in a 2026 abstract on alphaXiv, applies the same idea to LLM-as-a-judge evaluation. It covers 880 items, four evaluation tasks, and 25 judges from six providers. Rewording reduced agreement on all four tasks, and the effects met the authors’ practical-meaning threshold on two of them. The abstract describes this result for the benchmark it reports; it does not establish that every judge system behaves the same way.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to read these numbers

  • Each result is conditional on the models, cases and procedures tested. A finding about one model family is not a claim about all models, all people, or legal decision-makers.
  • A percentage only means something next to its noise floor. A 7.5% flip rate is close to the self-consistency noise the authors measured, while 24.3% is not.
  • Agreement figures from different protocols measure protocol choice as well as the judge. Kappa of 0.55 is moderate agreement, not reliability in general.
  • Dilemmas from one forum are a specific kind of input. They are narrative, informal and often one-sided, which may differ from formal cases.

Human and legal verdicts: a narrower claim

The title can also be read as a claim about juries or courts, but the sources examined here do not establish a specific court ruling or legal boundary that moves in this way. A 2018 analysis of expert witness testimony makes a related point. It argues that scientific evidence should be understood within the wider context of legal adjudication, and that fact-finders must connect evidence to legal concepts. It also notes that a scientifically validated general proposition does not guarantee the factual and normative rectitude of a particular verdict.

That is a reason for caution, not proof of human framing effects. The LLM experiments above should not be used as direct evidence about how people or courts decide.

How to test whether an LLM judge is prompt-sensitive

The following procedure follows the designs and caveats reported in the studies above. No single step eliminates bias, but together they separate framing effects from noise.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Write the case once. Record the facts and the moral conflict in a neutral, numbered list. Every variant you create must be logically equivalent to this record.
  2. Build surface variants. Reword sentences and reorder non-factual text. Keep names, actions and outcomes identical.
  3. Build perspective variants. Retell the same events from each party’s point of view. Check that no fact was added, dropped or softened.
  4. Counterbalance labels and order. Present the options as A/B and as B/A, and compare them with a yes/no framing. Record any difference that follows the label rather than the content.
  5. Ask for a graded rating as well as a binary verdict. If the two disagree, report which one you are interpreting.
  6. Repeat each condition several times. Use the same model, prompt and settings to estimate the model’s own noise. Only changes larger than that noise should be treated as meaningful.
  7. Record the full setup. Note the model name and version, the exact prompt text, the response format and the task structure, so another person can repeat the test.

Read the results with a simple rule. If a verdict changes only when the labels or order change, the evidence points to surface sensitivity. If it changes under a perspective shift and the change exceeds the noise you measured, the framing has altered the judgment in a way that deserves a closer look. In either case, the claim you can make is about that model, under that protocol, on that set of cases.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.