Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

A 1% sample is not automatically too small—but “we review 1% because the judge costs too much” is not a statistical justification. The fraction alone says nothing about how many cases were reviewed, whether they represent the cases you care about, whether the AI judge agrees with people, or how uncertain your conclusion is. Choose the sample to support a specific claim, then validate the judge and account for how cases were selected.

Start with the claim you need the evaluation to support

The right sample depends on the decision, not on a conventional percentage. Before setting a review budget, write down the quantity you want to estimate and what decision will follow from it. A sample designed to estimate average response quality may not be adequate for comparing two models or detecting rare failures.

Average quality

If the question is whether a model meets a quality target on a defined population of cases, specify the target population, quality measure, and acceptable uncertainty around the estimate. The cases available to review must be sampled in a way that supports inference to that population.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Comparing models or detecting a regression

A comparison needs enough information to distinguish the difference that matters to your decision from ordinary variation. Define the comparison and a practically meaningful difference in advance; then plan the human-reviewed sample around the expected uncertainty or statistical power. A percentage that looks substantial can still leave a small absolute number of paired, informative cases.

Estimating rare failures

If the concern is a low-frequency but consequential failure, a small random sample may contain too few examples to estimate its rate precisely—or may contain none. You can deliberately oversample relevant edge cases, but that changes the sampling design. Analyze those cases as a defined stratum or use a method that accounts for their selection; do not report the resulting mix as if it were a simple random sample of ordinary traffic.

Use human ratings to calibrate the judge, not just to spot-check it

One practical design in the 2026 paper Augmenting Human Evaluation with LLM Judges: How Many Human Reviews Do You Need? treats the judge as an aid to human evaluation rather than a substitute. The LLM rates all observations at the first stage; people rate a planned subset at the second stage; a doubly robust estimator combines the two sources. The authors also use asymptotic variance to plan human- and LLM-rating sample sizes for a target power.

This is a proposed design, not a universally optimal recipe. It does, however, make the roles clear: judge scores can provide broad coverage, while human ratings provide evidence about the relationship between those scores and the human assessment the evaluation is meant to represent. A small human sample used only to say “the judge seems fine” does not, by itself, establish that relationship or support a precise result.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Plan the human sample against uncertainty and power

Estimate the number of human ratings needed for the intended inference using the outcome variability, the comparison or target, and the uncertainty or power you require. The reviewed sources do not establish one optimal review percentage, universal minimum count, or general cost threshold for AI-reviewer benchmarks. If the available budget cannot support the planned inference, report the result as exploratory or narrow the claim rather than presenting a convenient fraction as proof.

Make the sampling frame visible

Record which cases could have been selected, how cases were drawn, and whether selection probabilities differed. Randomly selected human ratings and deliberately selected difficult examples answer different questions. For stratified or targeted review, preserve the strata and use an estimator that reflects the design; do not silently treat an intentionally non-random sample as random.

Test two different properties: human alignment and judge reliability

A judge can be stable without agreeing with people, or agree on average while changing its rating when the prompt changes. These are distinct checks. The ICML 2026 paper Diagnosing the Reliability of LLM-as-a-Judge via Item Response Theory distinguishes consistency under prompt variation from alignment with human assessments. Evaluate both for the judge configuration and task you actually use.

Alignment with human assessments

Compare judge ratings with human ratings on the human-reviewed cases using an appropriate agreement or association measure. Describe the measure, the rating scale, the reviewed sample, and uncertainty around the result. The 2026 evalstats preprint, How to Run Statistics over LLM Judges and Trust the Results: Calibrated Inference for Small-Sample AI Evaluation with evalstats, gives ρ² ≥ 0.4 as a rough range where meaningful gains from mixed judge-human designs begin, and ρ² < 0.2 as a range where its authors advise that the judge is too poor to use. These are that paper’s rules of thumb, not universal pass/fail thresholds; task, data, and metric still matter.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Stability under prompt variation

Check whether changing prompt wording or other relevant configuration details changes the judge’s ratings. If the evaluation depends on a particular prompt, preserve its exact text and configuration. A result that is sensitive to small prompt changes needs to be interpreted differently from one that is stable under the variations you tested.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Report enough for someone else to audit the result

The evalstats preprint recommends reporting details that let readers judge how much confidence to place in mixed human-AI estimates. Include the information needed to reconstruct the evaluation and understand its uncertainty:

  • Target and decision: the population, outcome or estimand, and the claim or decision the evaluation is intended to support.
  • Judge configuration: exact model identifier, prompt, and relevant settings, including any changes made during the evaluation.
  • Sampling design: the eligible cases, human-reviewed count, selection method, strata or oversampling, and how those choices enter the analysis.
  • Human-rating process: who rated the cases, the rubric and process, how disagreements were handled, and whether raters were blinded to judge scores or model identity where relevant.
  • Validation: the human–judge alignment measure and its uncertainty, plus the prompt-stability checks performed.
  • Inference: the estimator or test used, its assumptions, the uncertainty reported, and effective sample size where applicable.

The evalstats methods analyze a missing-completely-at-random setting that assumes random selection of the human-rated subset. A different selection scheme needs an analysis that reflects it. The paper also notes limited validation for some between-subjects pairwise confidence intervals, so do not assume every reported interval has been broadly validated across designs.

A practical decision rule for a constrained budget

  1. Define the claim. Decide whether you need an average, a model comparison, a regression signal, or a failure-rate estimate.
  2. Set the required precision or power. Specify the uncertainty you can tolerate or the effect you need to detect.
  3. Choose and document the human-review sample. Use a selection plan suited to the target population; if you oversample strata or edge cases, retain that design in the analysis.
  4. Validate before relying on judge scores. Measure human alignment and test stability under relevant prompt changes.
  5. Use an estimator and uncertainty calculation that match the design. A mixed human-judge approach can use information from both sources, but its assumptions must fit how the cases were selected.
  6. Match the conclusion to the evidence. If the budget cannot support the intended inference, call the result exploratory, collect more human ratings, or make a narrower claim.

There is no evidence-backed universal rule that 1% is wrong—or that 5% is right. The failure is stopping at the percentage instead of showing what the selected cases and validated judge can establish.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.