Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

In a between-subjects usability study, different participants test different versions of an interface; each person sees only one version. Use this approach when trying one design could teach someone how to use another. “Multi-user” can also mean testing several people at the same time, which is a separate decision: concurrent sessions do not change whether the study is between- or within-subjects.

What does between-subjects design mean in usability testing?

A between-subjects, or between-groups, design assigns different people to different study conditions. For example, one group tests the current interface and another tests a revised interface. Each participant uses only the version assigned to their group.

In a within-subjects, or repeated-measures, design, the same participants try multiple conditions. The distinction matters because it changes the risk of learning or carryover, the number of people to recruit, and how results should be analyzed. Nielsen Norman Group (NN/g) explains the trade-offs in its guide to between-subjects versus within-subjects study design.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Should you use between-subjects or within-subjects testing?

Choose based on what participants might learn from one condition and what conclusion the study needs to support. Neither design is universally better.

#1 Best Overall
Decision factor Between-subjects Within-subjects
Who tests each condition? Different participants test each version. The same participants test multiple versions.
Learning between versions Reduces the risk that experience with one version will affect another. Can be affected by learning, exposure, and order.
Recruiting Requires separate groups, so recruiting may take more effort. Can use fewer participants because each person contributes observations to multiple conditions.
Variation between people Differences in group composition can influence results. Comparisons involve the same people, which can reduce random variation between participants.
Analysis Compare outcomes between groups and account for how participants were allocated. Account for repeated observations and possible order effects.

Prefer between-subjects when exposure could change behavior

Use separate groups if completing a task, seeing a feature, or receiving feedback in one condition could make a participant more capable in a later condition. NN/g’s example of parallel design and testing used different participants for each version specifically to avoid skills learned in earlier tests affecting later ones. Its example included 10 people per design version; that number illustrates the study, not a universal recommendation. See NN/g’s explanation of parallel design and testing.

Prefer within-subjects when the comparison is fair and recruiting is constrained

Having the same people test multiple versions can make direct comparisons easier and reduce noise from differences among participants. But people may remember earlier tasks, transfer what they learned, or respond differently because of the order in which they saw the versions. If you use repeated measures, plan how to handle order and carryover; the cited guidance identifies these risks but does not prescribe one counterbalancing protocol for every study.

How many users do you need for each design?

There is no universal participant count that guarantees a reliable result. The needed number depends on whether the study is qualitative or quantitative, the outcomes and precision required, the conditions being compared, and how many user groups must be represented.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Qualitative issue finding: NN/g’s Jakob Nielsen describes five participants as a practical default for a qualitative usability study, with exceptions. This is not a promise that five people will find every problem, nor does it automatically mean five people per condition or user segment.
  • Quantitative comparison: NN/g says at least 20 participants may be needed for quantitative studies seeking statistically significant numbers, with more needed for tight confidence intervals. Treat this as general guidance, not a substitute for calculating sample size for the actual study and outcome.
  • Multiple user groups: NN/g’s planning checklist suggests 2–5 participants per group depending on how much groups overlap in experience and attitudes. This is a planning range, not a statistical sample-size rule.

These figures come from NN/g’s participant-count guidance for usability studies and its checklist for planning usability studies. A small qualitative sample can help identify usability problems, but it should not be treated as an estimate of population-wide task success. In particular, five users is not automatically enough for a statistical comparison between separate design groups.

How do you plan a between-subjects usability study?

  1. State the decision and outcome. Specify what the study needs to help you decide and what evidence would inform that decision. Distinguish exploratory problem discovery from a quantitative comparison.
  2. Define the conditions. Write down exactly what each group will use, such as the existing interface versus a revised one. Keep task wording and session procedures comparable so differences are not caused by unrelated changes.
  3. Set the participant criteria and segments. Recruit from the intended user population. If the product serves distinct groups, plan how each group will be represented in each condition rather than assuming one overall total covers them all.
  4. Choose a sample size suited to the goal. Use qualitative guidance for finding usability issues and a study-specific precision or power calculation for quantitative claims. Do not treat a rule of thumb as proof of statistical significance.
  5. Keep the comparison consistent. Use comparable tasks, moderator behavior, devices, context, and levels of assistance. Measure outcomes such as completion, errors, time, or observed breakdowns only when they answer the research question.
  6. Document allocation and analysis. Record who was assigned to which condition and how outcomes will be compared. In a between-subjects study, differences may reflect participant composition as well as interface design; allocation and sample size affect how confidently you can separate those influences.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Can you test multiple users at the same time?

Yes, but simultaneous testing is about session logistics, not the between- versus within-subjects design. If each participant independently tests one assigned version, multiple people can be tested at once while the study remains between-subjects. If everyone discusses the interface together, that discussion is not a substitute for observing independent task behavior.

NN/g recommends starting with one-on-one live-interface sessions before group discussion in its overview of multiple-user simultaneous testing. Digital.gov also notes that simultaneous independent testing followed by discussion requires enough note-takers to observe each participant. Without direct observation of each person, important differences in task behavior can be missed.

What conclusions can the results support?

Make claims that match the study design and sample. A qualitative round can surface observed problems and help prioritize follow-up work; it does not establish a population-wide success rate. A quantitative comparison needs enough participants and an analysis suited to the conditions, outcomes, allocation, and precision sought.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For separate groups, report meaningful differences in participant composition as a limitation. For repeated measures, account for the possibility that earlier exposure or task order shaped later behavior. In either case, comparable tasks and consistent procedures make the comparison easier to interpret.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.