Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Measure AI’s effect on the work it actually changes: define a task-level baseline, compare results with a credible control where possible, and track output alongside quality, rework, customer or stakeholder outcomes, and worker experience. Track access and usage separately. Adoption shows exposure to AI—not that performance improved.

Define what “better performance” means for the work

Start with the task or workflow where people use the AI, then write down the specific change you expect. For example: “Reduce minutes per resolved support case without lowering resolution quality,” or “Increase accepted drafts per week without increasing rework.” “Improve productivity” is too vague to test.

Choose measures that fit the task and its risks. A support team might track cases resolved per hour, resolution quality, escalations, and customer sentiment. A team producing documents might track cycle time, completed and accepted work, accuracy, rework, and stakeholder response. NIST’s AI RMF Core — Measure calls for context-appropriate metrics and documented evaluation methods.

Build a baseline and a credible comparison

Capture performance before rollout using the same definitions you will use afterward. A before-and-after comparison can show that results changed, but it cannot by itself establish that AI caused the change. Staffing, workload, seasonality, process updates, or other factors may also explain it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When practical, randomly assign access or rollout timing. If randomization is not feasible, introduce the tool in phases and compare teams or tasks that are similar, with one group not yet using AI. Record the benchmark, sample, observation window, uncertainty, tool version, and deployment conditions. The customer-support study used a staggered introduction, while the cross-firm knowledge-worker study randomly selected workers for access. These designs support stronger attribution than anecdotes, but their results do not automatically predict what will happen in another organization.

Track AI access and usage separately from results

Record who was eligible, who received access, who used the system, how often they used it, and which tasks they used it for. Where relevant, note whether AI output was accepted, edited, or discarded. These are exposure and use measures—not performance outcomes.

Usage data can help explain a result: low use may help explain why a rollout had little effect, while frequent use can coexist with neutral or negative results. In a six-month randomized experiment across 66 firms and 7,137 knowledge workers, frequent users in the second half spent two fewer hours per week on email; researchers did not detect changes in task quantity or composition attributable to individual access. The study illustrates why time saved should not be treated as proof of more output.

Balance speed and volume with quality and value

Pair productivity measures with checks that reveal whether faster work is also useful work. Select a manageable set tied to the task rather than treating every candidate below as mandatory.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Throughput and time: completed tasks, resolved cases, accepted deliverables, or cycle time.
  • Quality: accuracy, first-pass acceptance, error rates, escalations, rework, and defect severity.
  • Customer or stakeholder outcomes: satisfaction, resolution, adoption of a recommendation, or another downstream result.
  • Workforce effects: workload, worker experience, learning, retention, and how gains are distributed.
  • Relevant guardrails: privacy, security, safety, fairness, reliability, and rates of human review or override.

NIST’s AI measurement and evaluation guidance emphasizes context-specific metrics, documented methods, benchmark comparisons, uncertainty, and monitoring in production. Which measures matter depends on the task and potential harms.

Look beyond the team average

Break results out by task type and, where sample size and privacy permit, by relevant worker groups such as experience or skill level. An overall average can hide who benefits, who needs more support, or whether work has shifted to another role.

For example, a study of 5,179 customer-support agents reported a 14% average increase in issues resolved per hour, with a 34% increase for novice and lower-skilled agents and minimal impact for experienced and highly skilled agents. Those are findings from that company, tool, task, and study period—not a forecast for every support team. The paper was issued as NBER Working Paper 31161 in 2023 and published in the Quarterly Journal of Economics in 2025.

Interpret results in the context of the evidence

AI’s effects vary with the work, the people doing it, the system, and the outcome being measured. Different study settings illustrate why a single presumed “AI productivity effect” is a poor benchmark:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Customer support: In the 5,179-agent study, the reported average increase in issues resolved per hour was 14%, with larger gains among novice and lower-skilled agents. It measured a particular support setting. NBER Working Paper 31161.
  • Knowledge work across firms: In a randomized six-month experiment involving 66 firms and 7,137 workers, frequent users in the second half spent two fewer hours per week on email. Researchers did not detect a shift in task quantity or composition resulting from individual access. NBER Working Paper 33795, revised November 2025.
  • Product innovation teamwork: In a preregistered field experiment with 776 professionals at Procter & Gamble, individuals working with AI matched the performance of teams without AI on real product-innovation challenges. This finding concerns that specific creative collaboration setting. NBER Working Paper 33641.
  • Broader labor outcomes: A Denmark study estimated no effects larger than 2% on earnings or recorded hours two years after ChatGPT’s launch, while documenting task reorganization and occupational transitions. Aggregate labor statistics do not settle whether AI improved a particular team’s performance. NBER Working Paper 33777, revised March 2026.

These findings measure different outcomes and levels—from individual task performance to labor-market indicators—so they are not interchangeable. Use them as context, not substitutes for measuring the team and work in question.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Keep measuring after deployment

A pilot may not capture learning, workflow changes, or later shifts in how the system behaves. Repeat the core measures in production and compare live results with the baseline and expectations. Monitor changes in task mix, quality, use, human overrides, feedback, and incidents; investigate unexpected shifts rather than relying on the initial result.

Before rollout, decide what evidence would prompt an adjustment, additional review, or rollback. NIST’s AI RMF Measure function states that “AI systems should be tested before their deployment and regularly while in operation.” It also calls for documented tests and metrics, uncertainty measures, and production monitoring.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.