Recommended Free Tools
Measure AI value against a defined work outcome—not a universal productivity percentage. Establish what the workflow delivers today, compare AI-supported work with a credible alternative, and track quality, costs, adoption, and risk alongside time or throughput. Time saved is potential capacity; it becomes business value only when you can show how that capacity improves an outcome your organization values.
What should you measure to determine AI’s value?
Start by naming one bounded task or workflow, the people doing it, the AI system and version, and the result the organization wants. For example, “drafting first responses to routine support tickets” is measurable; “making customer service more efficient” is too broad to evaluate well.
The right measures depend on the setting. NIST notes that “How a given component is measured and evaluated can change based on the context in which the AI system operates.” Its AI measurement and evaluation guidance emphasizes context-specific evaluation.
Choose a balanced set of measures before rollout:
- Output and time: tasks completed, throughput, cycle time, or time per task.
- Quality: output judged against a stable rubric, plus error and rework rates.
- Relevant outcomes: for example, customer experience, shorter waits, or worker experience, depending on the use case.
- Adoption: who has access, who actually uses the system, and how often. Access or licenses alone do not show that AI is being used.
- Costs and risk: implementation and operating costs, oversight needs, and risks relevant to the workflow, such as accuracy, privacy, security, or bias.
Keep measures specific enough to reveal trade-offs. Faster completion is not an improvement if it also creates more errors, more rework, or a worse customer outcome.
#1 Best Overall
How do you establish a useful baseline?
Record how the workflow performs before AI is introduced, using the same definitions you intend to use afterward. Depending on the task, that baseline may include output per hour, time to completion, quality scores, errors, rework, service outcomes, and worker outcomes.
Document the observation period and the conditions that could affect results: workload mix, seasonality, staffing, process changes, and differences in task difficulty. Without that context, a change after rollout may be caused by something other than AI.
Rank #2
Keep unrelated work separate. Averaging very different workflows can conceal where AI helps, where it has no effect, or where it causes problems.
How should you compare AI-supported work with work without AI?
A before-and-after comparison is useful, but it cannot by itself establish that AI caused the change. When feasible, randomly assign access or use a phased rollout so comparable workers or teams can be compared over the same period. If random assignment is not practical, use a comparison group or a time-series design, and explain its limitations.
Rank #3
- Define the comparison: specify the tasks, users, AI system and version, outcome measures, and observation window.
- Choose a defensible design: use random assignment where feasible; otherwise, select a comparison group or time-series approach that can account for other changes.
- Keep measures consistent: apply the same quality rubric and task definitions before and after, and note changes in workload or process.
- Separate test conditions from ordinary use: a controlled task test can show what a system can do under those conditions; field performance shows how it works in day-to-day operations.
- Report uncertainty and coverage: state which people and tasks were included, what changed, and what the comparison cannot establish.
NIST’s AI RMF Core: Measure function and AI RMF Playbook: Measure describe documenting methods, metrics, uncertainty, limitations, and context-relevant field data. The AI Risk Management Framework is voluntary guidance; NIST says it is being revised, so check its current status when applying it.
What do published workplace AI studies show—and not show?
Studies show that gains are possible, but results differ by task, population, and measurement. They are evidence about their particular settings, not forecasts for a different organization.
| Study | Setting and result | What the result does not establish |
|---|---|---|
| Noy and Zhang, 2023, Science | A preregistered online experiment with 453 college-educated professionals completing incentivized, occupation-specific writing tasks found 40% lower average time and 18% higher output quality for participants using ChatGPT. | Those task results do not establish the same time or quality gains across other roles, workflows, tools, or workplaces. Study details. |
| Brynjolfsson, Li, and Raymond, Generative AI at Work | A study of 5,179 customer-support agents after staggered introduction of a conversational AI assistant reported 14% more issues resolved per hour on average. The paper reports a 34% productivity improvement for novice and lower-skilled workers and minimal impact for experienced and highly skilled workers. The NBER page lists a 2025 published version in the Quarterly Journal of Economics. | The average conceals differences by experience and skill; it is not a guaranteed effect for other teams or tasks. Study details. |
| Dillon, Jaffe, Immorlica, and Stanton, Shifting Work Patterns with Generative AI | A six-month randomized field experiment across 66 firms and 7,137 knowledge workers found that 80% of treated workers who used the tool spent two fewer hours per week on email in the second half of the experiment and reduced work outside regular hours. The researchers did not detect changes in the quantity or composition of tasks from individual-level AI access alone. The NBER page records a November 2025 revision; the AEA page lists the study as forthcoming in American Economic Review: Insights. | Email time reductions and tool use did not, by themselves, show that workers took on more tasks or changed the overall composition of their work. Study details. |
These studies use different tasks, tools, populations, and methods, so their percentages should not be combined into a single expected productivity uplift or ROI.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How do you translate saved time into business value?
Measure what happens to the released capacity. If a task takes less time, find out whether that time leads to more completed work, better quality, shorter waits, reduced overtime, or another outcome the organization values. If none of those changes is demonstrated, report the time reduction as a measured operational effect—not as cash savings or realized ROI.
Best Value
Include costs relevant to the use case, such as implementation, training, integration, operation, and oversight. NIST’s industrial AI evaluation procedure considers baseline risk, installation and operating costs, risks of operating the system, estimated value, and a risk-based investment analysis using business metrics. There is no universal AI value threshold, payback period, or ROI formula established by the sources cited here.
How should you monitor risk and decide whether to continue?
Track risks alongside benefits, using checks suited to the workflow. These may include accuracy and reliability, privacy and security, bias, or other effects on workers and customers. Document metric limitations, collect user feedback, and revisit the measures if the model, process, users, or operating context changes. NIST’s ARIA program states that “ARIA supports three evaluation levels: model testing, red-teaming, and field testing.”
Make the decision report explicit about:
- the measured effect and its uncertainty;
- the people, tasks, and period covered;
- actual adoption and use;
- quality, errors, rework, relevant outcomes, costs, and risks;
- how any released capacity was used; and
- what the evaluation cannot conclude.
That record makes it possible to decide whether to continue, adjust, expand, or stop the use case without mistaking access or faster task completion for proven organizational value.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

