Free tools Windows power users keep installed
One-click scans. No signup required.
OpenAI’s 2025 GDPval evaluation tested whether AI models could produce work products for specific, bounded tasks—not whether ChatGPT can replace entire jobs. The benchmark covered 44 selected occupations and 1,320 tasks, with a 220-task open set. OpenAI reported that frontier models could complete those tasks faster and at lower model-inference cost than experts, but the comparison leaves out workplace review, revision and integration.
What OpenAI actually released
GDPval is OpenAI’s benchmark for assessing how AI models perform on economically valuable work tasks. Its first version covers 44 occupations across nine U.S. industries and 1,320 specialized tasks. A 220-task “gold set” is open-sourced. Tasks are based on real work products or comparable constructed deliverables, and can require documents, slides, diagrams, spreadsheets or multimedia, often with reference files and context. OpenAI’s GDPval announcement describes the benchmark and its methodology.
Examples of the deliverables include a legal brief, engineering blueprint, customer-support conversation or nursing care plan. The benchmark’s unit is a task and its output—not an occupation’s full range of duties.
Which occupations are included
OpenAI selected occupations using 2024 U.S. Bureau of Labor Statistics wage and employment data and O*NET task classifications. It chose five occupations per industry based on wage and compensation contribution, then focused on occupations where at least 60% of tasks were classified as not requiring physical work or manual labor. The nine industries contribute more than 5% of U.S. GDP each, according to OpenAI.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problems#1 Best Overall
- Real estate and rental/leasing: concierges; property, real estate and community association managers; real estate sales agents; real estate brokers; counter and rental clerks.
- Government: recreation workers; compliance officers; first-line supervisors of police and detectives; administrative services managers; child, family and school social workers.
- Manufacturing: mechanical engineers; industrial engineers; buyers and purchasing agents; shipping, receiving and inventory clerks; first-line supervisors of production and operating workers.
- Professional, scientific and technical services: software developers; lawyers; accountants and auditors; computer and information systems managers; project management specialists.
- Health care and social assistance: registered nurses; nurse practitioners; medical and health services managers; first-line supervisors of office and administrative support workers; medical secretaries and administrative assistants.
- Finance and insurance: customer service representatives; financial and investment analysts; financial managers; personal financial advisors; securities, commodities and financial services sales agents.
- Retail trade: pharmacists; first-line supervisors of retail sales workers; general and operations managers; private detectives and investigators.
- Wholesale trade: sales managers; order clerks; first-line supervisors of non-retail sales workers; wholesale and manufacturing sales representatives for technical or scientific products and for other products.
- Information: audio and video technicians; producers and directors; news analysts, reporters and journalists; film and video editors; editors.
This is a deliberately selected set of knowledge-work occupations, not a representative census of all jobs. The selection criteria favor tasks that can be expressed as work products and exclude much work requiring physical activity or manual labor.
What the benchmark results say
OpenAI says expert graders blindly compared AI-generated deliverables with human-produced work across the 220 gold-set tasks. The company reported Claude Opus 4.1 as the top overall performer in that set, while GPT-5 was especially strong on accuracy. OpenAI also reported that performance more than doubled from GPT-4o to GPT-5.
Those are company-reported results for particular model versions and benchmark tasks. They are not a finding that ChatGPT is best at every occupation or that an employer can remove a worker without affecting the work.
Rank #2
OpenAI also estimated that frontier models completed GDPval tasks roughly 100 times faster and 100 times cheaper than industry experts. The speed figure refers to model inference time, and the cost figure to API billing rates. Both exclude workplace oversight, iteration and integration, so they should not be read as estimates of end-to-end business savings or the cost of replacing an employee.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchOpenAI describes its early finding this way: “Early GDPval results show that models can already take on some repetitive, well-specified tasks faster and at lower cost than experts.” The qualifier matters: repetitive and well-specified tasks are not the same as all the responsibilities attached to a job.
Examples of tasks—not proof of job replacement
Futurism’s September 30, 2025 coverage of GDPval cited examples including a financial analyst creating a competitor landscape for last-mile delivery, a registered nurse assessing skin-lesion images, and a real estate agent designing a sales brochure. These illustrate the benchmark’s narrow, deliverable-based approach: each is a particular piece of work that can be assessed, not a demonstration that a model can independently perform the whole occupation. Futurism’s report gives those examples.
Rank #3
A real job also involves deciding what needs doing, gathering missing context, communicating with people, handling exceptions and taking responsibility for outcomes. OpenAI itself notes: “However, most jobs are more than just a collection of tasks that can be written down.”
What GDPval does not test
GDPval is a one-shot evaluation. It does not measure a model building context over time, improving work through multiple drafts, responding to ambiguous requests or deciding which work product is appropriate for a particular client or situation. It also does not include the human review, iteration and systems integration required to use model output in a workplace.
Recommended Free Tools
That gap is important when interpreting the cost and speed figures. A fast first draft may still need expert checking, correction or substantial rework. The benchmark can help compare model output on defined tasks, but it does not establish the total time, cost, quality or risk of completing those tasks in a real organization.
Nor does the evaluation establish the net effect on employment. It examines selected tasks in selected occupations; it does not measure how employers will reorganize work, how demand may change, or how many positions may be created or eliminated.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How to interpret “Can AI do your job?”
A useful reading of the headline claim is that AI models can produce plausible deliverables for some bounded knowledge-work tasks. The evidence does not support the stronger claim that ChatGPT can already replace 44 occupations. Whether a model can assist with one task is separate from whether it can reliably perform the broader role, meet professional obligations or operate without human supervision.
GDPval is also a snapshot of specified model versions against a specified task set, not a live guarantee of what ChatGPT can do today. OpenAI’s release notes are continuously updated; current features, model access and availability can change. Check ChatGPT release notes for current product details rather than assuming that a 2025 benchmark result describes the current service.
Best Value
How to judge AI-work benchmarks
GDPval is most informative when read as a structured evaluation, not as a forecast of job displacement. When comparing it with another AI-work benchmark, look at:
- Whether tasks resemble real work and what deliverables are evaluated.
- Which occupations and industries are covered, and how they were selected.
- Who created the tasks and reference outputs, and how experts were selected.
- Whether grading is blinded and independent.
- Whether the test is one-shot or measures an interactive workflow with revisions.
- Which model versions were tested and when.
- Whether reported time and cost include human review, iteration and integration.
OpenAI’s published results describe its own benchmark; the cited material does not establish an independent replication of those comparisons.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

