AI Evaluation

OpenAI GDPval: The Evaluation of AI's Economic Potential

This post explores the methodology behind GDPval, its key findings, and what they might signal for the future of knowledge work. And the picture GDPval paints is far more interesting than any exam score.

Vishnu Vardhan Sai Lanka
Beginner
October 14, 2025
10 min read

For years, the story of AI progress has been told through a series of escalating challenges. First, we measured it with games. Chess, then Go. These were good milestones, but they taught us a narrow lesson. They showed us that machines could master complex, but closed, systems with perfect information and clear rules. Then, we moved on to academic tests. SATs, AP exams, the Bar. This seemed like a step up. It showed AIs could absorb and reason over vast amounts of human text. But this, too, was a trap. Standardized tests are, by definition, standardized. They are clean distillations of messy reality, designed for objective grading. Real work isn't like that. Real work is messy, ambiguous, and subjective. It’s not about finding the one correct answer. It’s about creating a deliverable that someone else, a client, a boss, a customer finds valuable. It’s about taste, aesthetics, and understanding unstated context.

The Anatomy of a GDPval Task

The core innovation of GDPval lies in its process built on deep human expertise from start to finish. Rather than having researchers invent tasks, the GDPval team sourced them directly from the real world. Let's walk through how a single data point is constructed.

Step 1: Selecting the Arena

The process begins at a macroeconomic level. The researchers first identified the top 9 sectors contributing to U.S. GDP, such as "Finance and Insurance" and "Professional, Scientific, and Technical Services". Within each sector, they selected the five highest-earning "knowledge work" occupations, jobs that are predominantly digital. An occupation was deemed "digital" if a GPT-4o model classified at least 60% of its official O*NET task descriptions as being computer-based. This led to a set of 44 occupations, from Accountants and Lawyers to Software Developers and Financial Analysts.

Step 2: Recruiting the Experts

For each of the 44 occupations, the team recruited seasoned industry professionals. The bar was set remarkably high: experts were required to have a minimum of four years of experience and a strong professional history, with the average participant having 14 years of experience in their field. These weren't junior employees; they were seasoned experts from firms like Goldman Sachs, Google, Disney, and major law firms.

Step 3: Creating the Task

An expert, say a Financial Analyst, is then prompted to create a task based on their actual work. An example from the paper is: "Create a competitor landscape for last mile delivery". This involves three components:

  • The Request: A detailed prompt that mirrors a real-world assignment.
  • Reference Files: The raw materials needed to complete the task, such as spreadsheets, reports, or internal memos. A typical GDPval task has around 2 reference files, but some have as many as 38.
  • The "Gold" Deliverable: The expert then completes the task themselves, producing a high-quality deliverable (e.g., a PowerPoint presentation) that serves as the human baseline. They also log the time it took, which averages around 7-9 hours for tasks in the dataset.

Step 4: Ensuring Quality and Representativeness

Multi-stage review process flowchart

Every task undergoes a rigorous, multi-stage review pipeline involving an average of five human reviews. This includes:

  • A generalist review to ensure it meets basic project guidelines.
  • An occupation-specific expert review to confirm the task is realistic, representative of the profession, and completable with the given materials.
  • A final iterative review loop with a lead reviewer to polish the task to a high standard.

This meticulous process ensures the final task is not an academic exercise but a genuine slice of professional work.

Step 5: The Blind Evaluation

Finally, the request and reference files are given to various AI models (e.g., GPT-5, Claude Opus 4.1). The resulting AI-generated deliverable and the original human-created deliverable are then presented, unlabeled, to a new set of occupational experts for a blind, pairwise comparison. The expert grader's job is simply to answer: "Which is better?" They consider not just correctness, but also subjective factors like structure, style, and aesthetics. The primary metric is the win rate of the model against the human expert.

Key Findings: A Picture of an Emerging Workforce

The results from GDPval provide a nuanced snapshot of where frontier AI stands today.

1. Approaching, but Not Exceeding, Human Parity

The top-performing model, Claude Opus 4.1, achieved a 47.6% win-or-tie rate against industry experts. This means in nearly half of the tasks, a seasoned professional preferred or saw no difference in quality between the AI's work and that of their peers. While this is an impressive feat, it's crucial to note that "parity" isn't the ceiling.

GDPval pairwise expert preferences bar chart

2. A Linear Trajectory of Improvement

When plotting the performance of OpenAI's frontier models over their release dates, a clear trend emerges: a roughly linear rate of improvement. The line doesn't appear to be flattening as it approaches the 50% "parity" mark, suggesting that current progress is not hitting a wall.

OpenAI frontier model performance over time

3. Different Models, Different Strengths

The evaluation revealed distinct "personalities" among the top models. GPT-5 high was found to excel in accuracy, such as carefully following instructions and performing correct calculations. In contrast, Claude Opus 4.1 excelled in aesthetics, like document formatting and slide layouts. This is corroborated by their failure modes. Experts most often rejected deliverables from models like Gemini and Grok for instruction-following failures. GPT-5 high had the fewest instruction-following issues but lost points on formatting, whereas Claude's weaknesses were more often related to accuracy. This suggests that choosing the "best" AI might soon become a matter of picking the right specialist for the task at hand.

Failure Modes bar chart comparing models

The Future Outlook: Speculations and Open Questions

While the findings above are grounded in the paper's data, they invite speculation about the future trajectory of AI in the workplace.

The Bottleneck is Shifting from Execution to Management

The paper includes a fascinating analysis of potential speed and cost improvements. While a model can generate a deliverable in minutes that takes a human hours (a "naive" speedup of over 90x for GPT-5), the real-world savings are tempered by the need for human review and the probability of failure. In a realistic workflow where an expert reviews the AI's output and re-does the work if it's unsatisfactory, the effective speedup for GPT-5 drops to a much more modest 1.4x.

This suggests a fundamental shift in the nature of knowledge work. The primary bottleneck may no longer be the time it takes to execute a task, but the time it takes to verify its quality and manage the AI agent. The most valuable skill in this new paradigm might not be raw execution, but the ability to write precise prompts, critically evaluate outputs, and effectively guide AI systems. The paper demonstrates this by showing a 5 percentage point increase in GPT-5's win rate just from improved prompting and scaffolding.

The "Context Gap" Remains

Perhaps the most telling experiment is one buried in the appendix, where researchers created an "under-contextualized" version of GDPval by shortening the prompts and removing explicit instructions. The model had to infer more context. The result? Performance dropped. This highlights what may be the final frontier for AI capability: navigating ambiguity. A senior professional's value often lies not in executing a perfectly defined brief, but in taking a vague goal and figuring out the necessary steps, inputs, and constraints. Today's models are brilliant junior assistants who need a clear set of instructions. The ability to independently close this "context gap" is what separates a good assistant from a true colleague.

AI's Performance by Occupation

By analyzing the occupation-specific win rates in GDPval, we can move beyond broad generalizations and identify the specific characteristics of professions most susceptible to AI disruption. The chart reveals a landscape of stark contrasts, where the underlying nature of a job's core tasks, rather than its prestige or complexity, determines its vulnerability.

Obvious Insights: The Highs and Lows

A direct reading of the chart shows clear performance disparities between professions:

  • High-Susceptibility Occupations: Roles like Software Developers show remarkably high win-or-tie rates, with top models approaching 70%. Similarly, occupations focused on structured data and process management, such as Financial & Investment Analysts, Order Clerks, and Sales Managers, also demonstrate strong performance for AI models. These are roles where the AI often performs at or above the level of an experienced human.
  • Low-Susceptibility Occupations: Conversely, AI models struggle significantly in professions that demand deep subjective judgment and creative generation. Lawyers, Film & Video Editors, and Producers & Directors all exhibit low win rates across the board, often falling well below the 50% parity mark. Engineering roles like Mechanical and Industrial Engineers also show lower AI performance.

Non-Obvious Insights: The Principles of Disruption

The variance between occupations like Software Developers and Lawyers, both highly skilled roles in the same economic sector, reveals deeper principles about what makes a job's tasks amenable to AI:

  1. Structured Logic vs. Ambiguous Interpretation: The core reason Software Developers have such a high win rate is that their primary medium—code—is a formal language with defined rules and syntax. The task is often to solve a logical problem within a structured system. Law, on the other hand, operates on ambiguous natural language. The most valuable legal work involves navigating subjectivity, crafting persuasive arguments, and interpreting intent, which are skills that current models find difficult to master.
  2. Information Synthesis vs. Novel Creation: Many of the high-performing roles involve synthesizing existing information. A financial analyst processes data to create a competitive landscape; an order clerk audits documents to find inconsistencies. AI excels at this rapid processing and restructuring of known data. In contrast, roles like Film Editor or Producer are centered on novel creation—building a narrative, evoking emotion, and making aesthetic judgments that don't have a single "correct" answer based on the inputs.
  3. Digitally Native vs. Physically Grounded Work: Even among computer-based tasks, there's a crucial distinction. The work of a software developer is purely digital. However, a Mechanical Engineer designing a 3D model of a machine part is performing a digital task that is ultimately grounded in the physics and constraints of the real world. This requires a form of spatial and physical intuition that is less developed in current models compared to their purely linguistic or logical reasoning abilities, likely contributing to lower win rates in engineering fields.

Conclusion

In conclusion, a profession's susceptibility to AI is not determined by its salary or required education level, but by the fundamental nature of its work. The more a job relies on formal logic, structured data, and the synthesis of existing information, the more likely it is that AI will become a powerful collaborator or competitor. The domains that remain most defensible are those built on ambiguity, subjective creativity, and a deep connection to the physical world. GDPval is a landmark study because it provides a map of this new territory. It shows us that the path to truly transformative AI is not just about building more powerful models, but about instilling them with the reliability, aesthetic sense, and contextual awareness that define expert human work.

Related Articles