Gdpval
Free
OpenAI's open source AI model economic value assessment framework covers 1,320 real economic tasks in 44 occupations in 9 major industries, and is used to measure the performance of AI in real work scenarios. Quantify the potential impact of AI on labor productivity and GDP with granular task-level assessments.
Gdpval: AI model economic value assessment framework
Core parameters and statistics
Gdpval (GDP Value Assessment of AI Models) is an AI model economic value assessment framework launched by OpenAI. Its goal is not to measure the model's score on academic benchmarks (such as MMLU, HumanEval), but to measure how much human labor can be replaced or enhanced by AI in real economic production activities. This is an important shift in the AI evaluation paradigm from "academic accuracy" to "economic impact."
| Project | Specifications |
|---|---|
| Model/API name | Gdpval (GDP Value Assessment Framework) |
| Product Type | AI Model Economic Value Assessment Framework |
| Delivery form | Open source evaluation benchmark + Web visual analysis platform |
| Assessment Dimensions | 9 major industries, 44 occupations, 1320 real economic tasks |
| Open source subset | 220 representative tasks (OpenAI GDPval Benchmark) |
| Task granularity | task-level, occupation-level |
| Data source | O*NET database, U.S. Bureau of Labor Statistics (BLS) |
| Evaluation method | LLM-as-judge + manual review double verification |
| Open License | MIT License (Open Source Subset) |
| Support Platform | Web (gdpval.com) |
Interpretation of core parameters: The evaluation granularity of Gdpval is "task" rather than "occupation". A profession (such as "software engineer") consists of dozens of specific tasks (such as "writing unit tests", "code review", "debugging production problems"). Gdpval evaluates the degree to which each task can be automated on a case-by-case basis and then aggregates it into occupation-level and industry-level economic impact inferences. The advantage of this fine-grained approach is that it does not assume that "AI will replace entire professions," but rather identifies precisely which specific tasks can be augmented or replaced by AI and which tasks still require human dominance.
Economic value quantification link: Task-level AI capability score → Occupational-level automation potential → Industry-level labor substitution/augmentation rate → Macro GDP impact deduction. Each layer is calibrated by an economic model rather than a simple sum.
User and market recognition
Since its release in September 2025, Gdpval has had a significant impact in the three intersection areas of AI policy research, labor economics, and AI safety.
Academic Citations and Discussions: The task-level assessment method proposed by Gdpval has been cited by economics and computer science departments in many universities. Its core methodological paper (in collaboration with OpenResearch and Wharton School) is used to discuss the structural impact of AI on labor markets. The problem that the framework solves is that although traditional AI benchmarks (such as MMLU, BIG-bench) can measure the breadth of knowledge of the model, they cannot answer the economic question of "how many human hours can this model replace?" Gdpval establishes a quantitative link from model performance to economic impact by mapping AI capabilities to the O*NET occupational task classification system.
Policy Level Attention: The release of Gdpval coincides with a window period for governments around the world to accelerate the formulation of AI policies. The "occupational automation potential" quantitative indicator provided by the framework is cited by multiple think tanks and policy research institutions to evaluate the impact of AI on employment in different industries. Task-level rather than occupation-level assessment results allow policymakers to identify "which specific work activities are most likely to be changed by AI" rather than discussing generally "which industries will be replaced by AI."
Industry analyst evaluation: AI industry analysts generally believe that Gdpval fills a key gap in the AI evaluation system - the gap from "what the model can do" to "how much the model is worth." Unlike traditional model rankings (such as LMSYS Chatbot Arena, Artificial Analysis) that focus on user experience and inference speed, Gdpval provides a value scale under the economics framework.
Open Source Community Feedback: A subset of 220 tasks publicly available on GitHub are used by researchers for replication and extended evaluation. Community feedback focuses on two directions: one is to expand the assessment task to non-US labor markets (such as European and Asian occupational classification systems), and the other is to connect the assessment method with more models (including open source models).
Clear boundaries of understanding are required: What Gdpval provides is not a prediction of "what will happen to AI in the future", but a snapshot assessment of "how many specific tasks in existing jobs can AI complete under the current technical level?" The actual economic impact also depends on the speed of technology adoption, the ability of organizations to change, policies and regulations, and labor market adaptations. The framework’s economic impact derivation tacitly assumes that technical bottlenecks are no longer a constraint – an assumption that needs to be verified in reality over time.
Cost advantage
The "cost" of Gdpval does not refer to the monetary cost of using the framework (it is completely free), but to the economic cost of traditional AI evaluation systems - the cost of high-risk decisions resulting from the lack of economic dimensions of evaluation.
| Cost Dimension | Traditional AI Valuation (without Gdpval) | Valuation using Gdpval |
|---|---|---|
| Evaluation indicators | Academic benchmark accuracy (MMLU, GSM8K) | Economic task completion rate + man-hour replacement rate |
| Basis for decision making | "Model A is 2% more efficient than model B" | "Model A can automate 37% of the task hours in this industry" |
| Investment return inference | Fuzzy qualitative judgment | Quantitative deduction (such as "Save X billion US dollars/year") |
| Policy-making inputs | Expert opinions + general forecasts | Task-level data + economic model calibration |
| Risk assessment granularity | Occupation level ("XX occupation will be replaced") | Task level ("YY tasks in XX occupation can be automated") |
| Cross-industry comparability | Low (each benchmark is independent) | High (unified economic value scale) |
Hidden Cost Savings: For a CTO who is evaluating "whether to deploy AI assistants across the company," traditional evaluation methods require a combination of multiple benchmark results, PoC reports, and vendor commitments to make a decision, and the evaluation cycle is usually 3-6 months. The industry-level and task-level evaluation data provided by Gdpval can shorten this evaluation cycle to 1-2 months, reducing the decision-making risk from "relying on supplier propaganda" to "based on economic models and task-level data."
Cost Dimensions for Policymakers: For government agencies, Gdpval provides a standardized AI economic impact assessment tool. Traditionally, assessing the impact of AI on the labor market has required commissioning a custom study from a consulting firm (typically $50K-$200K+ per report). Gdpval’s public framework enables internal teams to directly use standardized methods for initial assessments, reducing the cost of entry to close to zero.
The True Cost of Open Source: The evaluation framework and open source subset of Gdpval are completely free (MIT license), but the evaluation data for the full version of 1320 tasks is only available through the gdpval.com platform. For research requiring complete career coverage, platform access is a necessary path. The platform itself is free, but if batch API calls are required for custom evaluations, the cost of the API calls needs to be considered.
Main functions
The function of Gdpval is designed around the core proposition of "mapping AI model capabilities to economic value", and each function point follows the causal logic of "mechanism → effect → scenario".
-
Task-level economic value evaluation (core mechanism): Gdpval does not evaluate "how smart the model is", but evaluates "how many specific tasks in real work scenarios the model can complete". Mechanism: 1,320 tasks from 44 occupations in the O*NET database are brought into the AI model one by one, and LLM-as-judge and human reviewers double-judge "whether AI can complete this task under the current technical level." Effect: Output the automatability probability (0-100%) of each task, instead of a simple "can/cannot" binary judgment. Scenario: Economists can use these probabilities to infer the rate of labor-hour replacement by AI for specific occupations, and policymakers can use it to identify occupational groups that need focus.
-
Industry Aggregation Analysis: Mechanism: Task-level assessment results are aggregated layer by layer according to O*NET's industry classification system (9 major industries), taking into account the task's work-hour weight within the occupation and U.S. Bureau of Labor Statistics employment data. Effectiveness: Generates an industry-level panorama of AI's economic impact - e.g. "Approximately 42% of task hours in manufacturing can be automated with current AI capabilities, focusing on quality inspection, data entry and report generation." Scenario: Investment institutions use industry aggregate data to evaluate the potential efficiency impact of AI on listed companies in different industries.
-
Cross-model capability comparison: Mechanism: Run multiple models (GPT-4o, Claude, Gemini, etc.) on the same set of 1320 task sets, and output the differences in automation potential of each model in various industries and occupations. Effectiveness: Generates capability radar plots and industry coverage differences across models - e.g. "Model A outperforms Model B by 12% on medical paperwork tasks, but Model B is ahead by 8% on programming tasks". Scenario: When selecting models, enterprises can select the optimal model based on their own industry characteristics instead of relying on general rankings.
-
Economic Impact Quantification Model: Mechanism: Input the task-level assessment results into the economic model, and combine it with macro data such as labor productivity, wage levels, and industry employment distribution to estimate the potential contribution of AI to GDP. Effect: Output a quantified GDP impact range (e.g. "Optimistic estimate: AI can increase US annual GDP by 0.5-1.5 percentage points in 5 years"). Scenario: Government agencies and international organizations use these inferences for AI policy design and industry planning.
-
Visual analysis panel: Mechanism: gdpval.com provides an interactive dashboard where users can filter and drill into assessment results by industry, occupation, task type and other dimensions, and supports custom views and data export. Effect: Users can quickly locate "which specific tasks in their industry/occupation are most likely to be changed by AI" instead of reading long reports. Scenario: Individual workers can understand what proportion of their work content may be affected by AI, so as to develop skills improvement plans.
Model and version evolution
The version evolution of Gdpval reflects the transformation path from academic research to engineering platform.
| Version | Date | Key Changes |
|---|---|---|
| Research Preview (0.9) | ~2025-09 | Early research preview, 220 open source task subsets released, methodology papers made public. In the community verification stage, the core contribution is to establish the academic foundation of task-level evaluation methodology. |
| Public beta version (1.0) | ~2026-07 | The full version has 1,320 tasks covering 44 occupations and 9 major industries. The gdpval.com visualization platform is online, and the cross-model comparison function is released. Upgrade from "research tool" to "industry standard assessment platform". |
Version Evolution Features: The jump in Gdpval from 0.9 to 1.0 reflects two key changes. First, task coverage expanded from 220 to 1,320, occupation coverage expanded from about 10 to 44, and industry coverage expanded from 4 to 9—this changes the assessment results from "sample level" to "industry level." Second, from "static paper + data set" to "dynamic visualization platform", the launch of gdpval.com allows non-academic users (policy makers, corporate decision-makers, individual workers) to directly obtain evaluation results without having to read research papers.
Possible follow-up directions: Based on the extensibility design of the framework, subsequent versions may support: adaptation of the occupational classification system for non-US labor markets, more fine-grained task decomposition (extending the current task level to the sub-task level), dynamically updated capability assessment (updated in real time with the advancement of AI technology instead of static snapshots), and industry-customized economic impact models.
Technical advantages
The technical advantage of Gdpval does not lie in the performance of a single model, but in the systematic nature of the economic evaluation methodology - it solves the interdisciplinary problem of "how to translate AI technical capabilities into economic value in a repeatable and verifiable method."
Design logic of task classification system: Gdpval builds a task classification system based on ONET (American Occupational Information Network Database). ONET is the authoritative occupational database maintained by the U.S. Department of Labor and provides a detailed list of tasks, required skills, knowledge, abilities and work activities for each occupation. The advantage of Gdpval is that it does not create a classification system from scratch, but borrows an existing framework that has been iteratively validated for 30 years. This means Gdpval’s assessment results can be directly interfaced with wage data, employment data and skills requirements data in O*NET, automatically gaining context for economic impact.
Three-tier verification architecture for evaluation methods:
Task input (O*NET description)
↓
The first layer: AI model execution → output results
↓
Second level: LLM-as-judge evaluation → output automatable probability
↓
The third level: manual auditor sampling review → calibration and deviation correction
↓
Output: calibrated task-level automated scoring
The first layer consists of candidate AI models performing specified tasks and producing evaluable outputs. The second layer uses an independent evaluation model (LLM-as-judge) to score the output quality, combined with structured evaluation criteria (accuracy, completeness, security and other dimensions). The third layer samples 10-20% of the overall results for review by human reviewers to calibrate LLM-as-judge bias. This double verification mechanism of "AI evaluation AI + manual calibration" minimizes the false positive rate of single-point evaluation.
Economic impact quantification model: Gdpval's economic impact calculation is not a simple "task automatability rate × salary", but considers three adjustment factors: (1) Technical feasibility - Even if AI can complete the task under laboratory conditions, actual deployment still needs to consider integration costs, regulatory obstacles and user acceptance, so the "technology adoption delay" factor is introduced; (2) Task complementarity - Even if some tasks can be AI Completed alone, but tightly coupled with other human tasks in the workflow, complete automation will reduce efficiency; (3) Labor market dynamics - The working hours replaced by AI will be partially absorbed by newly created tasks (the "compensation effect" in economics), rather than a simple one-to-one reduction. The introduction of these factors makes Gdpval's economic deduction closer to reality than the "direct aggregation" method.
Open & Reproducible: An open source subset of 220 tasks is licensed under the MIT license, allowing free reproduction and extension by academia and industry. The code of the evaluation framework is also open source (GitHub), allowing researchers to examine the implementation details of each evaluation step. This kind of transparency is especially important in the field of AI evaluation—because the evaluation method itself (rather than just the evaluation results) should be able to withstand scrutiny by academic peers.
Adaptation boundaries and restrictions
Gdpval's assessment framework has significant value at both corporate and policy levels, but has a clear scope of applicability and known limitations.
Recommended usage scenarios:
- Enterprise Technology Strategic Planning: When enterprises need to evaluate "the actual economic value of AI in their own industries," the industry-level data provided by Gdpval can be used as strategic decision-making input.
- Policy Research and Labor Market Analysis: Gdpval’s methodology can serve as an analytical framework for government agencies and think tanks when assessing the impact of AI on employment and designing response policies.
- Cross-model selection decision: When industry suitability comparisons need to be made between multiple AI models, Gdpval's task-level comparison data has more reference value than the universal benchmark (MMLU).
- Academic Research and Economic Modeling: Academic teams studying the relationship between AI and labor productivity can use Gdpval's evaluation results as empirical data input.
Not recommended scenario:
- Real-time/automated decision-making: Gdpval is an analysis framework rather than an API service and is not suitable for scenarios that require real-time AI capability scoring.
- Individual Career Counseling: The output of the framework is industry-level and group-level statistical results, which is not suitable for providing individuals with personalized advice on "whether your job will be replaced by AI."
- Technical R&D Guidance: Gdpval evaluates the capability boundaries of existing AI models and does not provide model improvement directions or technical route suggestions.
Known limitations:
- US Labor Market Bias: The framework is based on O*NET (US Occupational Classification), and task descriptions and weight distributions reflect the structure of the US labor market. Differences in occupational distribution and task composition need to be taken into account when directly applying to other economies.
- Static snapshot nature: The evaluation results reflect the capability boundaries "under the current technical level". AI technology is developing rapidly, and the timeliness window for evaluating results may be only 6-12 months.
- Task Granularity Cap: Although Gdpval claims that task-level evaluation is superior to career-level, "task" itself in O*NET is also an aggregate concept. A "task" may contain multiple subtasks, and the degree to which they can be automated may vary significantly.
- Unable to measure implicit value: The framework quantifies the evaluable explicit task output and cannot measure the value of AI in implicit dimensions such as "innovation, leadership, teamwork, and interpersonal trust".
How to use
Gdpval offers two main paths to use: a web visualization platform and an open source benchmark dataset.
| Entrance | How to use |
|---|---|
| Web visualization platform | Visit gdpval.com → Select industry/occupation dimension → View evaluation results and comparative analysis |
| Open source data set | GitHub download 220 task subsets → Run local evaluation → Reproduce study results |
| API access | gdpval.com Obtain API Key after registration → Call task evaluation interface on demand |
Web platform usage process:
- Visit gdpval.com to browse aggregated industry- and career-level assessment results without registration.
- Select the industry (such as "information technology") or occupation (such as "software engineer") of interest and view the automatable probability distribution of 40+ specific tasks under this occupation.
- Use the "Cross-Model Comparison" function to select 2-4 AI models and view the differences in automation potential of each model in each occupation/industry side by side on the same page.
- Export assessment results (CSV/PDF) for internal reporting or further analysis.
Open source data set usage (taking Python as an example):
# Download and load the GDPval open source evaluation subset
import pandas as pd
from datasets import load_dataset
dataset = load_dataset("openai/gdpval", split="test")
print(f"Total number of tasks: {len(dataset)}")
#dataset contains: task_id, occupation, industry, task_description,
# ai_capability_score, human_benchmark
# Aggregate analysis by industry
df = dataset.to_pandas()
industry_summary = df.groupby("industry")["ai_capability_score"].agg(["mean", "std"])
print(industry_summary)
Researcher Custom Assessment: For researchers who need to use their own models or custom tasks, they can fork the GitHub repository and follow the steps below:
- Prepare the evaluation task list (JSON format, refer to the structure of the open source subset)
- Configure the model API endpoint to be evaluated
- Run the evaluation script to generate original results
- Use the built-in LLM-as-judge module for automated scoring
- Perform manual sampling calibration on the results
Product Pricing
Gdpval’s core evaluation framework and open source datasets are completely free. The pricing structure at the platform level is as follows:
| Billing items | Price |
|---|---|
| Open source benchmark dataset (220 tasks) | Free (MIT license) |
| Web visualization platform (basic browsing) | Free |
| Industry/occupation level aggregated report export | Free |
| Cross-model comparative analysis | Free |
| Customized evaluation API (batch call) | Billing based on usage (specifically, please refer to the official website) |
| Enterprise-level customized assessment project | Business quotation |
Free truth: The open source subset of Gdpval and the basic functions of the Web platform are completely free, covering the needs of most users (browsing evaluation results, cross-model comparison, exporting reports). The Custom Assessment API is intended for researchers and enterprise users who need to run assessments on their own sets of tasks, which may incur API call charges. For most users who only need to refer to the assessment results (policy makers, corporate decision-makers, individual workers), the full value can be obtained at zero cost.
Application scenarios
-
Scenario 1: Enterprise Technology Strategic Planning - The CTO of a manufacturing company needs to evaluate "how much labor costs can be saved in factory operations by introducing AI assistants in the next 3 years." Using Gdpval, select the Manufacturing industry to view data on the potential for task automation for each occupation in it (quality inspector, production planner, equipment maintenance technician, etc.). Integrate the company's job distribution and salary data to calculate the financial impact of automatable hours. Input→Processing→Output link: Enterprise position data + Gdpval industry assessment → Calculation of working hour replacement rate → Quantitative deduction of cost savings. Verification method: Compare Gdpval's automation potential assessment with the actual efficiency improvement data in the PoC trial, and calibrate the model parameters.
-
Scenario 2: Government AI policy formulation - A government department needs to assess the impact of AI on the local labor market in order to design a career transition training plan. Using Gdpval's industry aggregation panel, identify the occupational groups with the highest potential for automation in employment-intensive industries in the region. Verification method: Cross-validation with local labor market data - whether the occupational groups with the highest automation potential are also groups with slow employment growth/increasing unemployment rates.
-
Scenario Three: AI Model Procurement Decision - The AI platform team of a technology company needs to choose between GPT-4o, Claude Sonnet, and Gemini 2.5 Pro as the base model for the enterprise AI assistant. Use Gdpval's cross-model comparison function to view the differences in task completion rates between these three models in the occupations corresponding to "the company's core business scenarios". Verification method: Select the five most commonly used task types in the enterprise, run PoC tests on the three models, and compare the consistency of the Gdpval evaluation results with the actual PoC results.
-
Scenario 4: Personal career development planning - A data analyst wants to understand the specific impact of AI on his career. Using the Gdpval platform, select the "Data Analyst" occupation to view the specific automatability probability of 40+ tasks under this occupation. If "the probability of automation of data cleaning and preprocessing is > 90%", it means that the future value of this skill has declined, and you should shift to tasks that "are difficult to replace with AI" (such as business insights, stakeholder communication). Verification method: Return to the Gdpval platform regularly (quarterly) to track the scoring trends of each task in this profession.
Applicable people
-
Economists & Policy Researchers: Researchers who need to quantify the economic impact of AI to support policy recommendations. Gdpval provides a standardized analysis framework and ready-made industry-level data, avoiding the high cost of modeling from scratch. Prerequisite: Have basic knowledge of labor economics and understand the difference between automation potential and employment impact.
-
Corporate Strategy and Decision Maker: CTO, CIO, CEO and other executives who need to make decisions on "AI ROI". Gdpval's industry-level assessment data can be used to quantitatively demonstrate internal business cases. Prerequisite: Understand the company's job distribution and work processes, and be able to map Gdpval's industry-level data to the company's specific conditions.
-
AI Developers and Researchers: AI engineers who need to understand the capabilities of models in real-world scenarios. Gdpval's task-level evaluation results better reflect the model's performance in real applications than general benchmarks. Prerequisite: Ability to understand the limitations of evaluation methodologies and statistical biases.
-
Investors and Industry Analysts: Assess the market impact and investment opportunities of AI on specific industries. Gdpval's cross-model comparison and industry aggregate analysis provide quantitative reference. Prerequisite: Be able to distinguish the gap between "technical feasibility" and "actual market adoption rate".
-
Unfit Boundary: Gdpval is not suitable for individual career counseling (giving personalized suggestions on "whether your job will be replaced by AI"), for real-time AI capability evaluation, or as the only guide for technology research and development direction. The evaluation results of the framework are group-level snapshots and do not contain information on individual differences and dynamics over time.
Comparison of competing products
| Comparative Dimensions | Gdpval (OpenAI) | MMLU (Academic Benchmark) | BIG-bench (Academic Benchmark) | HumanEval (Code Benchmark) | LMSYS Chatbot Arena (Crowdsourced Evaluation) |
|---|---|---|---|---|---|
| Assessment Objectives | Economic Value (GDP Impact) | Breadth of Knowledge | Reasoning Skills | Code Generation | User Experience Satisfaction |
| Assessment granularity | Task level (1320 real work tasks) | Question level (57 subjects) | Task level (200+ tasks) | Function level (164 questions) | Conversation level (A/B comparison) |
| Evaluation dimensions | Automated probability + economic impact deduction | Accuracy rate | Accuracy rate/score | pass@k | Winning rate/ELO score |
| Economic relevance | ✅ Direct quantification (working hours replacement rate, GDP contribution) | ❌ None | ❌ None | ❌ None | ❌ None |
| Industry coverage | ✅ 44 occupations in 9 major industries | ❌ Subject classification | ❌ Task classification | ❌ Programming | ❌ No classification |
| Timeliness | Static snapshot (needs to be updated regularly) | Static | Static | Static | Dynamic (continuously updated) |
| Level of open source | Partially open source (220 task subset + evaluation code) | Fully open source | Fully open source | Fully open source | Partially open source |
| Applicable scenarios | Economic analysis, policy formulation, corporate strategy | Academic research, model comparison | Exploration of capability boundaries | Code capability assessment | User preference research |
Decision Suggestion: Gdpval is not a replacement for traditional benchmarks, but a complement from different dimensions. If you need to answer "how much is this model worth" or "what economic impact will AI have on my industry?" Gdpval is the only tool that provides a quantitative answer. If you need to compare the accuracy of models in a specific academic domain (e.g. mathematics, medicine), traditional benchmarks such as MMLU are more appropriate. The ideal evaluation strategy is to use Gdpval for economic value evaluation to determine strategic direction, and use traditional benchmarks for technical detail verification.
Summary and Outlook
Gdpval represents an important extension of the AI evaluation paradigm from "academic accuracy" to "economic impact." It attempts to answer the most fundamental business question in the AI industry—"How much is AI worth?"—by mapping model capabilities to 44 occupations, 1,320 real economic tasks, and using quantitative methods to deduce the potential impact on labor productivity and GDP.
Core Strengths: Gdpval’s differentiator is the systematic nature of the assessment framework—it’s not another benchmark ranking list, but an interdisciplinary tool that connects AI capabilities, labor economics, and public policy. The fine-grained assessment at the task level makes the output conclusion more valuable for action guidance than the general "occupational replacement rate". Open subsetting and evaluation code ensures that the methodology can be reproduced and extended by academia and industry.
Current limitations: (1) U.S. labor market bias - the classification system based on ONET cannot directly reflect the occupational structure and task distribution of other economies; (2) static snapshot nature - AI technology iterations accelerate, and assessment results may be outdated within 6-12 months; (3) upper limit of task granularity - a single task in ONET may contain multiple sub-tasks with different automation potential; (4) Hidden value blind spots - unquantifiable work dimensions such as innovation, leadership, and trust building cannot be measured.
Acquisition/Adoption Risk Assessment: For enterprise users, Gdpval has extremely low adoption risk - the core framework and open source subset are completely free, the web platform is free to use, and there is no risk of vendor lock-in. The biggest risk is not in adopting Gdpval itself, but in making decisions overly relying on the evaluation results of Gdpval and ignoring its limitations. It is recommended to use Gdpval's evaluation data as one of the decision-making inputs, combined with PoC experiments, industry experience and expert judgment. For government agencies, the methodological framework provided by Gdpval can serve as a starting point for autonomous assessments, but it is recommended to incorporate national labor market data for calibration and local adaptation.
Directions for follow-up observation: Whether Gdpval can evolve from an annual static snapshot into a dynamic tracking platform, whether it can be expanded to more countries' occupational classification systems, and whether OpenAI will integrate it into the standard evaluation suite for model releases are three key variables that determine the long-term impact of the framework.
Related tools: hugging-face, replicate
Version Info
- Public beta version :The AI model economic value assessment framework launched by OpenAI in September 2025 measures the performance of AI in real economic tasks. The full version covers 1320 tasks in 44 occupations, and the open source subset contains 220 representative tasks. The framework provides a task classification system, assessment methods and economic impact quantification models.
- Research Preview :An early research preview version, including a subset of 220 open source tasks (OpenAI GDPval Benchmark), for community validation and methodological demonstration.
User Reviews