OpenAI and Ironclad built 11 contracting workflow tasks, where GPT‑6 Astra scored 55.0% on average versus GPT‑5.6 Sol's 41.6%
Synopsis
OpenAI partnered with AI contracting platform Ironclad to turn legal, commercial, and procurement contracting workflows into 11 research tasks scored against 8 to 50 criteria each, training models with synthetic tasks and reinforcement learning inside Ironclad-hosted software environments; across the 11 tasks GPT‑6 Astra averaged 55.0% versus 41.6% for GPT‑5.6 Sol, with estimated average time per attempt falling from 37.0 to 19.2 minutes.
Interpretation
Researchers, together with Ironclad employees and people who use Ironclad at OpenAI, identified 11 tasks across legal, commercial, and procurement work, including setting up nondisclosure agreements, creating procurement approval processes, and updating a reusable legal clause to reflect the jurisdiction a requester selects. Unlike generic computer-use evaluations, these tasks were selected by people who know the work and paired with hosted Ironclad product environments where models could practice, bringing evaluation closer to real professional workflows. The number and types of tasks and their provenance are stated; each task was evaluated against 8 to 50 criteria, and the work is estimated to take an experienced user about 30 to 40 minutes per task on average.
GPT‑6 Astra is the first frontier model trained on Ironclad tasks, averaging 55.0% across the 11 research tasks versus 41.6% for GPT‑5.6 Sol, with estimated average time per attempt falling from 37.0 to 19.2 minutes. The report gives both a rubric score and a simulated time dimension, and states the comparison used Astra's Max reasoning and Sol's High reasoning, the settings where each model scored highest. The comparison rests on mean rubric scores and estimated average time across the 11 research tasks; a single-task example is also given, with Astra meeting about 94% of criteria in an estimated 20 minutes and Sol about 85% in an estimated 32 minutes.
An internal model used in the development of Astra reached 63.7% on these tasks, above Astra's 55.0%, and the team says it aims to bring these further gains to future models. This number indicates headroom within the training pipeline itself, beyond the gap between the released model and the previous generation. The result comes from the internal model's score on the same 11 tasks; the text does not give its timing or further configuration details.
Training data consisted of simulated tasks built from contracts publicly available in the SEC's EDGAR database after filters designed to remove personal information; OpenAI customer data, OpenAI internal contracts, and nonpublic Ironclad customer data or contracts were not used. Stating this provenance lets outside readers judge the data-compliance and reproducibility boundaries of the evaluation. The text states the data source and exclusions directly in a footnote, as a direct statement of the training and evaluation data scope.
Perspective
The results cover the 11 research tasks, not all Ironclad workflows; the text states explicitly that the times are simulated estimates based on assumed model processing and generation speeds, not measured customer time savings. Methodologically, tasks were identified with Ironclad employees and people who use Ironclad at OpenAI, models practiced in hosted Ironclad product environments, and training tasks were created from contracts publicly available in the SEC's EDGAR database after filters designed to remove personal information. This combination suits software companies that want to turn their own professional workflows into model training and evaluation problems, provided they can bring people who know the work deeply, a secure testing environment, and data that can be safely used for research.
The criteria number 8 to 50 per task, but the text does not show their specific content or weighting, so readers cannot tell which capabilities the scoring is most sensitive to. The times are simulated estimates whose assumed processing and generation speeds are not detailed, so they should not be converted directly into real labor savings. The internal model's 63.7% result lacks configuration and timing information, leaving the source of that gain to be explained. Beyond that, generalization past the 11 tasks, and whether a model may lose track of a business rule mid-workflow, are addressed by the statement that human oversight still matters rather than by quantitative evidence.
