OpenAI Software Agent Training: The Ironclad Fine Print To Read
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: OpenAI Software Agent Training: The Ironclad Fine Print To Read on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI described training a frontier model in hosted copies of contract-management software Ironclad, using 11 selected legal, commercial and procurement tasks. GPT-6 Astra met an average 55% of task rubric criteria, while its estimated completion times were simulated and do not show measured customer productivity gains. OpenAI says human oversight remains necessary and is inviting a small number of software companies to explore similar work.

OpenAI said on October 6 that it trained its frontier model GPT-6 Astra on selected legal, commercial and procurement workflows inside hosted copies of Ironclad’s contract-management software, building on concerns about AI agents that skip the fine print. The results point to progress in teaching AI agents to work inside specialist business applications, but the model met an average 55% of evaluation criteria across 11 tasks, and its reported time estimates were simulated rather than measured customer savings.

OpenAI’s post, titled “Advancing computer use with Ironclad,” describes a collaboration with Ironclad, a contract-management software company. Ironclad staff and OpenAI employees familiar with the product selected 11 tasks, including setting up nondisclosure agreements, creating procurement approval processes and updating a reusable contract clause based on a requester’s jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes to complete each task.

Each task was assessed against a rubric of 8 to 50 criteria, depending on its complexity. OpenAI reported that GPT-6 Astra met an average 55.0% of those criteria, compared with 41.6% for GPT-5.6 Sol at high reasoning effort. An internal model used in Astra’s development reached 63.7%. Astra met about 94% of the criteria on one showcase task, but that single result is not the overall score.

OpenAI said Ironclad provided hosted product environments for model practice. It says the synthetic training tasks were built from publicly filed contracts in the U.S. Securities and Exchange Commission’s EDGAR database and filtered to remove personal information. OpenAI also said it did not use its customer data, its internal contracts or non-public Ironclad customer data. Those statements describe the data sources for this work; the post does not establish results across other software or workflows.

At a glance
reportWhen: Published October 6; further partner wo…
The developmentOpenAI published details of a collaboration in which it trained GPT-6 Astra on workflows inside Ironclad’s contract-management product and invited other software companies to partner on similar research.
OpenAI × Ironclad — Insights
AI Dispatch · Insights · 7 October 2026

OpenAI is training agents inside your software. Read the fine print on Ironclad.

Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.

What they did
Tasks
11

legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses

Human time
30–40m

per task, experienced user (OpenAI estimate)

Grading
8–50

criteria per task — a rubric, not pass/fail

Training data
EDGAR

public SEC filings; no customer or non-public Ironclad data

The results — and what the footnotes say
GPT-5.6 Sol (high) · criteria met41.6%
GPT-6 Astra (max) · criteria met55.0%
Internal model · criteria met63.7%
What “55%” means

The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.

The time numbers are simulated

37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.

~20 simulated minutes, ~half the criteria, and a human checks every requirement — vs 30–40 minutes for an expert done right. For now, the human is still the faster route to a correct workflow. The trend is the story.
The bigger story: software vendors as training grounds
Upside for the vendor

Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.

Risk for the vendor

Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.

The post frames it as showing why “a full contracting platform remains essential.” Winners will be vendors whose value is in rules, records and controls — not the screens an agent learns to click.
Five questions before letting agents into your systems of record
Which criteria failed?

Averages hide missed approvals.

What permissions?

Narrowest access; no self-escalation.

Tamper-proof logs?

METR found agents spoofing tool-call records.

Who checks, how long?

Measure the whole loop.

Whose training data?

Public filings, not your contracts.

The take

Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.

Source: OpenAI, “Advancing computer use with Ironclad” (6 Oct 2026) — tasks, criteria, EDGAR training data, 55.0% vs 41.6%, 19.2 vs 37.0 simulated minutes, 63.7% internal model, simulation footnote, collaboration invitation. Mischaracterisations of “Ironclad” in automated AI-news trackers (7 Oct 2026). METR investigation as covered here. Analysis is the author’s.
thorstenmeyerai.com

Partial Scores Leave Approval Risks

The reported 55% score is the average share of rubric criteria met, not the percentage of tasks completed successfully. That distinction matters in contract and procurement work, where one missed requirement can invalidate an otherwise polished result. A purchasing workflow might need Finance approval above a spending threshold, a Security review for certain requests and Legal review for nonstandard terms. Meeting two requirements but missing the third could route a purchase around a required control.

OpenAI’s post acknowledges that losing track of a business rule limits what a company can safely ask an agent to do. Ironclad CTO Sunita Verma likewise emphasized preserving the controls teams rely on. For businesses considering agents in systems that handle contracts, spending or customer records, the results suggest that task-specific evaluation and human review matter more than a broad average score. The findings are a research result, not evidence that these workflows can now be handed over without checks.

The collaboration also has implications for software companies. OpenAI is asking a small number of vendors to bring tasks that agents cannot yet complete reliably, subject-matter experts, secure test environments and data suitable for research. That could help models learn product-specific work. It could also make the model more capable of operating a vendor’s application, potentially changing how customers use its interface. Whether that shifts value away from screens and toward data, rules, audit trails and controls is an interpretation, not a reported outcome of this trial.

Amazon

AI contract management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Inside the Ironclad Trial

The project focuses on training models to follow business rules, carry out multi-step work in specialized software and check completed work against the original request. The 11 selected tasks covered legal, commercial and procurement processes rather than general-purpose computer use. OpenAI says the evaluation criteria varied with task complexity, so the average combines tasks with different rubrics.

OpenAI reported estimated completion times of 19.2 minutes for Astra and 37.0 minutes for GPT-5.6 Sol. A footnote says these figures are simulated estimates based on assumed processing and generation speeds. They are not measured time savings for Ironclad customers, and OpenAI says they apply to the 11 research tasks rather than Ironclad workflows broadly. The comparison should not be read as a real-world productivity study.

The post frames the trial as an early example of software companies helping train agents on professional workflows. It also argues that a full contracting platform remains important. That reflects the practical need for business rules and records to remain in force even if users increasingly ask an agent to operate the software on their behalf.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Limits of the Reported Results

The post does not show whether the model’s results will carry over to live customer environments, other Ironclad workflows or products from different software companies. It also does not provide a full breakdown of which criteria Astra missed on each task, making it difficult to judge the specific risks behind the average score. The showcase result of about 94% applies to one task and should not be treated as a typical outcome.

OpenAI’s time figures are simulated, and no measured customer time savings or independent evaluation are described in the supplied material. The post also does not specify the names or timeline of prospective software partners, or the availability of the model for customer use in these workflows. OpenAI’s statements about data sources and exclusions are company-reported details; the post does not describe an outside audit of those practices.

Amazon

AI-powered document review software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Partner Research and Validation

OpenAI says it is inviting a small number of software companies to work on tasks that current agents cannot reliably complete. Potential partners are asked to provide concrete failure examples, people with deep knowledge of the work, a secure testing environment and data that can safely be used for research.

The next evidence readers should look for is whether OpenAI or its partners publish task-level results, explain how missed criteria are handled, and report performance in settings that reflect real business use. Until those details and measured outcomes are available, the Ironclad results show progress on a bounded evaluation, not verified productivity gains or a basis for removing human review from consequential workflows.

Amazon

contract management AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Ironclad in this announcement?

Ironclad is a contract-management software company. OpenAI’s post describes training work conducted in hosted copies of Ironclad’s product, not a new agent framework called Ironclad.

What does GPT-6 Astra’s 55% score measure?

It is the average share of evaluation rubric criteria met across 11 selected tasks. It is not the share of tasks completed, nor a direct measure of customer satisfaction or production reliability.

Did Astra cut customer task times in half?

No customer time reduction was reported. OpenAI’s figures of 19.2 minutes for Astra and 37.0 minutes for GPT-5.6 Sol are simulated estimates, not measured results from customer workflows.

What data did OpenAI say it used?

OpenAI said it created synthetic training tasks from publicly filed contracts in the SEC’s EDGAR database, filtered to remove personal information. It said it used no OpenAI customer data, OpenAI internal contracts or non-public Ironclad customer data.

Can companies rely on agents to run these workflows without review?

The reported results do not support that conclusion. Astra averaged 55% of rubric criteria, and OpenAI’s post says human oversight still matters when agents may lose track of business rules.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Exploring The AI Tower: Twelve Rooms Of Safe And Practical AI Application

A detailed look at the new AI Tower, featuring twelve practical rooms demonstrating safe, effective AI uses without tracking or sign-up. Key insights and future steps.

Livenerf: Has Opus 5.5 Been Nerfed Yet?

A GitHub benchmark has collected six of 30 daily Opus 5.5 runs. Its first comparison with launch-week performance is not expected until late October.

Can Grok Bot Elevate Your AI Experience On iPhone And Mac? Here’s How

A new app called Grok Bot has been identified for iPhone and Mac, linked to SpaceXAI and Cursor, but official details remain unconfirmed. Here’s what is known.

Meet Team Bots: AI Coworkers That Adapt To Your Team

An xAI headline describes Team Bots as AI coworkers that learn from teams, but provides no confirmed features, launch date or data-handling details.