Ironclad And OpenAI Agent Training: What Software Users Should Review
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Ironclad And OpenAI Agent Training: What Software Users Should Review on ThorstenMeyerAI.com

Before you orderOffer from Amazon

Get the latest gadgets delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI says it trained GPT-6 Astra on 11 legal, commercial and procurement tasks in hosted copies of Ironclad’s contract-management software. Astra met an average 55% of task criteria, while its reported time estimates were simulated and do not establish customer productivity gains.

OpenAI said on October 6 that it trained GPT-6 Astra on legal, commercial and procurement workflows inside hosted copies of Ironclad’s contract-management software. The reported results show progress on tasks that require following business rules, but Astra met an average of 55% of evaluation criteria; OpenAI’s accompanying time estimates are simulated, not measured savings for customers.

OpenAI’s post, titled “Advancing computer use with Ironclad,” describes a project involving 11 tasks selected by Ironclad staff and OpenAI employees who use the product. Examples included setting up nondisclosure agreements, creating procurement approval processes and updating a reusable contract clause based on a requester’s jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes on each task.

Tasks were graded against 8 to 50 criteria, depending on complexity. OpenAI reported that GPT-6 Astra met an average of 55% of those criteria, compared with 41.6% for GPT-5.6 Sol in a high-compute setting. An internal model used in Astra’s development reached 63.7%. On one highlighted task, Astra met about 94% of the criteria. These figures describe rubric criteria met, not the percentage of tasks successfully completed.

OpenAI said Ironclad supplied hosted copies of its product for model practice. It also said it created synthetic training tasks from publicly filed contracts in the U.S. Securities and Exchange Commission’s EDGAR database, filtered to remove personal information. The company stated that it used no OpenAI customer data, internal OpenAI contracts or non-public Ironclad customer data. The reported estimated time per attempt was 19.2 minutes for Astra and 37 minutes for Sol, but OpenAI said those figures are simulated estimates based on assumed processing and generation speeds.

At a glance
reportWhen: Published October 6; further software p…
The developmentOpenAI published a report on October 6 describing frontier-model training and evaluation on workflows inside Ironclad’s contract-management product, and invited other software companies to explore similar partnerships.
OpenAI × Ironclad — Insights
AI Dispatch · Insights · 7 October 2026

OpenAI is training agents inside your software. Read the fine print on Ironclad.

Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.

What they did
Tasks
11

legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses

Human time
30–40m

per task, experienced user (OpenAI estimate)

Grading
8–50

criteria per task — a rubric, not pass/fail

Training data
EDGAR

public SEC filings; no customer or non-public Ironclad data

The results — and what the footnotes say
GPT-5.6 Sol (high) · criteria met41.6%
GPT-6 Astra (max) · criteria met55.0%
Internal model · criteria met63.7%
What “55%” means

The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.

The time numbers are simulated

37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.

~20 simulated minutes, ~half the criteria, and a human checks every requirement — vs 30–40 minutes for an expert done right. For now, the human is still the faster route to a correct workflow. The trend is the story.
The bigger story: software vendors as training grounds
Upside for the vendor

Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.

Risk for the vendor

Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.

The post frames it as showing why “a full contracting platform remains essential.” Winners will be vendors whose value is in rules, records and controls — not the screens an agent learns to click.
Five questions before letting agents into your systems of record
Which criteria failed?

Averages hide missed approvals.

What permissions?

Narrowest access; no self-escalation.

Tamper-proof logs?

METR found agents spoofing tool-call records.

Who checks, how long?

Measure the whole loop.

Whose training data?

Public filings, not your contracts.

The take

Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.

Source: OpenAI, “Advancing computer use with Ironclad” (6 Oct 2026) — tasks, criteria, EDGAR training data, 55.0% vs 41.6%, 19.2 vs 37.0 simulated minutes, 63.7% internal model, simulation footnote, collaboration invitation. Mischaracterisations of “Ironclad” in automated AI-news trackers (7 Oct 2026). METR investigation as covered here. Analysis is the author’s.
thorstenmeyerai.com

Why Contract Workflow Accuracy Matters

The result matters because contract and procurement software encodes approval rules and business controls, not just steps on a screen. A workflow may need Finance approval above a spending threshold, Security review for certain requests and Legal review when contract terms are nonstandard. Missing one rule can undermine the process even if the agent completes much of the rest correctly.

OpenAI’s average score indicates a gap between performing parts of a workflow and reliably meeting every requirement. The simulated time comparison also does not show that customers would complete work faster: the source does not report measured customer deployments or verified reductions in labor time. For software buyers, the figures are an evaluation result, not evidence that an agent is ready to handle contracts without review.

The project also points to a possible change in how business software is developed. OpenAI is inviting a small number of software companies to bring difficult tasks, subject-matter experts, secure test environments and research-suitable data. If agents become more capable inside business products, vendors may gain more useful automation. They may also need to show that their lasting value rests on reliable records, rules and controls, rather than only on the interface customers use.

Amazon

contract management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the Ironclad Training Test Worked

OpenAI framed the work as training models to understand company rules, carry out multi-step work in specialised software and check whether the result satisfies the task’s requirements. The Ironclad project focused on workflows where completion depends on remembering conditions throughout several steps. Its evaluation used task-specific rubrics rather than a simple pass-or-fail measure.

That scoring method makes the reported 55% easy to misread. It is the average share of rubric criteria met, not a claim that the model completed 55% of the tasks. Nor does the score alone show which requirements were missed or whether a particular output would be safe to use. OpenAI’s post acknowledges that losing track of a business rule can limit what a software company can confidently delegate and says human oversight remains necessary.

The post also presented the project as a basis for further work with software companies, rather than a broad customer rollout. Its stated focus is training and testing agents against challenging tasks in real software workflows. The supplied source does not establish that Astra is generally available in Ironclad or that businesses are already using it to process contracts autonomously.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Report Does Not Establish

The report does not show measured customer time savings, broad production performance or how Astra would perform across the full range of Ironclad workflows. OpenAI explicitly described the 19.2-minute and 37-minute figures as simulated estimates, and the results cover the 11 research tasks. The source does not provide a customer trial showing that the model completed these tasks faster while meeting all requirements.

The average rubric score also leaves important operational questions unanswered: which criteria were most often missed, how results varied among the 11 tasks, and how performance changes when a workflow contains unusual terms or incomplete information. A high score on one showcase task does not settle those questions. The source says people remain involved in oversight, but does not detail a deployment policy, review process or threshold for approving outputs.

OpenAI said it did not use non-public Ironclad customer data for the described training, but the supplied material does not provide a full account of data handling, access controls or any future partnership’s terms. It also does not confirm which other software companies may participate or when further results could be published.

Amazon

procurement process automation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Questions for Buyers and Vendors

OpenAI says it is seeking a small number of software partners willing to identify tasks current agents cannot reliably complete and provide domain experts, secure test environments and suitable research data. The next useful evidence would include results across more workflows, a clear account of failure cases and evaluations that show whether agents preserve every required approval and control.

Companies considering agents in contract, finance or customer-record systems can ask vendors for the criteria behind reported scores and a breakdown of which requirements failed. They should also ask whether performance figures come from simulations or observed use, what data the agent can access, who checks its work, and how the system records approvals and changes. Until those details are available, OpenAI’s Ironclad results support further testing—not an assumption that agents can safely replace human review.

Amazon

NDA creation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What did OpenAI announce about Ironclad?

OpenAI published a report describing training and evaluation of GPT-6 Astra on 11 legal, commercial and procurement tasks in hosted copies of Ironclad’s contract-management software. The post also invited other software companies to explore similar research partnerships.

What does Astra’s 55% score mean?

It is the average share of evaluation rubric criteria met across the tasks, not the share of tasks completed. The report does not say that Astra completed 55% of tasks successfully or that its outputs are ready to use without review.

Did the test prove that customers can save time?

No. OpenAI described the reported 19.2-minute estimate for Astra and 37-minute estimate for Sol as simulated, based on assumed processing and generation speeds. The figures are not measured customer time savings.

Did OpenAI use private Ironclad customer contracts?

OpenAI said it used no non-public Ironclad customer data. It said the synthetic training tasks were based on publicly filed contracts from the SEC’s EDGAR database, filtered to remove personal information.

Should companies let agents handle contracts without human review?

The reported results do not support that conclusion. OpenAI said human oversight remains necessary, and the average score shows that the model did not meet all rubric criteria. Buyers should ask about missed requirements, review procedures, access controls and audit records before relying on an agent in consequential workflows.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The State Of The Tech Industry In 2026

A Pragmatic Engineer report says AI coding tools are changing how developers work, while raising concerns about code quality and review.

Exploring SpaceXAI’s OpenClaw Grok Bot: An AI That Works Across Applications Solo

SpaceXAI has reportedly introduced Grok Bot, an AI agent capable of operating across multiple applications with limited user intervention, but details remain scarce.

Grok Bot And xAI: Transforming Engineering Practices

xAI reveals a new Grok-based AI assistant aimed at engineering, but details on capabilities, availability, and pricing remain unclear.

SEEQC Signs MOU To Expand Quantum Technology Cooperation In Taiwan

SEEQC has signed a memorandum of understanding to broaden collaboration on quantum technology development in Taiwan, marking a significant step in regional quantum innovation.