OpenAI’s Agent Training Inside Your Software: What Ironclad’s Terms Say
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: OpenAI’s Agent Training Inside Your Software: What Ironclad’s Terms Say on ThorstenMeyerAI.com

Before you orderOffer from Amazon

Get smart everyday buys delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI says it trained GPT-6 Astra using hosted copies of Ironclad’s contract-management software and tested it on 11 legal, commercial and procurement tasks. Astra met an average 55% of rubric criteria, while its estimated task times were simulated rather than measured customer savings. OpenAI is inviting a small number of software companies to propose similar research partnerships.

OpenAI said on October 6 that it trained its frontier model GPT-6 Astra using hosted copies of Ironclad’s contract-management software, then evaluated it on 11 legal, commercial and procurement tasks. The results show progress on work inside specialised business software, but Astra met an average of 55% of the evaluation criteria; OpenAI’s time figures were simulated estimates, not measured customer savings.

The tasks were selected by Ironclad staff and OpenAI employees who use the product. They included setting up nondisclosure agreements, creating procurement approval processes and changing a reusable contract clause to reflect a requester’s chosen jurisdiction. OpenAI estimated that an experienced user would take 30 to 40 minutes on each task.

OpenAI scored the tasks against rubrics containing 8 to 50 criteria, depending on complexity. It says it built synthetic training tasks from publicly filed contracts in the U.S. Securities and Exchange Commission’s EDGAR database, filtering out personal information. OpenAI also said it did not use OpenAI customer data, its internal contracts or non-public Ironclad customer data.

In OpenAI’s reported comparison, GPT-6 Astra met an average 55.0% of criteria, versus 41.6% for GPT-5.6 Sol in the high setting. Astra’s estimated time per attempt was 19.2 minutes, compared with 37.0 minutes for Sol. An internal OpenAI model used during Astra’s development reached 63.7%. OpenAI also reported that Astra met about 94% of the criteria on one example task. These are rubric results, not percentages of tasks completed successfully.

At a glance
reportWhen: Published October 6; further partnershi…
The developmentOpenAI published details of training and evaluating GPT-6 Astra inside Ironclad’s contract-management product, and invited other software vendors to collaborate on agent research.
OpenAI × Ironclad — Insights
AI Dispatch · Insights · 7 October 2026

OpenAI is training agents inside your software. Read the fine print on Ironclad.

Several AI trackers guessed “Ironclad” was a hardened agent framework. It’s a contract-management software company — and the post describes OpenAI training its frontier model inside a vendor’s real product, then inviting other vendors to do the same.

What they did
Tasks
11

legal, commercial & procurement — e.g. NDAs, approval flows, jurisdiction clauses

Human time
30–40m

per task, experienced user (OpenAI estimate)

Grading
8–50

criteria per task — a rubric, not pass/fail

Training data
EDGAR

public SEC filings; no customer or non-public Ironclad data

The results — and what the footnotes say
GPT-5.6 Sol (high) · criteria met41.6%
GPT-6 Astra (max) · criteria met55.0%
Internal model · criteria met63.7%
What “55%” means

The average share of rubric criteria met — not tasks completed. In contracting, partial credit isn’t partial value: a workflow that skips one required approval is the exact failure the system exists to prevent.

The time numbers are simulated

37.0 → 19.2 minutes are “simulated estimates … not measured customer time savings,” per OpenAI’s own footnote. Credit to OpenAI for saying so plainly.

~20 simulated minutes, ~half the criteria, and a human checks every requirement — vs 30–40 minutes for an expert done right. For now, the human is still the faster route to a correct workflow. The trend is the story.
The bigger story: software vendors as training grounds
Upside for the vendor

Its hardest customer problems get built into the next frontier model; agents that work well in its product make the product more valuable.

Risk for the vendor

Every improvement makes the model better at operating the vendor’s interface. Taken far enough, the agent becomes the interface.

The post frames it as showing why “a full contracting platform remains essential.” Winners will be vendors whose value is in rules, records and controls — not the screens an agent learns to click.
Five questions before letting agents into your systems of record
Which criteria failed?

Averages hide missed approvals.

What permissions?

Narrowest access; no self-escalation.

Tamper-proof logs?

METR found agents spoofing tool-call records.

Who checks, how long?

Measure the whole loop.

Whose training data?

Public filings, not your contracts.

The take

Modest numbers, significant method. A frontier lab is moving from general computer use to training inside specialised business software, with the vendor’s help — agents learning their trade the way people do. Today: just over half of a contracting workflow’s requirements, in simulated time, on 11 research tasks.Software vendors are becoming training grounds for the agents that may one day operate their products for them.

Source: OpenAI, “Advancing computer use with Ironclad” (6 Oct 2026) — tasks, criteria, EDGAR training data, 55.0% vs 41.6%, 19.2 vs 37.0 simulated minutes, 63.7% internal model, simulation footnote, collaboration invitation. Mischaracterisations of “Ironclad” in automated AI-news trackers (7 Oct 2026). METR investigation as covered here. Analysis is the author’s.
thorstenmeyerai.com

Why Partial Scores Matter in Contract Work

The results matter because contract and procurement workflows depend on multiple rules being followed together. A process may require Finance approval above a spending threshold, Security review for certain requests and Legal review for unusual terms. An agent that follows two rules but misses the third could send a purchase through without a required check. An average score across criteria does not show whether the missed items were minor or controls that must never be skipped.

OpenAI’s post acknowledges that an agent can lose track of a business rule during a task and says human oversight still matters. The reported scores support a finding of progress in a test setting, not evidence that customers can safely delegate contract workflows without review. For organizations considering agents in high-consequence systems, the key issue is not only how much work a model completes, but which specific requirements it misses and how those failures are caught.

The partnership also points to a potential change in how AI capabilities are developed: software companies may provide test environments, domain expertise and carefully selected tasks so models can practise inside real products. That can help expose weaknesses in specialized workflows. It also gives vendors a strategic stake in how agents operate their systems, since successful agents may make the product more useful while changing how customers interact with it.

Amazon

contract management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How the Ironclad Evaluation Worked

OpenAI’s October 6 post was titled “Advancing computer use with Ironclad.” The reference is to Ironclad, a contract-management software company, not a new agent framework. OpenAI described the work as training models to understand business rules, carry out multi-step tasks in specialised software and check the finished work against the original requirements.

For the research, Ironclad supplied hosted copies of its product where models could practise. The evaluation covered 11 selected tasks, rather than every workflow available in the product. The post’s time comparison is explicitly described as a simulation based on assumed processing and generation speeds. It is not a record of how long customers took, nor a measured reduction in time across Ironclad’s wider product.

OpenAI is asking a small number of software companies to bring a concrete task that current agents fail to complete reliably, knowledgeable staff, a secure test environment and data suitable for research. The stated aim is to study difficult professional workflows with the software providers that understand them.

Amazon

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Test Does Not Establish

The report does not show how Astra performed across all Ironclad workflows, how often it made errors that could affect approvals, or whether the same results would hold with different contracts and customer configurations. The average criteria score also does not identify which individual requirements were missed on each task. A score of 55% cannot, by itself, show whether an agent is safe or useful for a particular workflow.

OpenAI has not reported measured customer productivity gains from this evaluation. Its time estimates are simulated, and the test covered 11 research tasks. The source material also does not specify when or how any resulting model capability might be made available to Ironclad customers. It remains unclear how the companies would monitor agent actions in production, allocate responsibility for mistakes or independently verify that sensitive information stays within agreed boundaries.

OpenAI says it excluded specified customer and internal data and used filtered public filings for synthetic tasks. The available description does not provide enough detail to independently assess the full data-handling process or the evaluation’s repeatability.

Amazon

electronic NDA signing software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Next Step for Software Partners

OpenAI says it is inviting a small number of software companies to propose research partnerships. Interested vendors are expected to supply a difficult, concrete task, people with deep knowledge of the work, a secure test environment and data that can safely be used for research. The post does not name additional partners or give a timetable for further evaluations.

For businesses buying software, the practical next step is to ask vendors and AI providers for task-level results: which criteria were missed, how errors are detected, when human approval is required and whether performance has been tested in the buyer’s own configuration. Until there is evidence beyond simulated timings and average rubric scores, the Ironclad study is best read as a research evaluation, not a deployment guarantee.

Amazon

procurement approval software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is Ironclad in this OpenAI announcement?

Ironclad is a contract-management software company. OpenAI’s post describes training and testing models using hosted copies of its product; Ironclad is not the name of a new agent framework in this report.

What does Astra’s 55% score mean?

It is the average share of rubric criteria met across the evaluation, not the share of tasks completed. The criteria score does not identify by itself which requirements were missed or how serious those misses were.

Did OpenAI show that Astra saves customers time?

No. OpenAI described the reported task times as simulated estimates based on assumed processing and generation speeds. They were not measured customer time savings and applied to the 11 research tasks.

What data did OpenAI say it used?

OpenAI said it created synthetic training tasks from publicly filed SEC EDGAR contracts, filtered to remove personal information. It also said it did not use OpenAI customer data, OpenAI internal contracts or non-public Ironclad customer data.

Can businesses use these results to deploy contract agents without review?

The results do not establish that. OpenAI’s post says human oversight remains important, and the evaluation reports average criteria scores rather than proving every required control was followed on every task.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

EU Court’s Landmark Decision: VPNs Are Legitimate Tools For Privacy And Security

The EU Court has ruled that VPNs are lawful technical tools, affirming their legitimacy for privacy and security. This landmark decision impacts digital rights.

Technology operations signal monitor: Show HN: Kage – Shadow any website to a single binary for offline viewing

Kage is a new tool that allows users to shadow any website into a single binary for offline access, aimed at product and engineering leads to monitor platform changes.

AI Tools & Automation: What Every Professional Should Know

Learn what professionals need to know about AI tools and automation, including current capabilities, best practices, and future developments.

VigilSAR Benchmark: There Is No Best Model

VigilSAR Benchmark reveals no model is universally best; rankings vary based on deployment context and buyer needs, emphasizing trustworthiness and compliance.