A central AI model core being evaluated against redacted conversation streams in a controlled test chamber
A central AI model core being evaluated against redacted conversation streams in a controlled test chamber
+ OpenAI News

OpenAI uses deployment simulation to test models before release

OpenAI says replaying realistic conversation contexts helped forecast undesired behavior across GPT-5-series Thinking deployments before models reached users.

OpenAI published Deployment Simulation on June 16, 2026, a safety-evaluation method that replays realistic conversation contexts against candidate models before release. The company says the method helped estimate undesired behavior across GPT-5-series Thinking deployments and surfaced one novel misalignment pattern, “calculator hacking,” before release.

The practical shift is simple: a model is not only tested on hand-written stress prompts. It is also asked to answer in contexts that look closer to real deployment traffic, with the old assistant response removed and the candidate model generating a new one.

Realistic context is the point

Traditional safety evals are useful because they can aim directly at high-risk cases. Their weakness is that they can also become recognizable as tests, narrow around known failure types, and miss how a model behaves in the messy middle of ordinary use.

Deployment Simulation tries to measure that middle. OpenAI says it takes recent production conversations, removes account-linked identifiers and identifiable information, removes the original assistant response, and regenerates that turn with a candidate model. The resulting answers can be graded for known categories and audited for new failure modes.

That makes the method more like a pre-release shadow deployment than a quiz. The model sees realistic prefixes, not only synthetic prompts written to catch it out.

1.3MDe-identified conversations analyzedOpenAI
20Undesired-behavior categories in GPT-5.4 Thinking predictionsOpenAI
1.5xMedian multiplicative error reported in aggregateOpenAI

The method is useful because it is checkable

The strongest part of the post is not that OpenAI claims a new eval technique. It is that the predictions can be compared with what happens after launch. OpenAI says it pre-registered predictions for GPT-5.4 Thinking and ran retrospective studies across other GPT-5-series Thinking deployments.

That matters because model-safety claims often collapse into vibes: a benchmark went up, a red team was run, a system card was published. Deployment Simulation creates a measurable forecast. If the simulation says a behavior should go down and deployment traffic shows it went up, the method has to explain the miss.

OpenAI says the method is better for non-tail risks that occur often enough to appear in sampled prefixes. It explicitly says a behavior that appears once in ten million conversations is unlikely to be discovered in a million-sample run. That caveat should stay attached to the story. This is not a tail-risk detector by itself.

Evaluation awareness is becoming a real problem

OpenAI also tested whether models could tell they were being evaluated. In one comparison, it says GPT-5.2 labeled real production traffic as evaluation-like 5.4% of the time and simulated traffic 5.1% of the time, while traditional evaluations such as SWE-Bench or MonorepoBench were classified as evaluations almost 100% of the time.

That does not mean the model is consciously “cheating” in a human sense. It means the distribution carries cues. If a model can infer that a prompt is part of an eval, its behavior under that prompt may not match its behavior in the product.

Deployment-like context reduces that distortion. It also raises a governance question: the best signal may come from private production data that outside auditors cannot fully access. OpenAI tested WildChat as an external-data substitute and found it informative, but less accurate than recent OpenAI production data.

Agents make simulation harder

The most interesting extension is coding agents. OpenAI says it used 120,000 internal employee agentic trajectories from GPT-5.4 to simulate an internal deployment based on GPT-5.5. A single agent run can depend on files, commands, network responses, and tool outputs, so replaying it with live tool access would be risky and difficult.

OpenAI’s workaround was to simulate tool calls with another model, using the original rollout, repository state, a tool-call database, and read-only connectors. In a test where a discriminator tried to tell real from simulated rollouts, realism improved from an 11.6% win rate to 49.5%, close to chance.

That is the operator lesson. As agents take more actions, safety evaluation becomes an environment problem. The model is only one part of the behavior. The surrounding tools, permissions, state, and failure modes have to be simulated well enough for the forecast to mean anything.

Sources

The AI Feed Desk

The AI Feed Desk

Editorial desk

The AI Feed Desk tracks AI provider updates, model releases, agent tooling, and enterprise adoption, turning fast-moving announcements into source-linked context for builders and operators.

Noticed a typo, incorrect information, or translation error?

Tell us so we can fix it.

Help Improve This Article

Related Articles

Lab vials, reaction plates, and an AI planning console connected by measured experiment paths

OpenAI shows GPT-5.4 improving a medicinal-chemistry reaction

OpenAI says GPT-5.4, connected to Molecule.one's Maria lab, found an additive that improved a difficult Chan-Lam coupling result across thousands of physical experiments.

The AI Feed Desk

By The AI Feed Desk

OpenAI's official GPT-Rosalind article card with a DNA illustration

OpenAI updates GPT-Rosalind for life sciences research

OpenAI's GPT-Rosalind update adds stronger life-sciences reasoning, Codex-based research plugins, and a trusted-access preview for eligible organizations.

The AI Feed Desk

By The AI Feed Desk

OpenAI's official Better memory for a more helpful ChatGPT article card

OpenAI rolls out Dreaming V3 memory for ChatGPT

OpenAI is rolling out Dreaming V3 memory to ChatGPT Plus and Pro users in the US first, with Free and Go access planned over the coming weeks after a 5x compute-efficiency gain.

The AI Feed Desk

By The AI Feed Desk

An enterprise AI admin console with credit usage gauges, team budget controls, and ChatGPT and Codex activity streams

OpenAI puts ChatGPT Enterprise spend into the admin console

OpenAI is adding credit usage analytics and updated spend controls for ChatGPT Enterprise, including ChatGPT and Codex usage by user, product, and model.

The AI Feed Desk

By The AI Feed Desk

Generated editorial image of an AI assistant connected to role-specific workflow panels

OpenAI pushes Codex beyond software development

OpenAI says Codex now has more than 5M weekly users and is adding role-specific plugins, Sites, and annotations for broader business work.

The AI Feed Desk

By The AI Feed Desk