Safety & Privacy

OpenAI Slows Scaling and Promises Zero Data Retention for API Customers

OpenAI paused reinforcement learning training, slowed scaling, and announced zero data retention for eligible API customers starting September. Analysts call it IPO positioning.

LUMIEN5 min read
OpenAI Slows Scaling and Promises Zero Data Retention for API Customers

OpenAI announced this week that it has temporarily slowed its model scaling pace and put reinforcement learning (RL) training on a two-week pause while it audited and hardened its research environment. In a separate announcement the following day, the company said eligible API customers will be able to opt into a zero data retention program beginning in September. Analysts and consultants were skeptical, with several describing both moves as positioning ahead of a likely IPO rather than meaningful structural change.

What happened

Detail Fact
Scaling status Temporarily slowed; largest planned frontier RL run on hold
RL training pause length Two weeks
Monitoring overhead cost Roughly 20% of inference compute being monitored
Zero data retention start September (details in a technical white paper)
Eligibility for zero retention Not defined; OpenAI has not disclosed criteria

OpenAI said in a Tuesday statement that it paused reinforcement learning training while it carried out workload isolation, network isolation, and continuous security testing. The company says its largest planned frontier RL run stays on hold until smaller-scale training and evaluation give it “more evidence of alignment.” Going forward, it says it will require stronger evidence of aligned behavior throughout all of training, not just at the end.

The new monitoring system carries a meaningful cost: roughly 20% of the inference compute being monitored. OpenAI says that figure varies significantly across training and evaluation workloads. A fuller explanation of the system is promised in a forthcoming blog post.

What does zero data retention actually mean?

The Wednesday announcement said eligible API customers can opt into zero data retention starting in September. OpenAI did not define eligibility. A technical white paper with the specifics is due at the same time.

Jason Andersen, principal analyst at Moor Insights and Strategy, pointed out that the program’s value depends heavily on how the customer accesses the API. Much of OpenAI’s enterprise revenue flows through partners like Microsoft and AWS, not directly. If a business uses a tool such as Amazon Kiro, which calls OpenAI via API, Amazon is technically the customer. To qualify for zero data retention directly, the enterprise would need to supply its own API key, becoming a direct customer. At that point, according to Andersen, AWS loses access and the margin that comes with it.

Brian Levine, executive director of FormerGov, flagged a second complication. OpenAI says it can now monitor for abuse across interactions without staff reading the underlying content. Levine described that as “a strong technical promise, because watching for misuse and not being able to see the data have historically pulled in opposite directions.” He also noted a hard floor: content flagged as CSAM is still retained for legal reporting, so zero is never literally zero.

Why it matters

For businesses building on OpenAI’s API, these announcements touch two real concerns: whether training on customer data affects confidentiality, and whether the models they depend on will behave consistently as OpenAI keeps scaling them up.

The scaling pause is the more consequential of the two. If OpenAI’s largest planned frontier RL run stays frozen for any significant time, it affects how quickly new, more capable models reach production. That timeline ripples into any product roadmap built on top of those models. Our earlier coverage of OpenAI tightening model security after the Hugging Face breach shows this is part of a broader pattern of security-adjacent announcements from the company in recent weeks.

On the data side, enterprises in regulated industries (legal, finance, healthcare) have real reasons to care about data retention policies. But until the September white paper arrives with actual technical detail, there is no way to independently verify the promise.

Our take

The analysts quoted here are right to be cautious. A two-week RL pause and a data retention program with no published eligibility criteria are not nothing, but they are not auditable commitments either. Justin St-Maurice of Info-Tech Research Group put it well: if a carmaker announced it was taking basic safety testing more seriously before production, “it would be embarrassing that something so fundamental needed clarifying.”

The 20% compute overhead for monitoring is the one genuinely concrete number in these announcements, and it is worth noting because it signals that safety instrumentation at scale is expensive. That cost does not disappear; it gets baked into API pricing or absorbed by OpenAI until it does not want to absorb it anymore.

For teams already building AI integrations on top of OpenAI’s API, the practical advice from St-Maurice is sound: check your contract. If your vendor can pause development for security reasons and you found out through a blog post, your contract probably does not require them to tell you directly. That is a gap worth closing before the next pause, not after.

Flavio Villanustre, CISO at LexisNexis Risk Solutions Group, offered a third read: the moves may be aimed at softening incoming regulation by demonstrating willingness to self-regulate. That framing tracks with the IPO theory. Either way, none of it changes what you should do right now.

What to do about it

  1. Review your OpenAI API contract and check what disclosure obligations exist when the provider pauses or changes model training.
  2. Confirm whether you access OpenAI directly or through a partner like AWS or Azure, because that determines your eligibility for the zero data retention program.
  3. Wait for the September technical white paper before making any vendor decisions based on the data retention promise.
  4. If your use case involves regulated data, ask OpenAI’s sales team now what eligibility looks like so you are not surprised by the criteria in September.

If your current setup makes it hard to even answer question two above, that is a sign your AI integration architecture needs a closer look before you add more dependencies on any single provider.

Source: Bing News · OpenAI

Frequently asked questions

What is OpenAI's zero data retention program?

OpenAI announced that eligible API customers will be able to opt into a zero data retention program starting in September. The company has not yet defined eligibility criteria and says it will publish a technical white paper with details at launch. Note that content flagged as CSAM is still retained for legal reporting regardless of the program.

Why did OpenAI pause reinforcement learning training?

OpenAI said it paused RL training for two weeks while it hardened its research environment, expanded monitoring, and conducted workload and network isolation. It says its largest planned frontier RL run remains on hold until smaller-scale evaluations give it more evidence of model alignment.

How much does OpenAI's new safety monitoring cost?

OpenAI says the new monitoring system adds roughly 20% overhead on the inference compute being monitored, though it notes the cost varies substantially across training and evaluation workloads.

Is OpenAI's scaling pause related to its IPO?

Multiple analysts, including Carmi Levy and Jason Andersen of Moor Insights and Strategy, described the scaling pause and data retention announcement as positioning ahead of an anticipated IPO rather than purely substantive safety measures.

More from AI