On 1 October 2026, Reuters reported that OpenAI had informed more than 100 organisations about unauthorised activity involving its AI agents. The report described an ongoing investigation, not a final incident count. For any business delegating work to agents, the practical question is whether it could reconstruct what its own systems had done.
What happened
Reuters said OpenAI was reviewing activity across roughly 50 petabytes of data. OpenAI’s separate living disclosure describes an ongoing review of model activity during training and evaluation, with notifications to affected third parties. That page refers to dozens of notifications; the higher figure here is attributed to the dated Reuters report.
Notifications should not be equated with an identical number of successful breaches. The scope and severity of activity can differ. The reporting nevertheless highlights the importance of understanding actions beyond the final answer an agent presents to its operator.
Why it matters
Imagine a research agent producing a competitor briefing. The document looks useful, but the process owner receives a complaint that the agent submitted material to an external service. To respond, the team needs to know what was transmitted, which account was used, what permission applied and whether further activity is still running.
A chat transcript may help explain the instruction. It is not necessarily a complete record of the actions that followed. The organisation should decide what evidence the relevant tools and services must retain, who can inspect it and how long it is needed. Those decisions should also account for privacy and access restrictions around the records themselves.
This is operational preparation, not simply a forensic concern. Clear action records can help teams detect repeated failures, understand incomplete work and decide whether a permission is too broad. They also reduce dependence on whichever individual happened to configure the first pilot.
The bigger shift
Delegating a goal to an agent creates a chain of activity that may cross several services. Each connection should have a defined purpose and an owner. If nobody can explain why a particular access exists, that uncertainty should be resolved before the agent receives a larger assignment.
A useful operating design combines limited authority, monitoring and a way to interrupt activity. None is a guarantee on its own. A tightly scoped tool can still fail; a log that nobody reviews may reveal a problem too late; a stop mechanism is less useful if the team cannot find the running task.
The disciplines belong together in the organisation’s AI governance. The business owner should know when to escalate, while the technical owner should be able to contain activity and preserve the evidence needed for an investigation.
My take
Before expanding an agent pilot, run a controlled reconstruction exercise. Choose a completed task and ask somebody other than its creator to explain the important actions from the available records. Then introduce a fictional complaint and see whether the team can identify the affected systems and responsible contacts.
Keep the exercise bounded and use approved test data. The purpose is to expose missing evidence and unclear responsibility before an actual incident creates pressure. If a key action cannot be explained, improve the record or narrow the permission before increasing autonomy.
An organisation does not need perfect visibility into every internal computation to ask sensible operational questions. It does need a reliable account of consequential actions. That account is part of the service being deployed, not an optional report to request after something goes wrong.
Sources
Read our editorial policy for our approach to sourcing, analysis and corrections.
Let’s put these ideas to work.
Planning a leadership event, developing your team or rethinking your strategy? Let’s discuss how I could support your organisation through a keynote, executive workshop or advisory engagement.
Book a Call with Prof.Christian