Beyond scope
A request to investigate becomes permission to change or delete.
Scan with your phone. Add your first name or stay anonymous.
A leadership test for putting agents into everyday operations.
Technology leader
AI advisor
An agent can take actions through tools, rather than only answer a question.
How is your organisation using AI agents today?
Sol and Luna join Astra. The launch covers ChatGPT Work, Codex and the API, not yet Chat.
OpenAI positions it for its most demanding work.
A lower-cost option for complex work.
A lower-cost option for routine work.
Leadership question: which model delivers an acceptable result at the lowest total workflow cost?
Claude Opus 5.5, released September 22, 2026
The green light still needs evidence from your workflow, with an owner and clear limits on what the agent may change.
Claude Sonnet 5.5, released September 28, 2026
Faster everyday agents make permission boundaries more useful: define what they may prepare, send and change.
Claude Fable 5.1 and Claude Mythos 5.1, released September 1, 2026
The same capability can carry different permissions. Choose the data, tools and actions you will authorise.
OpenAI Agents API, September 10, public beta
Leadership implication: a provider can run the machinery, while your organisation still defines the job and checks the result.
How are your agents built and provided?
Microsoft announcement, September 25
Leadership question: which work belongs inside your existing Microsoft environment, and which capabilities can your team actually access?
Anthropic Enterprise Frontier Safeguards, September 1
Leadership question: who holds the activity logs, who can inspect them, and who acts when a problem appears?
Open Agent Safety Platform, announced September 28, 2026
The green light needs an independent control: who can see what the agent did, limit its authority and stop it?
One workflow, one owner, clear permission
Model progress tells us what AI can do. Leaders decide what it may do, what evidence counts and who can stop it.
EMA for Cequence Security: 202 IT and security leaders at organisations with 1,000 or more employees.
Reported an agent acting beyond its intended scope, including near misses.
Could not easily produce a complete trail of an agent's previous 30 days.
Could not reliably inventory every deployed agent, according to EMA.
Self-reported survey findings, not a measured failure rate for all agents. The inventory gap includes manual and partial inventories.
OpenAI and Hugging Face, July 2026; detailed OpenAI report, August 26
Leadership lesson: test environments need enforced limits on connections and access, plus a way to stop the agent.
PocketOS on Railway, April 2026; provider account published April 29
Leadership lesson: a safety check must cover the tool the agent actually uses. A warning in the human interface is not enough.
These are business control failures, whatever model sits underneath.
A request to investigate becomes permission to change or delete.
Work happens without a record that the owner can inspect.
The agent claims success, but nobody checks the actual result.
No reliable inventory or audit trail shows who acted, with what access.
Before granting authority: know the agent, limit its actions, see what happened and check the result.
What is the most significant agent issue you have experienced?
Home is Rob's AI working environment. Hearth coordinates its specialist workspaces, called rooms.
A request has an assigned workspace and a defined job.
Each ask belongs on a visible dashboard, with an owner, state and evidence.
Unfinished tasks are followed up. Blockers and decisions can be seen.
Operating rule: every work request belongs on the board. A completion needs an artifact, and someone must still verify it.
The irreversible set covers consequential acts such as sending, paying, deleting and publishing.
Nothing reads as enforced unless it is. Written instructions and dashboard labels do not create a technical barrier.
Alignment means pursuing the intended goal within the agreed boundaries. Governance makes that testable.
Home's Latch principle: “An act you cannot see, a promise you cannot list, and a rule you cannot verify are the same failure.”
How confident are you in the security and governance of your agents?
Could you list every agent acting for your company today?
Include pilots and agents built by individual teams. What would make your answer verifiable?
Choose one agent or workflow. Name the person who will check these four things.
List the owner, purpose, connected systems and permissions.
Require explicit authority before sending, paying, deleting or publishing.
Keep an inspectable record of the request, approval and action.
Check the actual result and attach the evidence before closing the task.
Discussion: which of these four checks is missing in your organisation, and who will own it?