The lean compliance team in this example is two people. One leads the program; the other runs the operations. They report to a CFO. They carry SOC 2 plus a state privacy regime, support enterprise sales with security questionnaires, and respond to roughly 40 to 80 control events per week from the cloud-posture scanner and the configuration baseline.
At Tier 1 supervised AI — the suggest-only model the rest of the market ships — this team's bottleneck is the queue. The agent drafts beautifully. The humans approve slowly. The week ends with a backlog that re-becomes Monday's backlog.
At Tier 2 scoped autonomy under Trustworthy Autonomy™, the bottleneck shifts. Most events resolve without entering the queue at all. The humans do the work the agent is not permitted to do, and they do it with more context than they had before. The team size is unchanged. The outcome is not.
Walk through the working week to see what actually shifts.
What "scope" looks like in writing
Before the week starts, the program lead defines the scope of the agent's Tier 2 authority — in writing, signed, in the delegation record. The scope reads like a policy floor, not a description of capability. Three example clauses (the real scope-of-authority document for a customer of this size runs to roughly 20):
- Categories the agent may act on autonomously: attest_control for SOC 2 CC6.1, CC6.6, CC6.7, CC7.1 (drift detection on existing baselines, not initial certification). remediate_finding for cloud-posture findings in the public-data-exposure family. respond_questionnaire for questions that draw exclusively from the published answer library at confidence ≥ 0.92.
- Categories the agent must escalate: Any attestation that requires a control judgment not covered by an existing adjudicated case in the customer's strata. Any remediation that touches IAM, encryption keys, or production database configuration. Any questionnaire question requiring a net-new factual claim. Out-of-scope is the default, not the exception.
- Conditions that demote any in-scope category to escalation: The runtime monitor flagged the trajectory as anomalous. The certified wrong-attestation rate for this category climbed above the Tier 2 threshold in the last calibration run. The canary check against a Tier 1 supervised baseline disagrees on more than 3% of the last 200 actions.
Scope is the operational artifact a Tier 2 program lives or dies on. It is not "what the agent can do" — it is "what the agent is permitted to do without a human in the path, and under what evidence." Two-person teams that get Tier 2 right spend more time on the scope document than on the agent itself. The agent runs against the scope; the scope is the contract.
Monday morning — the queue
At 8:30 the program lead opens the agent's overnight summary. Forty-two events processed since 5pm Friday. Of those, 31 are in-scope attestations and remediations that the agent handled autonomously — the action records are in the chain, the runtime monitor saw nothing anomalous on any trajectory, and the canary comparison against the Tier 1 baseline (shadowing 10% of in-scope actions) agreed on all of them.
Eleven are escalations. The agent describes each one in two sentences: what was detected, why it escalated. Five are out-of-scope category escalations (an IAM finding, a production database configuration, three net-new questionnaire claims). Six are in-scope events the agent attempted and demoted — the runtime monitor caught one as anomalous, and the other five were below the confidence floor for the action category.
The program lead reviews the 11 escalations. Three are clear; she approves and the agent applies them. Four require a brief conversation with the on-call engineer (the IAM finding, two of the database items, one of the new questionnaire claims). Four are non-trivial and go on the week's roadmap.
The 31 autonomously-handled events were not in her queue. She reviews a sample of three of them from the action chain because that is the standard discipline at this customer — spot-check three per morning, log the result. All three match what she would have decided.
What the operations lead does at 10am
The operations lead's morning used to be triage and questionnaire drafting. The triage is now the agent's. The questionnaires are partially the agent's — the in-scope, library-backed answers are drafted and posted; the net-new claims are escalated and waiting for him.
He spends an hour on the four net-new claims, drafts the answers, and pulls in the security architect for one of them. The other three go back into the questionnaire flow as agent-drafted, human-approved answers and update the answer library so the next time the question appears the agent can handle it.
Then he reviews the runtime monitor's weekly drift report. This is the operational task that is genuinely new at Tier 2. The monitor is not free; it has to be tended. The report shows three trajectories the monitor halted in the last week, the reason for each halt, and whether the halt was the right call. Two were correct (one tool misuse, one looping pattern). One was a false halt — the agent was correctly proceeding on an edge case the monitor's model did not recognize. He marks it, the monitor's calibration set updates, and the threshold tightens slightly.
What changed, in numbers
For this customer profile — two-person team, 40-80 events per week, mid-mid-market control surface — the operational shape at Tier 2 versus Tier 1 looks like this:
The headcount line is the one the conversation usually wants to be about. It is not the right line. The point of scoped autonomy at this team size is not to shrink the team; it is to redirect the team from the routine to the consequential. The lead spends her recovered hours on the work that used to drop off the bottom of the priority list — the policy refresh, the questionnaire library buildout, the actual program improvements. The ops lead spends his on the net-new claims and on tending the monitor calibration.
Tier 2 is a scope document, not a license. Five categories of action that stay human-gated at this customer profile regardless of the agent's accuracy on adjacent categories: any change to IAM principals, any change to production encryption configuration, any attestation on an initial certification (versus drift on an established baseline), any net-new factual claim in a questionnaire, and any cross-account or cross-tenant action. The agent's correct behavior on these is to escalate, not to attempt. The certified evidence for moving any of these to Tier 2 does not exist yet at this customer.
What this requires of the team, beyond the agent
The team that runs scoped autonomy well looks operationally different from the team that does not. Three habits this customer's team developed in the first six weeks of running Tier 2:
The scope document is a living artifact. The lead updates it weekly. New categories enter scope when the certified evidence supports them; categories exit scope when the calibration run shows they no longer clear the threshold. The scope document is signed each time and the chain records the version that was in effect for each action.
The spot-check discipline is non-negotiable. The lead reviews three autonomously-handled actions every morning from the action chain. Two minutes per action, six minutes a day, half an hour a week. The point is not to catch anything — the monitor and the certified accuracy are doing that — but to keep the human's intuition calibrated against the agent's behavior.
The monitor's calibration is somebody's named job. Drift, edge cases, and false halts move the monitor's behavior over time. Without a named owner, the monitor gets stale; with a named owner, the runtime gate keeps clearing its floor. At this customer, the ops lead owns the monitor.
Where this works and where it does not
Scoped autonomy at Tier 2 is the right operational posture for a compliance team when three conditions hold: the action categories have certified evidence above the Tier 2 thresholds, the runtime monitor has been measured on the team's actual workload (not just a vendor benchmark), and the team has the discipline to run the scope document as a live operational artifact rather than a one-time setting.
It is the wrong posture when any of those conditions fail. A team that adopts Tier 2 on a category whose evidence has not cleared the threshold has shipped suggest-only-with-extra-steps. A team that cannot tend the scope document loses the contract that makes the whole arrangement defensible. A team without an honest runtime monitor is operating on faith, which is what the methodology was built to retire.
The next article in the series is about the destination past Tier 2 — goal-oriented compliance, where the human sets an objective and the agent plans the work to achieve it within guardrails. That tier is not shipping today. The final article in this series explains why, and what evidence would need to clear before it could.
