Skip to main content
A vast machine evidence vault sends controlled green and amber signal paths through a central mechanical authority gate while a red path remains contained.
OPERATING KNOWLEDGE // 001

The AI governance field manual.

A visual guide to the moment an AI recommendation becomes a real-world consequence—and the controls that keep speed from becoming unaccountable execution.

MANUAL PURPOSE

Translate standards, incident evidence, and operational controls into decisions a team can actually use.

FOUNDATIONAL CHAPTERS08
INCIDENT REPORTS ANALYZED104
PRIMARY FRAMEWORKS05
ORIENTATION // 01

Govern the action, not the adjective.

“AI,” “agent,” and “copilot” do not tell you how much authority a system holds. Start with the action, the consequence, and the person accountable for allowing it.

AI governance is an operating system for accountability. It connects named decision owners, context-specific evidence, technical boundaries, monitoring, intervention, and recovery across the AI lifecycle. NIST places governance across its Govern, Map, Measure, and Manage functions and explicitly calls for defined human-oversight processes.[1]

01 // PROPOSEMachine recommends

The system produces an answer, decision, tool call, or change.

02 // CLASSIFYConsequence is scored

Reversibility, reach, rights, privileges, and uncertainty are assessed.

03 // PROVEEvidence is assembled

Context, dependencies, telemetry, test results, and rollback state travel with the action.

04 // AUTHORIZEA named owner decides

The reviewer can approve, reject, narrow, delay, or escalate.

05 // VERIFYOutcome is observed

Deployment is staged, measured, contained, and recovered when necessary.

FOUNDATIONAL CHAPTERS // 02

Eight concepts. One control path.

Read in sequence for the complete model, or enter at the pressure point your organization needs to solve.

FM-01 // AUTONOMY

Agentic AI

Systems that pursue goals through planning, tool use, and actions with some degree of autonomy. Autonomy is a capability—not permission.

Read the chapter →
FM-02 // GOVERNED AUTONOMY

Augmented AI

An operating pattern in which AI capability is paired with accountable human judgment at consequential checkpoints.

Read the chapter →
FM-03 // ACCOUNTABILITY

Human in the Middle

A qualified decision owner placed at the point where machine output could become material consequence.

Read the chapter →
FM-04 // EXECUTION

Authority Gates

Executable policy that decides which actions may proceed automatically and which require named approval.

Read the chapter →
FM-05 // VALIDATION

AI Change Validation

Testing an action against your real dependencies, risk tolerances, evidence, and recovery objectives before broad deployment.

Read the chapter →
FM-06 // CONTAINMENT

Blast-Radius Control

Limit how far one wrong action can travel through systems, customers, environments, or decisions.

Read the chapter →
FM-07 // RECOVERY

Known-Good State & Rollback

Prove that restoration works inside the required time—and restore only what the dependency map says was affected.

Read the chapter →
FM-08 // OUTCOME

Operational Resilience

The ability to continue delivering important outcomes through disruption, intervention, and recovery.

Read the chapter →
INCIDENT EVIDENCE // 03

What the ledger contains.

These charts describe the ServantStack incident ledger as verified on August 12, 2026. They do not estimate the prevalence of AI failures across industry.

AI or automation’s causal role

Count of reports by the editorial role assigned in the ledger (n=104).

Autonomous actor47
Advisory output18
Decision system11
Fraud enabler10
Material contributor10
Operational automation8
Accessible data for causal-role chart.
RoleReports
Autonomous actor47
Advisory output18
Decision system11
Fraud enabler10
Material contributor10
Operational automation8

Primary harm signal

Each report receives one primary harm signal; a report can involve additional harms.

Data security29
Financial harm24
Physical safety19
Human welfare14
Rights & due process10
Public trust6
Operational disruption2
Accessible data for primary-harm chart.
Harm signalReports
Data security29
Financial harm24
Physical safety19
Human welfare14
Rights & due process10
Public trust6
Operational disruption2

Method note. Counts are generated from data/incidents.json, the source dataset for ServantStack’s canonical incident pages. Categories are editorial classifications based on cited records. Selection, documentation availability, recency, and publication bias make this dataset unsuitable for calculating population rates or comparative product safety. Inspect the machine-readable ledger.

OPERATING MODEL // 04

Oversight must change the path.

The UK ICO warns that merely inserting a human somewhere in a lifecycle does not create meaningful review; timing and actual agency over the final outcome matter.[5]

A red autonomous instruction enters an amber mechanical evidence gate; a green controlled instruction exits toward a production fleet.
EXECUTION GATE // CONTROL SURFACEA gate is effective only when it receives evidence, enforces policy, records the decision, and can stop the action before consequence.
01 // OWNER

Named authority

Assign the person or role accountable for the decision—not merely the team that operates the tool.

  • Competent in the affected domain
  • Available inside the decision window
  • Empowered to reject or narrow scope
02 // EVIDENCE

Decision packet

Make the reviewer’s evidence travel with the proposed action, including uncertainty and known omissions.

  • Intended outcome and scope
  • Dependencies and blast radius
  • Tests, telemetry, and conflicts
03 // BOUNDARY

Least privilege

Give the system only the access and duration required for the authorized action.

  • Scoped credentials
  • Separated trust zones
  • Short-lived execution authority
04 // VALIDATE

Real context

Validate against your business services and dependencies—not only a vendor’s generic sandbox.

  • Canary or staged cohort
  • Per-system health checks
  • Business outcome signals
05 // CONTAIN

Limited blast radius

Prevent one wrong action from automatically becoming an enterprise-wide event.

  • Rate and scope limits
  • Progressive deployment
  • Automatic halt conditions
06 // RECOVER

Known-good restoration

Prove recovery before execution and validate restoration after a fault.

  • Dependency-aware rollback
  • Tested recovery objective
  • Post-restore verification
A production fleet remains green while one red failed machine is isolated beside an amber validation cell and a blue known-good recovery archive.
RESILIENCE // SELECTIVE RECOVERYA fault in one validation cohort should stop propagation, restore the affected component, and leave unrelated validated systems operating.
DECISION GUIDE // 05

When should a human decide?

Not every output needs approval. The useful question is whether the action can create a consequence the system should not be allowed to authorize for itself.

SignalExampleGateMinimum evidence
Hard to reverseDelete data, publish externally, terminate serviceHUMAN REQUIREDBackup state, dependency map, rollback test, named owner
Rights, safety, or livelihoodHiring, benefits, medical triage, physical controlHUMAN REQUIREDQualified reviewer, explanation, affected-person recourse, audit trail
Crosses a trust boundaryUntrusted content reaches code, credentials, or private dataHUMAN REQUIREDSource provenance, scoped privileges, output inspection
Uncertain or novelNew failure mode, low-confidence recommendation, edge caseESCALATE BY RISKConfidence limits, comparable cases, expert review criteria
Reversible and boundedDraft, summarize, simulate, recommend without executionAUTOMATE WITH MONITORINGLogging, sampling, clear non-execution boundary
Repeated low-risk actionPreviously approved pattern inside a fixed scopePOLICY-BASED AUTO-APPROVALVersioned policy, scope limit, drift monitoring, kill switch
SOURCE LEDGER // 06

Standards behind the controls.

The manual synthesizes these sources into an operational model; it does not claim that “AugmentedAI,” “Human in the Middle,” or “Authority Gate” are terms defined by these institutions.

  1. NIST AI Risk Management Framework Core. Govern, Map, Measure, and Manage; includes defined human-oversight processes (MAP 3.5), testing before deployment and during operation, and decisions about whether deployment should proceed.
  2. NIST SP 800-53 Revision 5. A catalog of security and privacy controls supporting trustworthy and resilient systems.
  3. CISA, Shifting the Balance of Cybersecurity Risk. Secure-by-design principles emphasize customer security outcomes, transparency, and accountability.
  4. Regulation (EU) 2024/1689, Article 14. High-risk AI systems must support effective human oversight, including competence, authority, intervention, and stop mechanisms appropriate to risk.
  5. UK Information Commissioner’s Office, Guidance on AI and Data Protection. Explains why nominal human involvement is not necessarily meaningful review and why sequencing matters.
  6. OECD AI Principle: Robustness, Security and Safety. Calls for lifecycle risk management and mechanisms to override, repair, or safely decommission systems exhibiting harmful behavior.
  7. NIST AI 600-1, Generative AI Profile. Cross-sector companion to the AI RMF for generative-AI risks.
  8. ServantStack Incident Intelligence. A curated, source-cited editorial ledger of 104 AI and automation incidents, last verified per report.
DIRECT ANSWERS // 07

Questions teams ask first.

Short answers for orientation. Each linked chapter includes the nuance, examples, controls, and source trail.

What is AI governance?

The accountable roles, policies, evidence, technical controls, and review processes used to decide whether and how an AI system may act throughout its lifecycle.

Is agentic AI the same as generative AI?

No. Generative AI describes systems that generate content. Agentic AI describes a pattern in which a system pursues goals through planning, tools, and actions. A system may be both, either, or neither.

When is human approval necessary?

When an action is hard to reverse, crosses a trust boundary, affects rights or safety, expands privileges, moves money, changes production, or has a large and uncertain blast radius.

Does a person clicking “approve” create oversight?

No. The reviewer needs relevant competence, sufficient evidence and time, genuine authority to intervene, and a checkpoint before the consequential action.

Can low-risk actions remain autonomous?

Yes. Bounded, reversible actions can run automatically when policy defines their scope and monitoring, logging, and halt conditions remain active.

Are the incident graphs industry statistics?

No. They describe ServantStack’s curated 104-report ledger. The dataset is useful for studying patterns and controls, but it is not a representative census of all AI failures.

NEXT MOVE // 08

Move from the manual to the evidence.

See how the controls map to real incidents, or compare ungoverned AgenticAI with the validated AugmentedAI path.