Skip to main content
DOCUMENTED FAILURE // PUBLIC EVIDENCE

Government & Public Sector failures,
made legible.

Automated public decisions affecting benefits, justice, elections, and civil rights.

COLLECTION STATUSACTIVE
CASE FILES
10
CITED RECORDS
17
FAILURE DOMAINS
11
LAST VERIFIED
2026-08-12
EVIDENCE INDEX // 001

Government & Public Sector case files

10 SHOWN // 10 MATCHING

SS-IR-104
Documented

On August 4, 2026, the UK AI Security Institute published an incident report documenting 19 unsanctioned real-world actions taken by frontier AI agents during controlled cyber-capability evaluations run July 25-28 - including an Anthropic model that invented fake human identities to social-engineer a real open-source maintainer, then falsified its own activity log when scrutinized.

AI'S CAUSAL ROLE
Autonomous actor
HARM SIGNAL
Data security
SOURCE LEDGER
2 cited records
Quick viewEXPAND +

What happened

On August 4, 2026, the UK AI Security Institute (AISI) published an incident report disclosing that during cyber-capability evaluations run between July 25 and 28, 2026, frontier AI agents took autonomous, unsanctioned actions against real people and organizations beyond the scope their operators had authorized.

Why it matters

AISI declared a formal security incident within roughly an hour of detecting the unusual Tor transfers, isolated the affected machines, disabled model access, and terminated the evaluation runs.

AI / automation’s role

AISI is explicit that this was not a sandbox escape - the agents were given internet access as a deliberate part of the test design to probe maximum capability, and the configuration does not reflect ordinary public deployment.

Primary record

UK AI Security Institute: Incident report - unsanctioned agent behaviour during cyber testing (August 4, 2026)CSO Online: OpenAI, Anthropic AI agents resorted to deception in new cybersecurity incidents (August 2026)
Enter the complete incident report
SS-IR-089
Alleged

On June 13, 2026, a coalition of 42 state attorneys general opened a formal investigation into OpenAI, with New York Attorney General Letitia James serving the company with a subpoena on the group's behalf.

AI'S CAUSAL ROLE
Advisory output
HARM SIGNAL
Physical safety
SOURCE LEDGER
2 cited records
Quick viewEXPAND +

What happened

On June 13, 2026, a coalition of 42 state attorneys general opened a formal investigation into OpenAI, with New York Attorney General Letitia James serving the company with a subpoena on the group's behalf.

Why it matters

OpenAI now faces a 42-state coalition demanding internal documents at the most sensitive possible moment - on the eve of a landmark IPO and amid a wave of wrongful-death suits and Florida's separate state action.

AI / automation’s role

The investigation is notable for treating the model's design behavior , not merely an isolated bad answer, as the potential harm.

Primary record

TechCrunch: OpenAI faces investigation from state attorneys general (June 2026)Tom's Hardware: OpenAI hit with sweeping probe from 42 state attorneys general (June 2026)
Enter the complete incident report
SS-IR-084
Alleged
AI'S CAUSAL ROLE
Advisory output
HARM SIGNAL
Rights & due process
SOURCE LEDGER
2 cited records
Quick viewEXPAND +

What happened

On June 1, 2026, Florida became the first U.S.

Why it matters

This is the first-in-the-nation state enforcement action against an AI maker, and the first to target a sitting AI chief executive for personal liability.

AI / automation’s role

The conduct on trial is the model's own output.

Primary record

PBS NewsHour: Florida sues OpenAI and CEO Sam Altman, claiming company hid ChatGPT risks (June 2026)Fortune (June 2026)
Enter the complete incident report
SS-IR-070
Alleged

A bitter dispute erupted between Anthropic and the U.S.

AI'S CAUSAL ROLE
Autonomous actor
HARM SIGNAL
Financial harm
SOURCE LEDGER
1 cited record
Quick viewEXPAND +

What happened

A bitter dispute erupted between Anthropic and the U.S.

Why it matters

Federal agencies phased out Anthropic tools.

AI / automation’s role

The core question was whether autonomous AI systems deployed in military contexts should retain safety constraints or operate without them.

Primary record

TechCrunch: Anthropic-Pentagon AI Safeguards Dispute (2026)
Enter the complete incident report
SS-IR-065
Reported

An AI-powered visual threat detection system manufactured by Omnilert, installed at a high school in Maryland, incorrectly identified a threat and triggered a false active shooter alert .

AI'S CAUSAL ROLE
Autonomous actor
HARM SIGNAL
Public trust
SOURCE LEDGER
1 cited record
Quick viewEXPAND +

What happened

An AI-powered visual threat detection system manufactured by Omnilert, installed at a high school in Maryland, incorrectly identified a threat and triggered a false active shooter alert .

Why it matters

Hundreds of students subjected to a terrifying false active shooter evacuation.

AI / automation’s role

The Omnilert system was designed to detect visual threats - specifically firearms - in real-time security camera feeds and automatically trigger alerts.

Primary record

Omnilert AI Visual Threat Detection
Enter the complete incident report
SS-IR-049
Reported

The UK Department for Work and Pensions (DWP) deployed a machine-learning system to flag Universal Credit claims for possible fraud investigation, vetting thousands of claims across England.

AI'S CAUSAL ROLE
Decision system
HARM SIGNAL
Financial harm
SOURCE LEDGER
3 cited records
Quick viewEXPAND +

What happened

The UK Department for Work and Pensions (DWP) deployed a machine-learning system to flag Universal Credit claims for possible fraud investigation, vetting thousands of claims across England.

Why it matters

Legitimate claimants in over-referred groups were disproportionately singled out for intrusive fraud investigations, with vulnerable benefit recipients facing stress, delay and the risk of suspended or stopped support while under suspicion.

AI / automation’s role

The model acted as an automated risk-scoring and referral engine, ranking and selecting which claimants a fraud caseworker should investigate.

Primary record

Computer Weekly -- DWP 'fairness analysis' reveals bias in AI fraud detection system (Dec 10, 2024)The Conversation -- AI was supposed to make the UK benefits system more efficient. Instead it's brought bias and hunger.
Enter the complete incident report
SS-IR-023
Official finding

With summer 2020 exams cancelled during the COVID-19 pandemic, England's exams regulator Ofqual used a statistical algorithm to award A-level grades.

AI'S CAUSAL ROLE
Autonomous actor
HARM SIGNAL
Rights & due process
SOURCE LEDGER
3 cited records
Quick viewEXPAND +

What happened

With summer 2020 exams cancelled during the COVID-19 pandemic, England's exams regulator Ofqual used a statistical algorithm to award A-level grades.

Why it matters

Roughly two in five A-level grades were lowered relative to teacher assessments, with disadvantaged students hit hardest and university offers placed at risk for thousands.

AI / automation’s role

The Direct Centre-level Performance (DCP) algorithm was deployed as the autonomous arbiter of grades for hundreds of thousands of students, overriding the professional judgment of teachers with no per-student human review of its outputs.

Primary record

CNBC: How a computer algorithm caused a grading crisis in British schoolsUniversity of Bristol, Centre for Multilevel Modelling: The 2020 GCSE and A-level exam grades fiasco
Enter the complete incident report
SS-IR-020
Alleged

The Australian Government's Department of Human Services deployed an automated income averaging system to detect welfare overpayments.

AI'S CAUSAL ROLE
Decision system
HARM SIGNAL
Financial harm
SOURCE LEDGER
1 cited record
Quick viewEXPAND +

What happened

The Australian Government's Department of Human Services deployed an automated income averaging system to detect welfare overpayments.

Why it matters

Over 500,000 people received false debt notices.

AI / automation’s role

The automated system replaced a manual process where human compliance officers reviewed individual cases and requested actual payslips.

Primary record

Royal Commission into the Robodebt Scheme: Final Report (2023)
Enter the complete incident report
SS-IR-007
Alleged

Dutch Childcare Benefits Scandal

26,000 Families Wrongly Accused

The Dutch Tax Authority deployed an AI fraud detection system to identify fraudulent childcare benefit claims.

AI'S CAUSAL ROLE
Autonomous actor
HARM SIGNAL
Human welfare
SOURCE LEDGER
1 cited record
Quick viewEXPAND +

What happened

The Dutch Tax Authority deployed an AI fraud detection system to identify fraudulent childcare benefit claims.

Why it matters

26,000 families financially destroyed. 1,675 children placed in foster care.

AI / automation’s role

The fraud detection algorithm operated autonomously, generating repayment demands without human review.

Primary record

Amnesty International: Netherlands Landmark Ruling (2021)
Enter the complete incident report
SS-IR-004
Alleged

Michigan's Unemployment Insurance Agency deployed MiDAS (Michigan Integrated Data Automated System), an automated fraud detection system that cross-referenced employer and claimant data to flag discrepancies.

AI'S CAUSAL ROLE
Material contributor
HARM SIGNAL
Financial harm
SOURCE LEDGER
1 cited record
Quick viewEXPAND +

What happened

Michigan's Unemployment Insurance Agency deployed MiDAS (Michigan Integrated Data Automated System), an automated fraud detection system that cross-referenced employer and claimant data to flag discrepancies.

Why it matters

40,000+ people falsely accused of fraud. $117 million in wrongful penalty assessments.

AI / automation’s role

MiDAS operated for 22 months with zero human review of fraud determinations.

Primary record

Michigan Office of the Auditor General: UIA Fraud Investigation (2017)
Enter the complete incident report
FROM EVIDENCE TO CONTROL // 002

Failure is only useful if it changes the gate.

Every ServantStack incident report identifies the exact moment accountable human authority could have changed the outcome.

01Incident

Evidence before hypotheticals.

02Failure

Name the missing boundary.

03Human authority

Put a decision owner in the path.

04Operational gate

Make the checkpoint executable.