Skip to main content
Incident intelligence/SS-IR-022CASE FILE OPEN
Symbolic editorial illustration for SS-IR-022SERVANTSTACK // INCIDENT INTELLIGENCEFORENSIC IMAGE // VERIFIED FRAME
SS-IR-022 // INCIDENT REPORTOfficial finding

Clearview AI

A 3-Billion-Face Database Built by Scraping the Internet, Then Breached

EXECUTIVE BRIEF

Clearview AI quietly assembled a facial-recognition database of more than 3 billion images by scraping photos from Facebook, YouTube, Venmo, LinkedIn, Twitter and the wider web, all without the consent of the people pictured.

FAILURE CHAINTRACE COMPLETE
  1. 01TRIGGERClearview AI quietly assembled a facial-recognition database of more than 3 billion images by scraping photos from…
  2. 02MACHINE ACTIONAutonomous actor
  3. 03MISSING GATENamed SME review and decision audit trail
  4. 04IMPACTData security
01 // INCIDENT SUMMARY

The short version

Clearview AI quietly assembled a facial-recognition database of more than 3 billion images by scraping photos from Facebook, YouTube, Venmo, LinkedIn, Twitter and the wider web, all without the consent of the people pictured.

02 // KEY FACTS

Case telemetry

INCIDENT
SS-IR-022
DATE
February 26, 2020
SYSTEM
Clearview AI
LOCATION / SCOPE
Global (United States / European Union)
EVIDENCE
Official finding
AI ROLE
Autonomous actor
HARM
Data security
SOURCES
3 cited records
03ENTRY POINT // WHAT HAPPENED

The event

Clearview AI quietly assembled a facial-recognition database of more than 3 billion images by scraping photos from Facebook, YouTube, Venmo, LinkedIn, Twitter and the wider web, all without the consent of the people pictured. It sold the resulting "search any face" tool to over 600 law enforcement agencies and private companies. A New York Times investigation exposed the operation on January 18, 2020. Five weeks later, on February 26, 2020, an intruder exploited a flaw and stole Clearview's entire client list, including customer names, the number of accounts each had set up, and how many searches they had run. Regulators across Europe later ruled the scraping unlawful: France's CNIL and Italy's Garante each fined the company EUR 20 million, Greece added EUR 20 million, and the Netherlands imposed EUR 30 million, with multiple orders to delete the biometric data and stop processing it.

04CAUSAL TRACE // AI'S ACTUAL ROLE

What the machine did

The AI was a face-matching engine trained and operated on a dataset whose entire legal and ethical basis was never validated by anyone with the authority to say no. There was no human SME consent-and-lawfulness gate in front of the data ingestion pipeline: the scraper ran autonomously across the open web, vacuuming biometric data at machine scale, and the matching model was shipped to police on the assumption that "public photo" equals "fair game." No data-protection officer, no legal review, and no jurisdictional check stood between the crawler and 3 billion faces. The model worked exactly as built. The problem was that nothing in the build process required a human to confirm the source data was lawful to collect, lawful to retain, or lawful to sell, before it became a product used to identify real people.

Autonomous actorAutomation was a causal participant—not a decorative label for the system around it.
05BLAST RADIUS // CONSEQUENCES

Where the failure landed

Clearview's complete customer list was exfiltrated, exposing which police forces and companies were secretly using face surveillance. Cease-and-desist letters arrived from Facebook, Google, YouTube, Twitter and Venmo for terms-of-service violations. Regulators ruled the company had no lawful basis to process the biometric data of EU residents: CNIL (France) and the Garante (Italy) each levied EUR 20 million, Greece another EUR 20 million, and the Netherlands EUR 30 million, alongside binding orders to delete EU citizens' data and a EUR 100,000-per-day penalty for non-compliance in France. In the U.S., an ACLU lawsuit under Illinois's BIPA forced Clearview to permanently stop selling its database to most private companies. Millions of people had their faces enrolled into a police-grade identification system without ever being asked.

06 // EVIDENCE STATUS

Official finding

Supported by a court, regulator, inquiry, or other official record cited below.

SOURCE RECORD UPDATED 2026-07-09

07 // SOURCE LEDGER

3 cited records

  1. 01
  2. 02
  3. 03
08CONTROL FAILURE // MISSING GOVERNANCE

Named SME review and decision audit trail

The failure pattern in this case: Automated judgment without accountable review.

09INTERVENTION POINT // HUMAN IN THE MIDDLE

The moment the path could change

A qualified reviewer tests the basis, context, and disparate impact before the decision reaches a person.

AI PROPOSESHUMAN OWNS THE DECISIONSYSTEM EXECUTES
10CONTROL DEPLOYMENT // AUTHORITYGATE

SME routing · decision audit trail

The AuthorityGate Operational Resilience framework requires a human SME data-provenance and lawfulness validation gate before any dataset can enter a model-training or production pipeline. A scraping job that ingests biometric identifiers (faces) would be classified as a high-sensitivity data acquisition and blocked from execution until a named Data Protection SME signs off on three things: a documented lawful basis and consent status for each source, a jurisdictional review confirming the collection is legal where the subjects reside, and a retention-and-sale authorization. The same change-validation gate fires again at the point of sale or model release: shipping a biometric product to a new customer class such as law enforcement cannot proceed until a human reviewer validates that the downstream use is sanctioned. Clearview's pipeline ran with none of these checkpoints; an AuthorityGate gate would have halted the crawler at ingestion, demanded the lawful basis nobody could produce, and the 3-billion-face database would never have been built.

RELEVANT KEYSTONE CONTROLHuman-in-the-Loop ValidationHow high-risk actions route to a named subject-matter expert who owns the go or no-go decision.
12 // THE ALTERNATIVE

Autonomy is a design choice.

See the operating model that keeps AI useful while preserving human authority at consequential moments.

Compare AgenticAI and AugmentedAI →