How Accurate Is Facial Recognition, Really?

TL;DR

Top algorithms score above 99% in NIST's mugshot-vs-mugshot tests. That headline number hides three things that matter more than the average: (1) accuracy drops sharply when photos are taken off-angle or years apart, (2) false-match rates are not equal across faces, and NIST's own report describes a factor of 100 difference in false positives between the best- and worst-performing countries of origin for some algorithms, and (3) real-world deployments use a different camera, a different threshold, and a different subject than the lab. NIST publishes the underlying rates for every algorithm it tests.

What "99% accurate" actually means

The 99% number comes from NIST's Face Recognition Vendor Test (FRVT), the most-cited independent benchmark for the technology. NIST does not sell software, it does not pick winners for procurement, and it publishes the raw numbers for every submitted algorithm. Two facts from NIST's April 21, 2023 results, as summarised by the Bipartisan Policy Center, make the headline number less reassuring than it sounds.

First, 99% applies to constrained, cooperative photos. NIST's April 21, 2023 results showed 45 of 105 identification algorithms exceeded 99% accuracy comparing mugshot-quality probes against a gallery of 1.6 million mugshot templates. Only 3 algorithms kept above 99% accuracy when the gallery expanded to 3 million templates with twelve or more years between captures. And 3 algorithms got above 90% with 90-degree rotation off-angle probe images.[1]

Second, the threshold matters as much as the algorithm. NIST reports accuracy at fixed false-positive rates. For verification, the April 21, 2023 results held the false-match rate at FMR thresholds of 10^-6 to 10^-5 (one in a million to one in 100,000), and 233 of 503 algorithms exceeded 99% accuracy matching constrained cooperative probes against visa photos. Zero algorithms hit 99% when the probes were unconstrained and uncooperative matched against kiosk photos.[1] Push the threshold to favor catching bad guys, and false matches rise. Push the threshold the other way, and you miss real matches.

For one-to-many identification, the April 21, 2023 NIST runs measured false-negative identification rate (FNIR) while limiting false-positive identification rate (FPIR) to 0.003, or three in a thousand.[1] That is the setting NIST reports at; a deployment that tunes its threshold differently sees different error rates in both directions.

Where accuracy breaks: off-angle, aged, uncooperative

Lab accuracy is not field accuracy, and the gap is what gets innocent people arrested.

Three operational conditions account for most of the degradation NIST measures:

  • Off-angle or partial faces. A driver captured in profile through a car windshield, a face turned at a protest, a CCTV image of someone looking down. NIST's April 21, 2023 results showed only 3 of the top identification algorithms achieving above 90% accuracy with 90-degree rotation probe images.[1]
  • Time between probe and gallery photo. Aging changes faces. NIST's April 21, 2023 identification results showed that 45 of 105 algorithms held above 99% on a 1.6 million-template gallery, but only 3 kept above 99% when the gallery expanded to 3 million with 12+ years between captures.[1]
  • Unconstrained, uncooperative images. Airport checkpoint photos are cooperative and constrained. Cell-phone footage of a shoplifter, body-cam footage of a suspect, or a stadium CCTV still is unconstrained and uncooperative. NIST's April 21, 2023 verification results showed zero algorithms hitting above 99% accuracy at any tested threshold on unconstrained, uncooperative probes matched against kiosk photos.[1]

This is the gap vendors do not put in their sales decks. A police chief hears "99%" and assumes the technology will reliably identify a person of interest caught on a distant, off-angle camera three years after the booking photo. NIST's own numbers say it usually will not.

The demographic gap: same algorithm, different error rates

NIST's Face Recognition Vendor Test Part 3: Demographic Effects, published as NIST IR 8280 in December 2019, is the most-cited independent study on demographic performance. It ran 189 algorithms from 99 developers against four photograph collections holding 18.27 million images of 8.49 million people, drawn from State Department, Department of Homeland Security and FBI operational databases.[2]

The report's own executive summary carries the headline finding. On one-to-one matching with application photos, "false positive rates are highest in West and East African and East Asian people, and lowest in Eastern European individuals." False positives were "higher in women than men," an effect the report calls "consistent across algorithms and datasets" but "smaller than that due to race," and it found "elevated false positives in the elderly and in children." At the same time, "some developers supplied highly accurate identification algorithms for which false positive differentials are undetectable."[6]

The ratio is the part that matters. NIST's press release put the one-to-one false-positive differential for Asian and African American faces relative to Caucasian faces at "a factor of 10 to 100 times, depending on the individual algorithm."[2] The report describes the country-of-origin effect as "generally large, with a factor of 100 more false positives between countries," while noting that for a number of algorithms developed in China the effect is reversed.[6] A small fractional gap in a lab table multiplies into very different outcomes when a system runs millions of comparisons.

The U.S. Government Accountability Office reached the same conclusion in GAO-20-522, its July 2020 report on commercial uses of the technology, summarizing NIST's testing: "facial recognition technology generally performs better on lighter-skin men and worse on darker-skin women, and does not perform as well on children and elderly adults." GAO also noted the operational consequence: "These differences could result in more frequent misidentification for certain demographics, such as misidentifying a shopper as a shoplifter when comparing the individual's image against a data set of known shoplifters."[3]

GAO was careful to add the caveat the lab data deserves: "There is no consensus on what causes performance differences, including physical factors (such as lighting) or factors related to the creation or operation of the technology."[3] The honest reading is that top algorithms have closed most of the demographic gap, but the long tail of deployed systems has not.

Real-world performance: the DHS 2022 Biometric Technology Rally

NIST runs controlled tests against photograph collections. The Department of Homeland Security's Science and Technology Directorate, working through the Maryland Test Facility (MdTF), ran a parallel real-world test in 2022 that is closer to what a border checkpoint or a stadium entry actually looks like: 575 volunteers diverse in race, gender, age, and skin tone walked past 4 acquisition systems paired with 10 matching systems.[4]

The MdTF 2022 Rally reported results using True Identification Rate (TIR), the share of returning volunteers correctly identified as people who had enrolled. The acquisition-system ranking on overall TIR for groups of two walked through in sequence was: Bison at 97.4%, Longs at 96.5%, Wilson at 93.2%, and Borah at 74.1%. One matching system ("Row") consistently scored 62-85% across demographics.[4]

The demographic numbers from the same rally are the ones that travel. In groups of two, 17 of the 40 system combinations met the 95% TIR threshold for female volunteers, 17 for male, 17 for Asian, 14 for Black, and only 5 for White at a stricter 99% TIR goal.[4] For darker-skinned volunteers, 9 of 40 combinations cleared 95% TIR, against 17 for lighter-skinned volunteers at the same threshold. The top system hit 98.4% TIR on darker-skinned volunteers and 96.3% on lighter-skinned volunteers, inverting the typical pattern but at the cost of more total failures for darker-skinned subjects.[4]

The 2021 Rally, run on the same MdTF infrastructure, reported that 26 of 50 system combinations achieved above 99% matching-TIR without masks, but only three did so with masks on. Without masks, median performance was lower for female volunteers, volunteers who self-identified as "Black or African-American," and volunteers "with relatively darker skin tones," even though, as BPC notes, "Some system combinations were able to meet the 95% Rally TIR threshold for all demographic groups."[1]

What "false positive" means in a police workflow

Two different errors get called "inaccuracy." A false negative is when the system fails to match a probe to a real entry in the gallery (a guilty person walks free). A false positive is when the system matches a probe to the wrong entry in the gallery (an innocent person gets flagged). NIST measures both. Police deployments care about both, but for the citizen on the street, only one is dangerous.

The false arrests of three Black men, Nijeer Parks, Robert Williams and Michael Oliver, in investigations that involved face recognition are the cases most often cited against the technology. The Bipartisan Policy Center's review adds a caution the headlines usually drop: not all the details of the role the software played have been made public, and Detroit's police chief attributed Williams's arrest to "sloppy, sloppy investigative work" rather than to a face recognition failure.[1]

Both readings lead to the same place. A candidate list is a lead, not an identification, and a false positive only becomes an arrest when the people downstream treat it as one. The error is rare per match. At the scale of a real police system, running thousands of searches against a gallery of millions, the long tail of bad matches is a regular feature rather than an edge case, which is why the human review step carries the weight.

How to read a face recognition vendor's claim

Three questions cut through vendor marketing:

  1. What is the false-positive rate at the threshold they are quoting? A 99% true-positive rate sounds great until you ask what false-positive rate comes with it. A vendor quoting 99% accuracy at an unstated threshold is hiding the operating point. Ask for the FRVT submission and read the curve.
  2. What is the probe-to-gallery photo gap? Booking photos are constrained, frontal, and recent. Cell-phone footage is none of those. NIST publishes accuracy by probe type; in the April 2023 verification results 233 of 503 algorithms cleared 99% on constrained, cooperative probes and none did on unconstrained, uncooperative ones.[1]
  3. What is the demographic breakdown? The most recent NIST demographic testing is the only reliable place to find this. The vendor will not volunteer it. If the vendor will not share the demographic breakdown, assume it is the lower-performing end of the curve.

For police use, two additional questions matter: how big is the watch list gallery, and what is the threshold the system uses in production. An FPIR of 0.003 means three false positives for every thousand searches of people who are not in the gallery, or 3,000 per million, before any human review. A bigger gallery makes the task harder; it does not change that arithmetic. That is the math NIST runs, and it is the figure to ask a vendor for.

What this means for you

The technology is real and improving. It is also uneven in ways that matter if your face is on the wrong end of a match. Three rules of thumb from the NIST and DHS data:

  • If the photo of you is good (passport, ID, frontal, recent) and the gallery photo is good (the same), top algorithms are reliably above 99%. This is the case for unlocking your phone, where the match is against an enrollment you control.
  • If the photo of you is bad (CCTV, off-angle, masked, low light) and the gallery photo is good, top algorithms fall well short of the headline figure: in NIST's April 2023 runs only three identification algorithms stayed above 90% on 90-degree off-angle probes.[1] This is the shape of most police deployments.
  • If the gallery is large (millions of faces) and the probe is bad, even top algorithms produce false matches at a rate that turns into wrongful arrests at scale. NIST reports this. Vendors do not.

On December 19, 2023 the Federal Trade Commission announced an order against Rite Aid, saying the retailer's facial recognition technology "falsely tagged consumers, particularly women and people of color, as shoplifters." The order bans Rite Aid from using facial recognition surveillance for five years and requires it to delete, and direct third parties to delete, the images its system collected.[5] It is the clearest regulatory statement so far that the number the public is exposed to is the operational false-match rate, not the lab benchmark.

Sources

  1. Bipartisan Policy Center: Face Recognition Technology Accuracy and Performance (May 24, 2023), summarising NIST FRVT 1:1 verification and 1:N identification results from April 21, 2023, the 2021 and 2022 DHS Biometric Technology Rallies, and the three false-arrest cases.
  2. NIST: NIST Study Evaluates Effects of Race, Age, Sex on Face Recognition Software (news release, December 19, 2019, on NISTIR 8280).
  3. U.S. Government Accountability Office: GAO-20-522, Facial Recognition Technology: Privacy and Accuracy Issues Related to Commercial Uses (published July 13, 2020, publicly released August 11, 2020). Direct quotes on demographic performance gaps: "facial recognition technology generally performs better on lighter-skin men and worse on darker-skin women, and does not perform as well on children and elderly adults" and "These differences could result in more frequent misidentification for certain demographics, such as misidentifying a shopper as a shoplifter."
  4. Maryland Test Facility: 2022 Biometric Technology Rally Results, run by DHS S&T at MdTF in 2022 with 575 volunteers across 4 acquisition systems and 10 matching systems. Acquisition ranking (overall TIR, groups of 2): Bison 97.4%, Longs 96.5%, Wilson 93.2%, Borah 74.1%. Demographic breakdowns at the 95% and 99% TIR thresholds as quoted in the article.
  5. U.S. Federal Trade Commission: Rite Aid Banned from Using AI Facial Recognition After FTC Says Retailer Deployed Technology without Reasonable Safeguards (press release, December 19, 2023).
  6. NIST: NISTIR 8280, Face Recognition Vendor Test (FRVT) Part 3: Demographic Effects (December 2019, PDF; executive summary quoted).