# The Camera Recorded It. That Doesn't Mean the AI Saw It.

2026-09-13 · Somerset County, New Jersey · Reported Feature

Princeton researchers found that leading AI systems can miss basic events in police body-camera footage. The larger problem is what happens when we begin delegating observation itself.

Princeton researchers found that leading AI systems can miss basic events in police body-camera footage. The larger problem is what happens when we begin delegating observation itself.

---

A body camera can record something without anyone truly seeing it.

That was already one of the strange contradictions built into police body-camera programs. Officers can accumulate enormous amounts of video, creating a record of thousands upon thousands of encounters, but the existence of a record does not mean that record is routinely watched. There is simply too much of it.

Now artificial intelligence is being proposed as part of the answer: use machines to scan the footage, identify important events, flag the moments worth reviewing, and reduce an impossible archive into something humans can actually examine.

The obvious appeal is efficiency. The less obvious problem is that once the machine becomes the filter, the question is no longer simply whether the camera recorded what happened. The question becomes whether the system responsible for finding it noticed it at all.

New research from Princeton Engineering suggests that, at least for now, that is not a safe assumption.

A Record Is Not the Same Thing as a Witness

Princeton researchers built a benchmark called EgoPolice using 185 hours of publicly available police body-camera footage from departments in Illinois, California, Texas and the District of Columbia. Thirty-three students annotated clips at one-second, 10-second and one-minute intervals, marking whether officers performed actions such as running, handcuffing someone, drawing a weapon or providing medical attention.

The team then tested 12 artificial-intelligence models, including Gemini 2.5 Flash, GPT-4.1, Llama-VID and Qwen 2.5VL. The best-performing systems correctly identified what was happening in one-minute clips about 77% of the time. The worst were correct only 11% of the time.

On an ordinary computer-vision benchmark, 77% might sound respectable. Jihoon Chung, a Princeton doctoral student and co-first author of the paper, made the problem plain: this is not an ordinary computer-vision benchmark. Misidentifying a dog breed and failing to notice that an officer drew a weapon are not equivalent mistakes. Chung told Princeton Engineering that even 95% accuracy would still be inadequate for deployment in a high-stakes setting like this.

That distinction matters because percentages have a way of sounding abstract until we ask what the missing percentage contains. Here, the misses can include basic facts that may become central to an investigation, an accountability review or an understanding of how an encounter unfolded.

The Real World Is Bad Training Data

The problem is partly environmental. Much of the video used to train modern vision systems is cleaner than the world those systems are eventually asked to interpret. Police body-camera footage is not a carefully framed image. It is worn on a moving person. It shakes, turns, gets blocked, loses subjects, catches poor lighting and records interactions that unfold quickly and unpredictably.

Olga Russakovsky, an associate professor of computer science at Princeton and a lead researcher on the project, told the university that many AI systems are trained on video that is staged, well-lit and clear. Body-camera footage is often the opposite: grainy, blurry, irregular and fast-moving.

That gap is easy to underestimate because we have become accustomed to AI systems that appear extraordinarily capable in controlled demonstrations. The machine can identify objects, describe photographs and summarize polished video. Then it encounters a human body moving through a chaotic room with an obstructed camera and suddenly the apparent certainty falls apart.

This is not necessarily evidence that computer vision is useless. It is evidence that reality is more difficult than a benchmark designed around tidy examples.

The Archive Became Too Large for Its Own Purpose

There is a second problem underneath the accuracy question, and it may be the more important one. Body cameras were introduced in part to create a durable record of police-civilian encounters and provide greater oversight and transparency. But recording is cheap compared with reviewing. Departments can produce extraordinary amounts of footage, while the human time available to watch it remains limited.

Princeton sociologist Brandon Stewart, a senior author on the paper, described the practical consequence: extreme incidents are likely to receive close attention, but lower-level encounters can simply disappear into the archive because there is no realistic way for people to review everything.

This is where AI becomes genuinely tempting. If a model can reduce thousands of hours of footage to a manageable set of clips, it does not need to replace an investigator to transform the system. It only needs to determine what reaches the investigator.

That is a quieter kind of power.

We tend to think about automated decision-making in terms of machines making final judgments: approve the loan, reject the application, identify the suspect, score the risk. But filtering can be just as consequential. A system that decides what deserves human attention is shaping the universe of evidence from which the human judgment will eventually be made.

Delegating Observation

For years, the most familiar criticism of generative AI has been hallucination: the machine produces a statement that sounds confident but is not true. Computer vision introduces a related problem that feels almost inverted.

Instead of inventing something that did not happen, the system may fail to identify something that did.

That matters because a body-camera recording feels intuitively objective. The event happened. The device captured light and sound. The file exists. We therefore feel as though the truth has been preserved somewhere inside it.

But preservation and access are different things. A library can contain a book nobody opens. A database can hold a fact nobody queries. A camera can capture an event nobody watches.

The addition of AI creates another layer: the event may be recorded, available and technically retrievable, yet still functionally invisible because the system assigned to locate it failed to recognize what it contained.

We are not merely delegating analysis at that point. We are delegating observation.

The Human in the Loop Has to Be More Than a Slogan

The Princeton researchers are careful about this distinction. Their stated goal is not to hand responsibility to an algorithm. They describe a system in which AI helps people sort and search enormous collections of video while human reviewers remain involved, potentially pairing footage with police reports and other information.

Max Gonzalez Saez-Diez, a co-first author of the paper, put an important boundary around the technology. A model might identify that a fight is happening, that a weapon is present or that someone is running away. Determining responsibility, he said, is a human problem rather than a technical one.

That boundary sounds obvious until real systems are built around it. “Human in the loop” can mean a human carefully reviewing machine recommendations. It can also mean a human moving quickly through a queue that the machine has already narrowed, categorized and prioritized. In both cases a person remains present, but the machine may still determine what that person sees first, what appears important and what never enters the queue.

The question, then, is not simply whether a human remains somewhere in the workflow. It is whether the workflow preserves enough independent human attention to catch what the machine misses.

When Seeing Becomes a Systems Problem

There is something almost backwards about the trajectory.

We introduced cameras because human memory, testimony and observation can be incomplete. The camera would create a record. Then the volume of records became too large for humans to examine. So we began looking toward machines to tell us which parts of the record deserve examination.

Each step solves a real problem. Each also creates another layer between the event and the person eventually asked to understand it.

That does not make the technology inherently dangerous or the research pessimistic. In fact, the EgoPolice benchmark exists precisely because better tools could make body-camera archives substantially more useful. Russakovsky told Princeton Engineering that AI could reduce the human review burden by a hundredfold or more. A reliable filtering system could surface patterns and incidents that today are effectively buried.

But usefulness is not the same as infallibility, and the distinction becomes especially important when systems operate upstream of human attention.

The most consequential AI mistakes may not always be the spectacular hallucinations we notice immediately. They may be the quiet omissions: the clip that was never flagged, the action that was never recognized, the piece of recorded reality that existed the entire time but never reached the person looking for it.

The camera recorded it but that does not mean the AI saw it.

SOURCE NOTES

• Princeton Engineering, “AI tools can miss the mark on police bodycam footage,” Sept. 8, 2026 • Gonzalez Saez-Diez et al., “EgoPolice: A Benchmark for Egocentric Video Understanding in High-Stakes Police Body-Worn Camera Footage,” arXiv, submitted July 7, 2026

---

ProbleMattic is written and maintained by Matthew Kulcsar, a software engineer, project manager, technologist, platform builder, emergency-services-trained helper, grandfather, and lifelong collector of broken systems, odd behaviors, and useful nonsense.
