Flock tells WIRED that officers cannot run searches using prohibited attributes including religion and nationality, and that an attempt to search prohibited terms “will be blocked.” Asked what circumstances produce a warning in those two categories instead, the company did not say.
When a T-shirt or a bumper sticker draws a warning because the text includes speech with “constitutional protections,” officers are told the search will be logged and that administrators will be notified. Officers then have to tick an acknowledgement box and leave a comment before they’re free to click “Continue With Search Anyway.”
Flock says the warning leaves room for legitimate work, offering as an example a victim who describes a suspect in a biker gang jacket bearing “a certain gang logo or emblem containing a flag or other insignia.” An officer who proceeds is reported to an administrator at their own department for review.
Kate Ruane, who directs the Center for Democracy and Technology’s free expression project, has spent years studying how automated moderation systems work. Political, social, and cultural expression is “an incredibly amorphous category,” she says, and one reason an officer’s search might contain language including it is to see who attended a protest. That’s already happened. The category worries Flock enough that the system returns a warning, she points out, “but you can still get the results if you just click through.”
“No content moderation done at scale is necessarily accurate,” she says. Analysis of moving video, she says, is harder still, and it goes wrong more often.
The categories that do get blocked can be worked around, Ruane says. The system blocks searches based on religion but not clothing, meaning in practice that an officer could get around the block by searching for the distinctive dress worn by some members of a particular faith. “Lots of people are having their images returned in response to these types of queries that would probably be upset if they knew about it.”
In a search last year an officer in California typed, “American flag.” The search drew a block when aimed at a person, then ran across 11,000 cameras when aimed at vehicles instead.
A warning only stops the officers who are not already determined to run the search, says Deepak Kumar, an assistant professor of computer science and engineering at the University of California San Diego. Kumar, who studies trust and safety systems, points to interfaces built to address online harassment, where motivated users sometimes did more harm after being warned, and to browser alerts meant to steer people away from malicious websites, where the effect depends on how warnings are presented.
“This kind of logging can be useful administratively, to see which officers or end-users bypass warnings often, but that value depends entirely on whether anyone administrates it,” Kumar says. “It functions less as a deterrent and more as a record.”
Reviewing search terms before a query runs will inevitably catch many attempts at misuse, but in isolation, such safeguards are known to fail without equal scrutiny of what the model sends back. “Best practice examines both inputs and outputs,” Kumar says, “which is why so many AI safety efforts now check both to prevent models from generating harmful outputs.”