
On January 29, 2025, a 22-year-old student teacher named Kristen Volpe sent a Snapchat message to her boyfriend and two roommates. She was frustrated: a student at John L. Hensey Elementary in Washington, Illinois, had closed her laptop mid-lesson-plan, and she vented about it the way people vent to the three or four humans they trust with their bad moods. It wasn’t a threat. It wasn’t public. It was a private message to a private group chat.
An hour later, sheriff’s deputies were at her school.
Snapchat’s automated content-scanning system had flagged the message, generated an emergency disclosure, and routed it to the FBI, which passed it to the Tazewell County Sheriff’s Office. Deputies searched Volpe, found nothing, and arrested her on a disorderly conduct charge. Bodycam footage shows her trying to explain, mid-arrest, that it was a bad joke. Investigators concluded she made the comment out of exasperation and posed no danger to the school. She was held overnight and released on $0 bond under Illinois’s SAFE-T Act; no formal charges have been filed, and no further court action had been reported as of this writing. But the damage that mattered most had already happened: the school district told parents, and Volpe’s student-teaching placement was over.
That’s the human story, and it’s worth sitting with. But the part I want to walk through is the mechanical one, because it’s the part most coverage of this case skips, and it’s the part that actually explains why this happened and what would need to change to stop it happening again.
What “emergency disclosure” actually means

Snapchat’s safety policy allows the company to voluntarily hand information to law enforcement when it believes there’s “an imminent risk of death or serious physical harm,” without waiting for a subpoena or court order. That policy exists for real reasons. If someone messages a specific, credible plan to hurt themselves or someone else, waiting for paperwork can cost a life. Nobody serious argues platforms should sit on that kind of signal.
The problem in Volpe’s case isn’t that the policy exists. It’s what triggered it, and what didn’t happen before it fired.
Snapchat runs automated scanning across message content looking for language patterns associated with risk: violence, self-harm, threats. When the system’s confidence crosses some threshold, it can generate a report and send it out through the emergency disclosure channel. Volpe’s message apparently used the word “shooting” in the context of her frustration with the laptop-closing student. The system read the word. It did not read the room.
And here’s the detail buried in the public record that matters most: nothing in the documentation indicates a human moderator looked at the message before it went to the FBI. Not a supervisor, not a trust-and-safety analyst, not anyone whose job is to tell the difference between a genuine threat and a 22-year-old typing “I could have shot that kid” to her boyfriend after a bad day. The record shows a flag, then a federal referral, then a sheriff’s response, in under an hour. No documented human review sits between any of those steps until the deputies who arrived in person.
The gap isn’t “AI moderation.” It’s “AI moderation with no checkpoint before consequences get physical.”
It’s worth being precise here, because this story is easy to flatten into either “AI is scary and platforms are watching everything” or “this is a rare glitch, nothing to see.” Both miss the actual, fixable problem.
Automated scanning at Snapchat’s scale isn’t optional. Billions of messages move through the platform; no army of human moderators reads all of them, and nobody wants a platform that only responds to threats a human happened to notice. Language-pattern detection catching a message that mentions violence is the system doing exactly what it’s built to do.
The failure is architectural, not conceptual: there’s no checkpoint between “the model flagged this” and “armed deputies are dispatched.” A flagged message with a risk score above some threshold apparently routes straight through to an external emergency disclosure with no human review step calibrated to the actual stakes of what happens next. Compare that to how flagged financial transactions work, where an algorithm can freeze a card instantly, but a human still reviews the account before anything more serious happens. Compare it to how most credible threat-assessment protocols in schools and workplaces work, where an automated or anonymous tip triggers a human evaluation, not an immediate law enforcement dispatch. The emergency disclosure mechanism skips that step entirely, because it was built for genuine imminent-danger cases where speed is the point and false positives were presumably assumed to be rare enough not to design around.
They aren’t rare enough. A model trained to catch “shooting,” “kill,” “hurt,” and similar words in proximity to a name is going to catch venting, hyperbole, dark jokes, and screenwriting dialogue at a nontrivial rate, because that’s how those words actually get used by non-dangerous people constantly. The system isn’t malfunctioning when it flags Volpe’s message. It’s functioning exactly as designed, and the design has no room in it for “flagged, but almost certainly not real.”
Why this case is a preview, not an outlier
Volpe’s case got attention because it had all the elements that make a story travel: a young teacher, a joke, bodycam footage, a job lost over a misunderstanding a computer made in seconds. But the underlying mechanism, automated content scanning tied to a fast-path law enforcement channel, isn’t unique to Snapchat. Meta, TikTok, and Discord all operate comparable systems, and all of them cooperate with law enforcement and organizations like the National Center for Missing & Exploited Children on flagged content, particularly around child safety, self-harm, and violence.
What made this case visible is that the consequence was immediate and physical: an arrest, in person, at a school, within an hour. Most flagged-content pipelines don’t produce a story this clean, because most of the time either nothing happens, or something happens quietly enough that nobody outside the person affected ever hears about it. That should worry you more than the version where you did hear about it. Volpe’s arrest is evidence the mechanism exists and can misfire. It is very unlikely to be the only time it has.
What would actually fix this
Not “less AI.” Not “no automated scanning,” which isn’t realistic at platform scale and isn’t obviously desirable given the real cases these systems do correctly catch. What’s missing is a calibrated checkpoint: a human review step sized to match the severity of what happens next. Freezing an account, or restricting a message’s reach, is reversible and can happen instantly. Sending armed law enforcement to physically confront someone can’t be walked back once it starts, and it deserves a proportionally higher bar before it fires, ideally a human being who can read tone, context, and the difference between three private recipients and a public threat, in the seconds before deputies get dispatched rather than the hour after.
Platforms building these systems have the data to know their false-positive rate. They should publish it. They should publish what human review, if any, sits between a flag and a law enforcement handoff. Right now, users are told their messages are private and encrypted, and separately, in policy language most people never read, that the platform can hand a message to the FBI without a warrant if its own model decides the stakes are high enough. Those two facts sitting side by side, mostly undisclosed, is the actual scandal here, more than any single false positive.
Kristen Volpe lost a job over a joke a machine couldn’t parse. The fix isn’t pretending the machine shouldn’t exist. It’s making sure a person looks at what the machine flagged before anyone gets arrested for it.
Chris Meredith writes about AI, technology, and what it actually means for real people. Follow along on Substack: monkeyattack.substack.com