Pre-publish, not post-publish
Most platforms moderate reactively: content goes live, someone reports it, maybe it gets removed. We don't do that. Every message is held in a private queue until it clears review. The public archive only contains messages we've affirmatively approved. This slows down publication — typically 2–6 hours depending on queue depth — but it means we've never had harmful content survive to the public feed because we were too slow to catch it.
This is a deliberate product choice with a real cost. If you call at 2am, your message might not go live until morning. We think that's an acceptable tradeoff. An anonymous platform without pre-publish moderation is not an anonymous platform — it's a harassment infrastructure with a delay timer. We are not building that.
Step one: automated transcription and classification
When your recording lands in our queue, it's automatically transcribed and run through a classification model. The model checks for a set of categorical signals: credible threats, doxxing content, hate speech targeting protected characteristics, CSAM indicators, and a handful of others. This is not sentiment analysis and it is not a vibe check — it is a categorical classifier trained on specific harm patterns.
Messages that score below threshold on all categories are flagged as low-risk and move to a fast-track human review lane, which is shorter. Messages that score above threshold on any category are escalated immediately to a human reviewer. The automated step does not publish or reject anything — it routes. Every final decision is made by a person.
Step two: human review
A human being reads every transcript and listens to every message before it goes live. This is not optional, not something we'll ever automate away, and not something we outsource to a low-wage content farm in a jurisdiction with different legal standards. Right now, that human is us. At scale, it will be a small, well-compensated team with written guidelines, regular calibration sessions, and an explicit right to refuse to review content that affects their wellbeing.
We are aware that content moderation is hard, poorly paid, and damaging to the people who do it when it's done badly. We intend to do it well or not at all. If we reach a volume where we cannot do human review responsibly, we will pause publishing before we will compromise on this.
What gets rejected
- Credible threats of violence against identifiable people or groups
- Doxxing: home addresses, workplaces, or identifying personal information shared without consent
- Hate speech that dehumanizes people based on race, religion, gender, sexuality, disability, or national origin
- Any content involving the sexual exploitation of minors — reported to NCMEC regardless of whether the message is otherwise published
- Targeted harassment of a specific named individual with no broader public interest component
- Fabricated "evidence" of crimes presented as factual and real
These categories are not an exhaustive list. They are the core of it. Reviewers use judgment in edge cases — that's the point of having reviewers.
What doesn't get rejected
Uncomfortable truths. Opinions. Confessions of legal acts. Criticism of institutions, employers, employers' lawyers, or public figures. Anger. Grief. Awkward silences. Political speech in any direction. Religious speech in any direction. Things that make the reviewer uncomfortable personally. Allegations against powerful people that cannot be immediately verified. Messages that are messy, inarticulate, or emotionally raw.
The bar for rejection is harm to an identifiable person or group — not discomfort, not controversy, not the reviewer's personal opinion about the content's merit. We are a platform for speech that is hard to say with your name attached. That means protecting the speech that makes people uncomfortable. If we only publish the safe stuff, we are not a public service. We are a press release machine.
Anonymity is not a license to harm. It's a license to be honest. We will protect the first use of this platform. We will not protect the second.
Notification
If your message is rejected, you'll receive a callback to the number you called from. The callback is automated and will explain the category of rejection — threats, doxxing, hate speech, etc. — without reading back the content of your message. We don't explain what specifically crossed the line, and we don't invite you to revise and resubmit. The message violated our rules and will not be published.
If you called from a number that cannot receive incoming calls (some VoIP lines, some international numbers), the notification will not reach you. We have no other way to reach you — that's the point. Check the archive: if your message isn't there within 24 hours, it was rejected.
Appealing a decision
If you believe your message was wrongly rejected, email hello@psa.xyz with your message ID, which is included in the callback notification. A second human reviewer — not the person who made the original call — will review the message against our written guidelines within 48 hours. If they agree the message should be published, it goes live. If they agree it should be rejected, it stays down.
Decisions are final after the second review. We will not adjudicate indefinitely. If you believe our guidelines themselves are wrong, we want to hear that argument — write to us. If you want us to publish something that violates those guidelines, the answer is no.
Our moderation guidelines are maintained as an internal document reviewed quarterly. If you're a researcher studying content moderation and want more detail, reach out at hello@psa.xyz.