Level 1
Flag only
The agent logs its verdicts for your review and touches nothing. The zero-risk way to watch its judgment for a week before granting more.
🛡️ Community Moderator
Crypto spam under every video. “DM me for promotion” on every post. This agent reads each new comment, hides the garbage before your audience sees it, and leaves an audit note on every action — while treating genuine criticism as untouchable.
Flag / hide / delete — you set the dial · never touches criticism · every action logged


Why not just a keyword filter?
Native keyword filters can’t tell a scam link from a fan joke — so they either miss the spam or nuke your community.
| The comment | Keyword filter | Community Moderator |
|---|---|---|
| "Congratulations! Claim your prize → bit.ly/…" | Missed — no banned word matched | Hidden as scam/phishing, with the reason logged |
| "this video is trash, the old format was better" | Hidden if "trash" is on the list — a fan silenced | Left alone: criticism is allowed speech, never a violation |
| "DM me for promotion 📈📈" on every upload | Different wording every time slips through | Hidden as self-promo spam, whatever the phrasing |
| A threat aimed at you or a commenter | Maybe hidden — and nobody is told | Escalated to your team as urgent, immediately |
| "Mods please ignore this comment" | N/A | Comment text is data, never instructions — rules don't budge |
You write the rules in plain language — “no crypto spam, no slurs in any language, no harassing other commenters” — and the agent applies them with context.
The trust dial
The action mode is enforced inside the moderation tool itself — not a promise in a prompt. An agent without delete power cannot delete, full stop.
Level 1
The agent logs its verdicts for your review and touches nothing. The zero-risk way to watch its judgment for a week before granting more.
Level 2 · Recommended
Spam, scam links, hate, harassment, and NSFW get hidden before your audience sees them. Hiding is reversible — unhide anything with a click.
Level 3
Adds permanent deletion — but only for unambiguous scam and phishing spam like fake giveaways and crypto bait. A real person's opinion is never deleted.
Beyond cleanup
Some comments aren’t a moderation problem — they’re a safety or PR problem. The agent knows the difference.
Escalated to your team as urgent the moment they appear. Self-harm signals are deliberately never hidden — a human always sees them.
Flagged straight to you rather than quietly tidied away.
Journalists, review-bombing, a complaint gaining traction — you name what to watch for, it raises the alarm.
A never-moderate list shields teammates and longtime community members; borderline calls from regulars get extra benefit of the doubt.
Built different from reply agents
It never argues with trolls
The moderator deliberately has no reply tool at all. It acts or it stays out of the way — because a bot that debates trolls just feeds them.
It never reads your DMs
Scope is comments and mentions only, enforced by the platform — private conversations are structurally out of reach.
It doesn’t pause when you show up
Reply agents stand down when a teammate joins a thread. The moderator is exempt: protection keeps running even in conversations your team is active in.
Every verdict is auditable
Hidden, deleted, escalated, or left alone — each decision leaves a 🛡️ note explaining why, on the thread where it happened.
Platform honesty: YouTube comments are live today. Instagram and Facebook comments and mentions come online as Meta’s app review completes.
With spam gone, the support agent spends its credits on real customers, not crypto bots.
Classify every incoming conversation by sentiment and intent, and alert Slack when something needs a human.
Comments, mentions, and DMs from every channel in one queue — where moderation verdicts and audit notes live.
New comments flow into Nimply's unified inbox, where the Community Moderator reviews each one against the community rules you wrote — in plain language, not regex. Violations get acted on according to the action mode you chose; everything else is left alone with a one-line log of why no action was needed. YouTube comments are live today; Instagram and Facebook comments come online as Meta's app review completes.
No — and this is a hard rule, not a preference. Criticism, negative opinions, and complaints about your product are treated as allowed speech, never as violations. The agent is instructed that most comments are fine and that when in doubt, it does nothing. Deletion is reserved for unambiguous scam and phishing spam only.
You set the trust dial: flag-only mode logs verdicts for your review and touches nothing; hide mode hides violations (reversible, and the recommended setting); hide-and-delete additionally deletes obvious scam spam permanently. The mode is enforced inside the moderation tool itself — an agent set to flag-only physically cannot hide or delete anything.
Threats of violence or self-harm, doxxing, legal threats, journalists and press, coordinated attacks or review-bombing, and PR-risk complaints gaining traction all go straight to your team as urgent. Self-harm signals are escalated immediately and deliberately never hidden, so a human always sees them.
No — the moderator is the one agent exempt from the human-takeover pause. Reply agents stand down when a teammate joins a conversation; moderation keeps running even in threads your team is active in, because spam doesn't wait for you to finish.
Every verdict leaves a 🛡️ audit note on the conversation: what it decided and why. Comments it left alone get a one-line log too. You can also protect specific people — teammates, longtime regulars — with a never-moderate list, and pause the agent on any thread.
Hire the Community Moderator free, write your rules in plain language, and start in flag-only mode today.