Session

Beyond the Ban Hammer: Designing a Trust and Reputation System for Automated Moderation

Most moderation systems make one decision about one post: remove it or keep it. That treats a first-time heated reply the same as a coordinated harassment campaign, and it misses the subtle forms of toxicity (dehumanising language, moral manipulation, dog-whistles) that keyword filters and generic toxicity APIs miss. This talk walks through a moderation system that detects around 100 forms of hateful and divisive rhetoric and combines them with a dynamic trust-and-reputation score for each participant. Interventions then match a user's behaviour over time: gentle nudges, reduced reach, or human review.

I'll cover the label taxonomy, how classifier outputs feed the reputation model, how scores recover so people aren't punished forever, how to defend against gaming, and where a human reviewer stays in the loop. You'll leave with an architecture you can adapt for any community platform.

Key takeaways:
1. Why per-post moderation fails and what a time-decayed, per-user reputation model adds
2. Designing a toxicity taxonomy beyond "offensive / not offensive"
3. Human-review gates and anti-gaming measures that hold up in production

Abdul Aleem Khan

AI Enthusiast & Serial Entrepreneur · Applied AI, LLMs & edge computer vision

City of London, United Kingdom

Actions

Please note that Sessionize is not responsible for the accuracy or validity of the data provided by speakers. If you suspect this profile to be fake or spam, please let us know.

Jump to top