content-moderation

3 posts

pinterest

How Pinterest Built a Real‑Time Radar for Violative Content using AI (opens in new tab)

Pinterest built an AI-assisted prevalence measurement system to estimate how often users actually see policy-violating content, rather than relying only on user reports. The system samples daily impressions, uses production risk scores to improve efficiency, labels content with a multimodal LLM, and applies statistical reweighting to preserve unbiased estimates. This enables daily, segmented monitoring with substantially lower cost and latency than human-only review. ## Why Prevalence Matters - User reports miss important harms because: - Some sensitive issues, such as self-harm, are under-reported. - Users seeking harmful content may not report it. - Rare policy categories provide too few reports for reliable trend detection. - Human review of reports is expensive and slow. - Prevalence measures exposure: the share of total views directed to violating content. - This helps Pinterest identify under-reported harms, evaluate interventions, and detect changes earlier. - Human-only prevalence studies were previously conducted only about every six months and required multiple reviewers plus adjudication. ## What Pinterest Measures - Daily prevalence is calculated as: - **Views of content violating a policy ÷ total views** - For example, 10 violating views in a sample of 100,000 produces an estimated prevalence of 0.01%. - Results include 95% confidence intervals to communicate statistical precision. - Metrics can be segmented by: - Policy area, such as Adult Content, Self-harm, or Graphic Violence - Sub-policy, such as nudity versus explicit sexual content - Surface, including Homefeed, Search, and Related Pins - Content age, geography, and user-age groups where relevant ## Risk-Aware, Unbiased Sampling - Pinterest samples from the daily user-impressions stream. - Production enforcement risk scores are used to prioritize likely high-risk and high-exposure content, but they are not treated as labels or eligibility rules. - Missing scores are replaced with the day’s median so that new content remains eligible. - Weighted reservoir sampling approximates probability-proportional-to-size sampling, considering impressions and risk scores. - Inverse-probability weighting removes the bias introduced by risk-based sampling, ensuring estimates represent impressions rather than model thresholds. - Pinterest uses Hansen–Hurwitz ratio estimators for sampling with replacement and Horvitz–Thompson ratio estimators for sampling without replacement. - Pure random sampling is also available for validation studies. ## LLM-Based Labeling - A multimodal LLM analyzes sampled content using both images and text. - Prompts are reviewed by policy subject-matter experts and can return structured label hierarchies such as `safe`, `not_safe`, and `unsure`. - Each decision records: - The label and brief rationale - Policy version - Prompt and model identifiers - Token usage and run cost - Human validation is performed on strategically selected samples to identify edge cases and AI blind spots. - The LLM is tested against human-reviewed gold sets before launch and periodically afterward to detect drift. - The workflow is reportedly 15 times faster and far cheaper than human-only labeling while maintaining comparable decision quality and statistical governance. ## Production System and Monitoring - Inputs include entity-by-day engagement data such as impressions, clicks, hides, and reports, alongside current production risk scores. - The system stores prevalence estimates, sampling weights, labels, diagnostics, and lineage for audits. - Dashboards display: - Daily prevalence and 95% confidence intervals - Confidence-interval width and effective sample size - Sample positive rate - Risk-score distributions - Prompt, model, taxonomy, and metric versions - Teams can pivot results by policy, sub-policy, and surface. - Validation samples and run-health information help monitor both statistical quality and operational reliability. Pinterest’s approach combines probability sampling, inverse-probability estimation, and continuously calibrated multimodal AI labeling to create a daily radar for harmful exposure. The practical recommendation is to use AI to scale measurement, but retain rigorous sampling, human validation, confidence intervals, and full model and policy lineage so that faster estimates remain trustworthy.

discord

ROOST Announces “Coop” and “Osprey”: Free, Open-Source Trust and Safety Infrastructure for the AI Era (opens in new tab)

ROOST, a non-profit dedicated to digital safety, has launched two open-source tools, Coop and Osprey, to provide enterprise-grade content moderation and threat investigation capabilities to organizations of all sizes. By open-sourcing technology previously developed by industry leaders like Discord and Cove, ROOST aims to democratize access to the infrastructure required to detect, triage, and respond to online harms. This initiative shifts Trust and Safety from a proprietary competitive advantage to a shared public resource, enabling platforms to prioritize user protection without the burden of expensive enterprise software. ### Content Review and Compliance with Coop Built on technology acquired from Cove and utilized by platforms like Notion, Coop focuses on the human-in-the-loop aspect of content moderation. * The platform provides robust tools for content review, allowing teams to route specific cases to subject-matter experts for deeper analysis. * It includes built-in integration with the National Center for Missing & Exploited Children’s (NCMEC) API, automating the mandatory reporting process for child sexual abuse material (CSAM). * The interface is designed to surface relevant context and metadata, ensuring moderators can make informed decisions and take immediate action against policy violations. ### Incident Response and Investigation with Osprey Osprey is a lightweight investigation tool originally developed by Discord to manage large-scale safety incidents and platform-wide threats. * It serves as a foundation for incident response, helping safety teams understand platform trends and investigate coordinated threats like phishing or harassment campaigns. * The tool is designed to be user-friendly and accessible for grassroots communities while remaining powerful enough for established platforms. * Early adopters, including the decentralized social network Bluesky, are implementing Osprey to demonstrate that effective safety infrastructure can be scalable and resource-efficient. ### A Collaborative Model for Safety Infrastructure The launch of these tools represents a strategic shift toward a collaborative "public-interest" model for digital defense. * ROOST acquired the intellectual property of Cove and received the donation of Osprey from Discord to ensure these tools remain available as a public good. * The initiative is backed by philanthropic funding and legal support from Perkins Coie, removing the financial barriers that often prevent smaller platforms from implementing high-level safety measures. * Major industry players like Notion and Bluesky are championing the move, signaling an industry-wide push to share safety innovations rather than silo them. Platforms and developers should prepare to integrate these tools into their safety stacks as they become publicly available in the coming months. By adopting open-source infrastructure for routine tasks like NCMEC reporting and incident triage, organizations can focus their internal resources on platform-specific innovations while maintaining a high standard of digital safety.

discord

Mental Health & Discord: Promoting Well-Being All Year Long (opens in new tab)

Discord’s Mental Health team describes its efforts to make the platform safer, more user-controlled, and better connected to mental-health support. The company highlights new research, crisis-resource partnerships, and initiatives exploring how gaming can support social connection and well-being. It concludes that Discord intends to use these findings to improve product design, policies, and partnerships. ## Discord’s Approach to Well-Being - Discord aims to provide a safe, welcoming place for people to connect around shared interests and gaming. - The platform emphasizes user control, allowing people to choose who they interact with and which communities they join. - Discord says it is not designed primarily to maximize engagement, but to support real-time interaction among friends. - The company plans to continue refining its products and safety efforts to better support users. ## Research on Users’ Mental Health - Over five months, Discord worked with research organization BSG to study users’ mental-health perceptions and the platform features that affect well-being. - The research included more than 1,200 participants aged 13 and older in the United States, using both qualitative and quantitative methods. - Participants generally viewed online gaming and social media as distinct, with gaming considered more positive for mental health. - Online gaming was the most frequently selected activity for supporting well-being, followed by exercise, in-person relationships, creative activities, and online connections with family and friends. - Users identified moderators, blocking, muting, and reporting as Discord features that support their mental health. - Participants expressed interest in more proactive mental-health tools and believed Discord cares about user well-being, while also acknowledging that more work is needed. - Discord intends to apply the findings to its policies, product development, and future partnerships. ## Crisis and Mental-Health Resources - Discord partnered with ThroughLine Care, a global network of vetted helplines. - At **discord.findahelpline.com**, users can find free, confidential support in more than 100 countries. - The directory supports chat, text, and phone services in 11 languages and provides topic filters, operating hours, and contact details. - In the United States, users can text **DISCORD** to **741741** to reach Crisis Text Line, a free, confidential, 24/7 service available in English and Spanish. - Discord defines a crisis broadly to include stress, anxiety, school difficulties, relationship problems, and other situations—not only life-threatening emergencies. ## Gaming and Mental Health - Discord cites industry research suggesting that gaming can improve mood, foster community, provide healthy outlets, and help people through difficult periods. - The company sponsored Safe In Our World’s inaugural Mental Health Game Dev Champions event. - The event supports developers and gamers creating personal games based on lived experiences with mental health. - Discord planned to host a master class on building healthy communities in Safe In Our World’s Discord server on October 17, 2024. - The post also highlights Hero Journey Club, which uses games such as *Stardew Valley*, *Animal Crossing*, and *Final Fantasy* to facilitate emotional support and track feelings. - Discord partnered with Xbox Game Studios’ *Gears of War* on the “Never Fight Alone” suicide-prevention initiative, raising funds for Crisis Text Line and Aqui Estoy Chat. Discord’s overall recommendation is to treat gaming communities as potential sources of connection while continuing to provide accessible, confidential professional support. Its research and partnerships are intended to make mental-health assistance a more integrated part of the Discord experience.