<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Content Moderation | Jason S. Lucas</title><link>https://jsl5710.github.io/tag/content-moderation/</link><atom:link href="https://jsl5710.github.io/tag/content-moderation/index.xml" rel="self" type="application/rss+xml"/><description>Content Moderation</description><generator>HugoBlox Kit (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Wed, 08 Apr 2026 00:00:00 +0000</lastBuildDate><image><url>https://jsl5710.github.io/media/icon_hu_1b2044c02ce09a43.png</url><title>Content Moderation</title><link>https://jsl5710.github.io/tag/content-moderation/</link></image><item><title>DIA-HARM: Dialectal Disparities in Harmful Content Detection Across 50 English Dialects</title><link>https://jsl5710.github.io/publication/conference-paper-dia-harm/</link><pubDate>Wed, 08 Apr 2026 00:00:00 +0000</pubDate><guid>https://jsl5710.github.io/publication/conference-paper-dia-harm/</guid><description>
&lt;blockquote class="border-l-4 border-neutral-300 dark:border-neutral-600 pl-4 italic text-neutral-600 dark:text-neutral-400 my-6"&gt;
&lt;p&gt;🏆 &lt;strong&gt;Social Impact Paper Award — ACL 2026&lt;/strong&gt;, San Diego, July 2–7, 2026.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;
&lt;figure &gt;
&lt;div class="flex justify-center "&gt;
&lt;div class="w-full" &gt;
&lt;img alt="Award ceremony at ACL 2026"
srcset="https://jsl5710.github.io/publication/conference-paper-dia-harm/award_hu_80ed752c554010bb.webp 320w, https://jsl5710.github.io/publication/conference-paper-dia-harm/award_hu_60d9b5863648e4e5.webp 480w, https://jsl5710.github.io/publication/conference-paper-dia-harm/award_hu_ed1787f737c5acbb.webp 760w"
sizes="(max-width: 480px) 100vw, (max-width: 768px) 90vw, (max-width: 1024px) 80vw, 760px"
src="https://jsl5710.github.io/publication/conference-paper-dia-harm/award_hu_80ed752c554010bb.webp"
width="760"
height="570"
loading="lazy" data-zoomable /&gt;&lt;/div&gt;
&lt;/div&gt;&lt;/figure&gt;
&lt;figure &gt;
&lt;div class="flex justify-center "&gt;
&lt;div class="w-full" &gt;
&lt;img alt="Award recipients slide showing DIA-HARM"
srcset="https://jsl5710.github.io/publication/conference-paper-dia-harm/award-slide_hu_6c685cfd29050812.webp 320w, https://jsl5710.github.io/publication/conference-paper-dia-harm/award-slide_hu_a0e9c278522b6ee6.webp 480w, https://jsl5710.github.io/publication/conference-paper-dia-harm/award-slide_hu_906922782ba7a4e0.webp 760w"
sizes="(max-width: 480px) 100vw, (max-width: 768px) 90vw, (max-width: 1024px) 80vw, 760px"
src="https://jsl5710.github.io/publication/conference-paper-dia-harm/award-slide_hu_6c685cfd29050812.webp"
width="760"
height="570"
loading="lazy" data-zoomable /&gt;&lt;/div&gt;
&lt;/div&gt;&lt;/figure&gt;
&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;DIA-HARM&lt;/strong&gt; investigates the robustness of harmful content detection systems across &lt;strong&gt;50 English dialects&lt;/strong&gt;, addressing critical equity gaps in automated content moderation.&lt;/p&gt;
&lt;p&gt;Key contributions include:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Dialect-Diverse Detection (D3) Corpus&lt;/strong&gt;: Over 195K samples derived from benchmark harmful content datasets, transformed using 189 morphosyntactic rules from eWAVE covering 50 dialects&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Comprehensive Model Evaluation&lt;/strong&gt;: 16 detection models tested — 10 fine-tuned, 5 zero-shot, and 1 in-context learning — revealing systematic performance disparities across dialect groups&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;D-PURIFY Validation&lt;/strong&gt;: Quality filtering pipeline ensuring linguistic validity of dialect transformations with 97.2% average F1 for the best model (mDeBERTa)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Key Finding&lt;/strong&gt;: Detection degradation correlates with density of morphosyntactic transformations rather than specific dialect features, with 2,450 dialect pairs analyzed&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Resources:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;/ul&gt;</description></item><item><title>AI Robustness &amp; Adversarial Safety</title><link>https://jsl5710.github.io/project/ai-robustness/</link><pubDate>Sun, 01 Feb 2026 00:00:00 +0000</pubDate><guid>https://jsl5710.github.io/project/ai-robustness/</guid><description>&lt;p&gt;AI systems deployed in the real world must withstand adversarial manipulation and perform reliably across the full spectrum of human language variation. This project investigates how dialect diversity, authorship obfuscation, and expert-level text editing expose critical vulnerabilities in content detection systems. From stress-testing harmful content classifiers across 50 English dialects to evaluating robustness against sophisticated evasion techniques, this work reveals that the Digital Language Divide is not only a gap in language coverage but also a security vulnerability—one that adversaries can exploit when AI systems are brittle to linguistic variation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Related Publications:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;DIA-HARM&lt;/strong&gt; (2026) — Harmful content detection robustness across 50 dialects&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Authorship Obfuscation in Multilingual MGT Detection&lt;/strong&gt; (2024, EMNLP)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;BEEMO&lt;/strong&gt; (2025, NAACL) — Expert-edited machine-generated outputs benchmark&lt;/li&gt;
&lt;/ul&gt;</description></item></channel></rss>