<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Low-Resource Languages | Jason S. Lucas</title><link>https://jsl5710.github.io/tag/low-resource-languages/</link><atom:link href="https://jsl5710.github.io/tag/low-resource-languages/index.xml" rel="self" type="application/rss+xml"/><description>Low-Resource Languages</description><generator>HugoBlox Kit (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Mon, 13 Jul 2026 00:00:00 +0000</lastBuildDate><image><url>https://jsl5710.github.io/media/icon_hu_1b2044c02ce09a43.png</url><title>Low-Resource Languages</title><link>https://jsl5710.github.io/tag/low-resource-languages/</link></image><item><title>Position: Breaking the Dual Curse of Multilingual AI Requires Socio-Technical Guardrails, Not Post-Hoc Alignment</title><link>https://jsl5710.github.io/publication/conference-paper-dual-curse-multilingual/</link><pubDate>Mon, 13 Jul 2026 00:00:00 +0000</pubDate><guid>https://jsl5710.github.io/publication/conference-paper-dual-curse-multilingual/</guid><description>&lt;p&gt;&lt;strong&gt;Position:&lt;/strong&gt; Multilingual AI safety cannot be retrofitted through post-hoc alignment. We identify a &lt;strong&gt;dual curse&lt;/strong&gt; in current systems and argue for socio-technical guardrails built in from pre-training.&lt;/p&gt;
&lt;p&gt;Key contributions:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Dual Curse documented&lt;/strong&gt;: Harmful content generation rises to &lt;strong&gt;35% in low-resource languages&lt;/strong&gt; (vs. 1% in English), while instruction-following capability declines sharply across the same languages.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Systematic review of 207 studies&lt;/strong&gt;: Reward models achieve only &lt;strong&gt;49–50% accuracy in low-resource languages&lt;/strong&gt; — equivalent to random chance — undermining post-deployment safety pipelines.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Socio-technical prescription&lt;/strong&gt;: Pre-training interventions, &lt;strong&gt;community-led harm specification&lt;/strong&gt;, and multilingual evaluation metrics that balance security and usability jointly, rather than trading one off for the other.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Call to action&lt;/strong&gt;: Treat multilingual safety as a first-class design constraint, not a downstream patch applied through RLHF or filtering.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Resources:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;/ul&gt;</description></item><item><title>BLUFF: Benchmarking in Low-resoUrce Languages for detecting Falsehoods and Fake news</title><link>https://jsl5710.github.io/publication/conference-paper-bluff/</link><pubDate>Sun, 01 Feb 2026 00:00:00 +0000</pubDate><guid>https://jsl5710.github.io/publication/conference-paper-bluff/</guid><description>&lt;p&gt;&lt;strong&gt;BLUFF&lt;/strong&gt; is the largest multilingual fake news detection benchmark to date, spanning &lt;strong&gt;79 languages&lt;/strong&gt; (20 high-resource &amp;ldquo;big-head&amp;rdquo; + 59 low-resource &amp;ldquo;long-tail&amp;rdquo;) with over &lt;strong&gt;202,000 samples&lt;/strong&gt;. The benchmark combines human-written fact-checked content from 130 IFCN-certified organizations with LLM-generated content from 19 diverse models.&lt;/p&gt;
&lt;p&gt;Key contributions include:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;AXL-CoI&lt;/strong&gt; (Adversarial Cross-Lingual Agentic Chain-of-Interactions): A multi-agentic framework using 10 fake chains and 8 real chains for controlled multilingual content generation&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;mPURIFY&lt;/strong&gt;: A 4-stage quality filtering pipeline with 32 features across 5 dimensions, ensuring dataset integrity through asymmetric evaluation thresholds&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Bidirectional translation&lt;/strong&gt;: English↔X coverage across 70+ languages with 4 prompt variants&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Comprehensive evaluation&lt;/strong&gt;: State-of-the-art detectors suffer up to 25.3% Macro-F1 degradation on low-resource versus high-resource languages&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Resources:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;/ul&gt;</description></item><item><title>Multilingual NLP</title><link>https://jsl5710.github.io/project/multilingual-nlp/</link><pubDate>Sun, 01 Feb 2026 00:00:00 +0000</pubDate><guid>https://jsl5710.github.io/project/multilingual-nlp/</guid><description>&lt;p&gt;Language technologies often fail beyond a handful of well-resourced languages, leaving billions of speakers vulnerable to disinformation and harmful content. This project develops multilingual approaches for detecting fake news, false claims, and machine-generated text across diverse linguistic landscapes—spanning 70+ languages. By building large-scale benchmarks and leveraging transfer learning techniques, this work directly addresses the Digital Language Divide, ensuring that information integrity tools are not limited to English or other high-resource languages but extend protection to the communities most susceptible to unchecked disinformation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Related Publications:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;BLUFF&lt;/strong&gt; (2026) — Benchmarking falsehoods/fake news in low-resource languages&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Beyond Speculation&lt;/strong&gt; (2026, IEEE) — LLM-generated texts in multilingual disinformation&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;MULTITuDE&lt;/strong&gt; (2023, EMNLP) — Multilingual machine-generated text detection benchmark&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Fighting Fire with Fire (F3)&lt;/strong&gt; (2023, EMNLP) — LLMs&amp;rsquo; dual role in crafting/detecting disinformation&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Detecting False Claims in Low-Resource Regions&lt;/strong&gt; (2022, ACL) — Caribbean false claim detection&lt;/li&gt;
&lt;/ul&gt;</description></item><item><title>Equity, Inclusion &amp; the Digital Language Divide</title><link>https://jsl5710.github.io/project/equity-inclusion/</link><pubDate>Fri, 01 Nov 2024 00:00:00 +0000</pubDate><guid>https://jsl5710.github.io/project/equity-inclusion/</guid><description>&lt;p&gt;The benefits and harms of generative AI are not distributed equally. This project examines how AI systems disproportionately impact long-tail users—speakers of underserved languages and members of marginalized communities who are often excluded from model training and evaluation. By quantifying these disparities and analyzing how generative AI amplifies existing inequities in the information ecosystem, this work makes the case that equitable AI is not optional but essential. It bridges AI for Social Good with Safe and Ethical AI, centering the voices and needs of communities that current technologies routinely overlook.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Related Publications:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Generative AI Disproportionately Harms Long Tail Users&lt;/strong&gt; (2024, Computer)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The Longtail Impact of Generative AI on Disinformation&lt;/strong&gt; (2024, IEEE IS)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Detecting False Claims in Low-Resource Regions&lt;/strong&gt; (2022, ACL) — cross-listed&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;BLUFF&lt;/strong&gt; (2026) — cross-listed&lt;/li&gt;
&lt;/ul&gt;</description></item></channel></rss>