<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Multilingual NLP | Jason S. Lucas</title><link>https://jsl5710.github.io/tag/multilingual-nlp/</link><atom:link href="https://jsl5710.github.io/tag/multilingual-nlp/index.xml" rel="self" type="application/rss+xml"/><description>Multilingual NLP</description><generator>HugoBlox Kit (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Mon, 13 Jul 2026 00:00:00 +0000</lastBuildDate><image><url>https://jsl5710.github.io/media/icon_hu_1b2044c02ce09a43.png</url><title>Multilingual NLP</title><link>https://jsl5710.github.io/tag/multilingual-nlp/</link></image><item><title>Position: Breaking the Dual Curse of Multilingual AI Requires Socio-Technical Guardrails, Not Post-Hoc Alignment</title><link>https://jsl5710.github.io/publication/conference-paper-dual-curse-multilingual/</link><pubDate>Mon, 13 Jul 2026 00:00:00 +0000</pubDate><guid>https://jsl5710.github.io/publication/conference-paper-dual-curse-multilingual/</guid><description>&lt;p&gt;&lt;strong&gt;Position:&lt;/strong&gt; Multilingual AI safety cannot be retrofitted through post-hoc alignment. We identify a &lt;strong&gt;dual curse&lt;/strong&gt; in current systems and argue for socio-technical guardrails built in from pre-training.&lt;/p&gt;
&lt;p&gt;Key contributions:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Dual Curse documented&lt;/strong&gt;: Harmful content generation rises to &lt;strong&gt;35% in low-resource languages&lt;/strong&gt; (vs. 1% in English), while instruction-following capability declines sharply across the same languages.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Systematic review of 207 studies&lt;/strong&gt;: Reward models achieve only &lt;strong&gt;49–50% accuracy in low-resource languages&lt;/strong&gt; — equivalent to random chance — undermining post-deployment safety pipelines.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Socio-technical prescription&lt;/strong&gt;: Pre-training interventions, &lt;strong&gt;community-led harm specification&lt;/strong&gt;, and multilingual evaluation metrics that balance security and usability jointly, rather than trading one off for the other.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Call to action&lt;/strong&gt;: Treat multilingual safety as a first-class design constraint, not a downstream patch applied through RLHF or filtering.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Resources:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;/ul&gt;</description></item><item><title>Dagstuhl Seminar 26252 — From Speech Translation to Multilingual Communication</title><link>https://jsl5710.github.io/event/seminar-dagstuhl-26252/</link><pubDate>Sun, 14 Jun 2026 09:00:00 +0000</pubDate><guid>https://jsl5710.github.io/event/seminar-dagstuhl-26252/</guid><description>&lt;p&gt;I am honored to be invited to &lt;strong&gt;Dagstuhl Seminar 26252: &amp;ldquo;From Speech Translation to Multilingual Communication – New Research Challenges&amp;rdquo;&lt;/strong&gt; taking place June 14–17, 2026 at Schloss Dagstuhl in Germany.&lt;/p&gt;
&lt;p&gt;This invitation-only seminar will bring together an interdisciplinary group of researchers from speech translation, interpretation studies, and human-computer interaction to explore how AI can better support multilingual communication in real-world scenarios.&lt;/p&gt;
&lt;p&gt;The seminar&amp;rsquo;s focus on bridging the gap between technical advances and end-user needs resonates deeply with my research on addressing the digital language divide.&lt;/p&gt;
&lt;p&gt;Organized by &lt;strong&gt;Marine Carpuat&lt;/strong&gt;, &lt;strong&gt;Claudio Fantinuoli&lt;/strong&gt;, &lt;strong&gt;Ge Gao&lt;/strong&gt;, and &lt;strong&gt;Jan Niehues&lt;/strong&gt;.&lt;/p&gt;</description></item><item><title>BLUFF: Benchmarking in Low-resoUrce Languages for detecting Falsehoods and Fake news</title><link>https://jsl5710.github.io/publication/conference-paper-bluff/</link><pubDate>Sun, 01 Feb 2026 00:00:00 +0000</pubDate><guid>https://jsl5710.github.io/publication/conference-paper-bluff/</guid><description>&lt;p&gt;&lt;strong&gt;BLUFF&lt;/strong&gt; is the largest multilingual fake news detection benchmark to date, spanning &lt;strong&gt;79 languages&lt;/strong&gt; (20 high-resource &amp;ldquo;big-head&amp;rdquo; + 59 low-resource &amp;ldquo;long-tail&amp;rdquo;) with over &lt;strong&gt;202,000 samples&lt;/strong&gt;. The benchmark combines human-written fact-checked content from 130 IFCN-certified organizations with LLM-generated content from 19 diverse models.&lt;/p&gt;
&lt;p&gt;Key contributions include:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;AXL-CoI&lt;/strong&gt; (Adversarial Cross-Lingual Agentic Chain-of-Interactions): A multi-agentic framework using 10 fake chains and 8 real chains for controlled multilingual content generation&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;mPURIFY&lt;/strong&gt;: A 4-stage quality filtering pipeline with 32 features across 5 dimensions, ensuring dataset integrity through asymmetric evaluation thresholds&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Bidirectional translation&lt;/strong&gt;: English↔X coverage across 70+ languages with 4 prompt variants&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Comprehensive evaluation&lt;/strong&gt;: State-of-the-art detectors suffer up to 25.3% Macro-F1 degradation on low-resource versus high-resource languages&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Resources:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;/ul&gt;</description></item><item><title>Multilingual NLP</title><link>https://jsl5710.github.io/project/multilingual-nlp/</link><pubDate>Sun, 01 Feb 2026 00:00:00 +0000</pubDate><guid>https://jsl5710.github.io/project/multilingual-nlp/</guid><description>&lt;p&gt;Language technologies often fail beyond a handful of well-resourced languages, leaving billions of speakers vulnerable to disinformation and harmful content. This project develops multilingual approaches for detecting fake news, false claims, and machine-generated text across diverse linguistic landscapes—spanning 70+ languages. By building large-scale benchmarks and leveraging transfer learning techniques, this work directly addresses the Digital Language Divide, ensuring that information integrity tools are not limited to English or other high-resource languages but extend protection to the communities most susceptible to unchecked disinformation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Related Publications:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;BLUFF&lt;/strong&gt; (2026) — Benchmarking falsehoods/fake news in low-resource languages&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Beyond Speculation&lt;/strong&gt; (2026, IEEE) — LLM-generated texts in multilingual disinformation&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;MULTITuDE&lt;/strong&gt; (2023, EMNLP) — Multilingual machine-generated text detection benchmark&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Fighting Fire with Fire (F3)&lt;/strong&gt; (2023, EMNLP) — LLMs&amp;rsquo; dual role in crafting/detecting disinformation&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Detecting False Claims in Low-Resource Regions&lt;/strong&gt; (2022, ACL) — Caribbean false claim detection&lt;/li&gt;
&lt;/ul&gt;</description></item><item><title>Penn State connections lead doctoral student to interdisciplinary College of IST</title><link>https://jsl5710.github.io/event/media-iconnect_web/</link><pubDate>Mon, 19 May 2025 00:00:00 +0000</pubDate><guid>https://jsl5710.github.io/event/media-iconnect_web/</guid><description>&lt;div class="callout flex px-4 py-3 mb-6 rounded-md border-l-4 bg-blue-100 dark:bg-blue-900 border-blue-500"
data-callout="note"
data-callout-metadata=""&gt;
&lt;span class="callout-icon pr-3 pt-1 text-blue-600 dark:text-blue-300"&gt;
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m16.862 4.487l1.687-1.688a1.875 1.875 0 1 1 2.652 2.652L6.832 19.82a4.5 4.5 0 0 1-1.897 1.13l-2.685.8l.8-2.685a4.5 4.5 0 0 1 1.13-1.897zm0 0L19.5 7.125"/&gt;&lt;/svg&gt;
&lt;/span&gt;
&lt;div class="callout-content dark:text-neutral-300"&gt;
&lt;div class="callout-title font-semibold mb-1"&gt;Note&lt;/div&gt;
&lt;div class="callout-body"&gt;Click on the &lt;strong&gt;Read Full Article&lt;/strong&gt; link above to read the complete story on Penn State News.&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;h2 id="academic-journey"&gt;Academic Journey&lt;/h2&gt;
&lt;p&gt;This feature story traces the interdisciplinary path that led Jason Lucas from the West Indies to Penn State&amp;rsquo;s College of Information Sciences and Technology, where he&amp;rsquo;s pursuing groundbreaking research in multilingual AI and harmful content detection.&lt;/p&gt;
&lt;h2 id="key-highlights"&gt;Key Highlights&lt;/h2&gt;
&lt;h3 id="multidisciplinary-foundation"&gt;Multidisciplinary Foundation&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Information Technology&lt;/strong&gt;: Bachelor&amp;rsquo;s degree at St. George&amp;rsquo;s University in Grenada, West Indies&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Health Informatics&lt;/strong&gt;: Master&amp;rsquo;s in computer information systems at Boston University&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Public Health&lt;/strong&gt;: Master of public health back at St. George&amp;rsquo;s University&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Current Focus&lt;/strong&gt;: Doctoral degree in informatics at Penn State&amp;rsquo;s College of IST&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="mentorship-network"&gt;Mentorship Network&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Theodore Hollis&lt;/strong&gt;: Professor emeritus from Penn State&amp;rsquo;s Eberly College of Science, mentor at SGU&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Dongwon Lee&lt;/strong&gt;: Professor and director of doctoral programs, current graduate adviser&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;LinDiV Program&lt;/strong&gt;: Transdisciplinary fellowship enhancing linguistic knowledge&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;PIKE Research Group&lt;/strong&gt;: Collaborative environment for data mining and security applications&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="research-evolution"&gt;Research Evolution&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Perfect alignment&lt;/strong&gt;: Informatics field matching his interdisciplinary background&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;AI and NLP solutions&lt;/strong&gt;: Addressing societal challenges in healthcare and information integrity&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Multilingual focus&lt;/strong&gt;: Developing technologies for languages with limited digital resources&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Social impact&lt;/strong&gt;: Protecting vulnerable communities from manipulation and harmful content&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="current-research-impact"&gt;Current Research Impact&lt;/h2&gt;
&lt;h3 id="pike-research-group"&gt;PIKE Research Group&lt;/h3&gt;
&lt;p&gt;Lucas works within the Penn State Information Knowledge and Web (PIKE) Research Group, studying data management and mining across diverse forms with focus on social and security applications.&lt;/p&gt;
&lt;h3 id="research-goals"&gt;Research Goals&lt;/h3&gt;
&lt;blockquote class="border-l-4 border-neutral-300 dark:border-neutral-600 pl-4 italic text-neutral-600 dark:text-neutral-400 my-6"&gt;
&lt;p&gt;&amp;ldquo;Success in my work would look like advancing research for social good and developing deployable technologies that protect vulnerable populations from manipulation and harmful content. My ultimate goal is to further the development of inclusive language technology that not only bridges the digital language divide but also protects people from hidden manipulations that exploit psychosocial biases.&amp;rdquo;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h3 id="lindiv-fellowship-impact"&gt;LinDiV Fellowship Impact&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Filling gaps&lt;/strong&gt;: Providing depth in linguistics and language sciences&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Transdisciplinary approach&lt;/strong&gt;: Combining computational methods with linguistic theory&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Community building&lt;/strong&gt;: Connecting with researchers across disciplines&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Research enhancement&lt;/strong&gt;: Strengthening multilingual NLP capabilities&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="looking-forward"&gt;Looking Forward&lt;/h2&gt;
&lt;h3 id="lawrence-livermore-national-laboratory-opportunity"&gt;Lawrence Livermore National Laboratory Opportunity&lt;/h3&gt;
&lt;p&gt;Lucas is preparing for the Data Science Summer Institute at LLNL, where he will:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Apply research&lt;/strong&gt;: Work on projects of national importance&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Collaborate&lt;/strong&gt;: Partner with leading scientists and engineers&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Bridge academia&lt;/strong&gt;: Connect academic research with practical applications&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Gain experience&lt;/strong&gt;: Access high-performance computing resources&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="vision-for-impact"&gt;Vision for Impact&lt;/h3&gt;
&lt;p&gt;Lucas aims to develop inclusive language technologies that can detect and counter harmful content across multiple languages, with particular relevance for defending communities whose languages aren&amp;rsquo;t prioritized by major tech platforms.&lt;/p&gt;
&lt;h2 id="the-power-of-interdisciplinary-education"&gt;The Power of Interdisciplinary Education&lt;/h2&gt;
&lt;blockquote class="border-l-4 border-neutral-300 dark:border-neutral-600 pl-4 italic text-neutral-600 dark:text-neutral-400 my-6"&gt;
&lt;p&gt;&amp;ldquo;Throughout this journey, I was unconsciously building a multidisciplinary background that would later become my greatest strength. Learning directly from leading clinicians, public health professionals and technology experts across these institutions gave me a unique perspective on solving complex problems at the intersection of these fields.&amp;rdquo;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Lucas&amp;rsquo;s story demonstrates how diverse educational experiences, combined with strong mentorship, can create researchers uniquely positioned to tackle complex, real-world challenges.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;em&gt;Read the complete story to learn more about Jason Lucas&amp;rsquo;s academic journey and his vision for developing inclusive language technologies that protect vulnerable communities worldwide.&lt;/em&gt;&lt;/p&gt;</description></item></channel></rss>