<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Benchmark | Jason S. Lucas</title><link>https://jsl5710.github.io/tag/benchmark/</link><atom:link href="https://jsl5710.github.io/tag/benchmark/index.xml" rel="self" type="application/rss+xml"/><description>Benchmark</description><generator>HugoBlox Kit (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Sun, 01 Feb 2026 00:00:00 +0000</lastBuildDate><image><url>https://jsl5710.github.io/media/icon_hu_1b2044c02ce09a43.png</url><title>Benchmark</title><link>https://jsl5710.github.io/tag/benchmark/</link></image><item><title>BLUFF: Benchmarking in Low-resoUrce Languages for detecting Falsehoods and Fake news</title><link>https://jsl5710.github.io/publication/conference-paper-bluff/</link><pubDate>Sun, 01 Feb 2026 00:00:00 +0000</pubDate><guid>https://jsl5710.github.io/publication/conference-paper-bluff/</guid><description>&lt;p&gt;&lt;strong&gt;BLUFF&lt;/strong&gt; is the largest multilingual fake news detection benchmark to date, spanning &lt;strong&gt;79 languages&lt;/strong&gt; (20 high-resource &amp;ldquo;big-head&amp;rdquo; + 59 low-resource &amp;ldquo;long-tail&amp;rdquo;) with over &lt;strong&gt;202,000 samples&lt;/strong&gt;. The benchmark combines human-written fact-checked content from 130 IFCN-certified organizations with LLM-generated content from 19 diverse models.&lt;/p&gt;
&lt;p&gt;Key contributions include:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;AXL-CoI&lt;/strong&gt; (Adversarial Cross-Lingual Agentic Chain-of-Interactions): A multi-agentic framework using 10 fake chains and 8 real chains for controlled multilingual content generation&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;mPURIFY&lt;/strong&gt;: A 4-stage quality filtering pipeline with 32 features across 5 dimensions, ensuring dataset integrity through asymmetric evaluation thresholds&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Bidirectional translation&lt;/strong&gt;: English↔X coverage across 70+ languages with 4 prompt variants&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Comprehensive evaluation&lt;/strong&gt;: State-of-the-art detectors suffer up to 25.3% Macro-F1 degradation on low-resource versus high-resource languages&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Resources:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;/li&gt;
&lt;/ul&gt;</description></item></channel></rss>