<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Text Detection | Jason S. Lucas</title><link>https://jsl5710.github.io/tag/text-detection/</link><atom:link href="https://jsl5710.github.io/tag/text-detection/index.xml" rel="self" type="application/rss+xml"/><description>Text Detection</description><generator>HugoBlox Kit (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Tue, 24 Jun 2025 14:00:00 +0000</lastBuildDate><image><url>https://jsl5710.github.io/media/icon_hu_1b2044c02ce09a43.png</url><title>Text Detection</title><link>https://jsl5710.github.io/tag/text-detection/</link></image><item><title>Beemo - Benchmark of Expert-edited Machine-generated Outputs</title><link>https://jsl5710.github.io/event/talk-munich_invited/</link><pubDate>Tue, 24 Jun 2025 14:00:00 +0000</pubDate><guid>https://jsl5710.github.io/event/talk-munich_invited/</guid><description>&lt;div class="callout flex px-4 py-3 mb-6 rounded-md border-l-4 bg-blue-100 dark:bg-blue-900 border-blue-500"
data-callout="note"
data-callout-metadata=""&gt;
&lt;span class="callout-icon pr-3 pt-1 text-blue-600 dark:text-blue-300"&gt;
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m16.862 4.487l1.687-1.688a1.875 1.875 0 1 1 2.652 2.652L6.832 19.82a4.5 4.5 0 0 1-1.897 1.13l-2.685.8l.8-2.685a4.5 4.5 0 0 1 1.13-1.897zm0 0L19.5 7.125"/&gt;&lt;/svg&gt;
&lt;/span&gt;
&lt;div class="callout-content dark:text-neutral-300"&gt;
&lt;div class="callout-title font-semibold mb-1"&gt;Note&lt;/div&gt;
&lt;div class="callout-body"&gt;Click on the &lt;strong&gt;Video&lt;/strong&gt; link above to watch the full presentation on YouTube.&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;h2 id="talk-overview"&gt;Talk Overview&lt;/h2&gt;
&lt;p&gt;This presentation introduces &lt;strong&gt;Beemo&lt;/strong&gt;, a groundbreaking benchmark that addresses a critical gap in machine-generated text (MGT) detection research. Unlike traditional benchmarks that only consider single-author scenarios, Beemo captures the reality of human-AI collaboration in text creation.&lt;/p&gt;
&lt;h2 id="key-contributions"&gt;Key Contributions&lt;/h2&gt;
&lt;h3 id="1-multi-author-benchmark-design"&gt;1. Multi-Author Benchmark Design&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;19.6k total texts&lt;/strong&gt; across five use cases: open-ended generation, rewriting, summarization, open QA, and closed QA&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Expert-edited content&lt;/strong&gt;: 2,187 machine-generated texts refined by professional editors&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;LLM-edited variants&lt;/strong&gt;: 13.1k texts edited by GPT-4o and Llama3.1-70B-Instruct using diverse prompts&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="2-comprehensive-evaluation"&gt;2. Comprehensive Evaluation&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;33 MGT detector configurations&lt;/strong&gt; tested across multiple scenarios&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Zero-shot and pretrained detectors&lt;/strong&gt; including Binoculars, DetectGPT, RADAR, and MAGE&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Novel task formulations&lt;/strong&gt; examining detection performance on edited content&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="3-critical-findings"&gt;3. Critical Findings&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Expert editing evades detection&lt;/strong&gt;: AUROC scores drop by up to 22% for edited content&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;LLM-edited texts remain detectable&lt;/strong&gt;: Less likely to be classified as human-written&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Category-specific challenges&lt;/strong&gt;: Detection performance varies significantly across text types&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="implications-for-ai-safety"&gt;Implications for AI Safety&lt;/h2&gt;
&lt;p&gt;This research reveals significant vulnerabilities in current MGT detection systems, particularly relevant for:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Content moderation&lt;/strong&gt; and misinformation detection&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Academic integrity&lt;/strong&gt; in educational settings&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Publishing and journalism&lt;/strong&gt; authenticity verification&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Legal and regulatory&lt;/strong&gt; compliance for AI-generated content&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="future-directions"&gt;Future Directions&lt;/h2&gt;
&lt;p&gt;Beemo opens new research avenues in developing more robust detection methods that account for the collaborative nature of human-AI text creation in real-world applications.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;em&gt;This work was conducted at The Pennsylvania State University in collaboration with Toloka AI, MIT Lincoln Laboratory, and University of Oslo.&lt;/em&gt;&lt;/p&gt;</description></item></channel></rss>