<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Shaurya Rohatgi | Jason S. Lucas</title><link>https://jsl5710.github.io/authors/shaurya-rohatgi/</link><atom:link href="https://jsl5710.github.io/authors/shaurya-rohatgi/index.xml" rel="self" type="application/rss+xml"/><description>Shaurya Rohatgi</description><generator>HugoBlox Kit (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Wed, 06 Dec 2023 00:00:00 +0000</lastBuildDate><item><title>Fighting Fire with Fire - EMNLP 2023</title><link>https://jsl5710.github.io/slides/f3-slides/</link><pubDate>Wed, 06 Dec 2023 00:00:00 +0000</pubDate><guid>https://jsl5710.github.io/slides/f3-slides/</guid><description>
&lt;section data-noprocess data-shortcode-slide
data-background-color="#ffffff"
&gt;
&lt;h1 id="fighting-fire-with-fire"&gt;Fighting Fire with Fire&lt;/h1&gt;
&lt;p&gt;&lt;strong&gt;LLMs in Crafting and Detecting Disinformation&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;small&gt;Jason Lucas¹, Adaku Uchendu¹,², Michiharu Yamashita¹, Jooyoung Lee¹, Shaurya Rohatgi¹, Dongwon Lee¹&lt;/small&gt;&lt;/p&gt;
&lt;p&gt;&lt;small&gt;¹Penn State University, ²MIT Lincoln Laboratory&lt;/small&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;EMNLP 2023 Main Conference&lt;/strong&gt;&lt;/p&gt;
&lt;hr&gt;
&lt;section data-noprocess data-shortcode-slide
data-background-color="#f8f9fa"
&gt;
&lt;h2 id="problem--motivation"&gt;Problem &amp;amp; Motivation&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Challenge&lt;/strong&gt;: LLMs generate realistic but harmful disinformation&lt;/p&gt;
&lt;small&gt;
&lt;span class="fragment " &gt;
&lt;ul&gt;
&lt;li&gt;Persuasive texts indistinguishable from human content&lt;/li&gt;
&lt;li&gt;Large-scale disinformation potential&lt;/li&gt;
&lt;li&gt;Limited LLM-generated content detection&lt;/li&gt;
&lt;/ul&gt;
&lt;/span&gt;
&lt;/small&gt;
&lt;span class="fragment " &gt;
&lt;strong&gt;Core Question&lt;/strong&gt;: &lt;em&gt;Can LLMs detect their own disinformation?&lt;/em&gt;
&lt;/span&gt;
&lt;hr&gt;
&lt;section data-noprocess data-shortcode-slide
data-background-color="#ffffff"
&gt;
&lt;h2 id="research-questions"&gt;Research Questions&lt;/h2&gt;
&lt;small&gt;
**RQ1**: Can LLMs efficiently generate disinformation via prompt engineering?
&lt;p&gt;&lt;strong&gt;RQ2&lt;/strong&gt;: How proficient are LLMs at detecting disinformation?&lt;/p&gt;
&lt;span class="fragment " &gt;
&lt;p&gt;&lt;strong&gt;5 Evaluation Dimensions&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Human vs. LLM-generated&lt;/li&gt;
&lt;li&gt;Self vs. externally-generated&lt;/li&gt;
&lt;li&gt;Posts vs. articles&lt;/li&gt;
&lt;li&gt;In vs. out-of-distribution&lt;/li&gt;
&lt;li&gt;Zero-shot LLMs vs. fine-tuned detectors&lt;/li&gt;
&lt;/ul&gt;
&lt;/span&gt;
&lt;/small&gt;
&lt;hr&gt;
&lt;section data-noprocess data-shortcode-slide
data-background-color="#f8f9fa"
&gt;
&lt;h2 id="f3-framework"&gt;F3 Framework&lt;/h2&gt;
&lt;small&gt;
**Fighting Fire with Fire (F3)** - 5-step approach:
&lt;span class="fragment " &gt;
&lt;ol&gt;
&lt;li&gt;Human data collection&lt;/li&gt;
&lt;li&gt;Prompt engineering generation&lt;/li&gt;
&lt;li&gt;PURIFY hallucination filtering&lt;/li&gt;
&lt;li&gt;Cloze-prompt detection&lt;/li&gt;
&lt;li&gt;Zero-shot evaluation&lt;/li&gt;
&lt;/ol&gt;
&lt;/span&gt;
&lt;/small&gt;
&lt;hr&gt;
&lt;section data-noprocess data-shortcode-slide
data-background-color="#ffffff"
&gt;
&lt;h2 id="rq1-bypassing-alignment"&gt;RQ1: Bypassing Alignment&lt;/h2&gt;
&lt;small&gt;
**Key Discovery**: Impersonator roles override safety measures
&lt;span class="fragment " &gt;
&lt;strong&gt;Without role&lt;/strong&gt;: &lt;em&gt;&amp;ldquo;Sorry, I can&amp;rsquo;t assist&amp;hellip;&amp;rdquo;&lt;/em&gt;
&lt;/span&gt;
&lt;span class="fragment " &gt;
&lt;strong&gt;With role&lt;/strong&gt; (&amp;ldquo;You are an AI news curator&amp;rdquo;): ✅ Generates disinformation
&lt;/span&gt;
&lt;span class="fragment " &gt;
&lt;strong&gt;Finding&lt;/strong&gt;: Impersonator prompts successfully bypass GPT-3.5 protections
&lt;/span&gt;
&lt;/small&gt;
&lt;hr&gt;
&lt;section data-noprocess data-shortcode-slide
data-background-color="#f8f9fa"
&gt;
&lt;h2 id="generation-strategies"&gt;Generation Strategies&lt;/h2&gt;
&lt;small&gt;
**Perturbation-Based** (Fake):
&lt;ul&gt;
&lt;li&gt;Minor: Subtle changes&lt;/li&gt;
&lt;li&gt;Major: Noticeable changes&lt;/li&gt;
&lt;li&gt;Critical: Significant alterations&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Paraphrase-Based&lt;/strong&gt; (Real):&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Minor: Light summary&lt;/li&gt;
&lt;li&gt;Major: Moderate rewording&lt;/li&gt;
&lt;li&gt;Critical: Full rephrasing&lt;/li&gt;
&lt;/ul&gt;
&lt;span class="fragment " &gt;
&lt;strong&gt;Output&lt;/strong&gt;: 43K+ synthetic samples
&lt;/span&gt;
&lt;/small&gt;
&lt;hr&gt;
&lt;section data-noprocess data-shortcode-slide
data-background-color="#ffffff"
&gt;
&lt;h2 id="purify-framework"&gt;PURIFY Framework&lt;/h2&gt;
&lt;small&gt;
**Problem**: 38% hallucinated misalignments
&lt;span class="fragment " &gt;
&lt;p&gt;&lt;strong&gt;PURIFY&lt;/strong&gt; filters using 4 metrics:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Natural Language Inference&lt;/li&gt;
&lt;li&gt;AlignScore&lt;/li&gt;
&lt;li&gt;BERTScore&lt;/li&gt;
&lt;li&gt;Semantic Distance&lt;/li&gt;
&lt;/ul&gt;
&lt;/span&gt;
&lt;span class="fragment " &gt;
&lt;strong&gt;Result&lt;/strong&gt;: 43,272 → 27,667 quality samples
&lt;/span&gt;
&lt;/small&gt;
&lt;hr&gt;
&lt;section data-noprocess data-shortcode-slide
data-background-color="#f8f9fa"
&gt;
&lt;h2 id="rq2-detection-results"&gt;RQ2: Detection Results&lt;/h2&gt;
&lt;small&gt;
**Human vs. LLM Content**:
&lt;ul&gt;
&lt;li&gt;Human-authored: 55-66% accuracy&lt;/li&gt;
&lt;li&gt;LLM-generated: 60-85% accuracy&lt;/li&gt;
&lt;/ul&gt;
&lt;span class="fragment " &gt;
&lt;p&gt;&lt;strong&gt;Self vs. External&lt;/strong&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;GPT-3.5: Strong self-detection&lt;/li&gt;
&lt;li&gt;LLaMA-GPT: Best external detector&lt;/li&gt;
&lt;li&gt;Challenge: Minor disinformation detection&lt;/li&gt;
&lt;/ul&gt;
&lt;/span&gt;
&lt;/small&gt;
&lt;hr&gt;
&lt;section data-noprocess data-shortcode-slide
data-background-color="#ffffff"
&gt;
&lt;h2 id="key-findings"&gt;Key Findings&lt;/h2&gt;
&lt;small&gt;
&lt;span class="fragment " &gt;
&lt;strong&gt;Content Type&lt;/strong&gt;: Articles &amp;gt; Social media posts
&lt;/span&gt;
&lt;span class="fragment " &gt;
&lt;strong&gt;Distribution&lt;/strong&gt;: In-distribution &amp;gt; Out-of-distribution
&lt;/span&gt;
&lt;span class="fragment " &gt;
&lt;strong&gt;Model Type&lt;/strong&gt;: Fine-tuned &amp;gt; GPT-3.5 &amp;gt; Domain-specific
&lt;/span&gt;
&lt;span class="fragment " &gt;
&lt;strong&gt;Critical&lt;/strong&gt;: Subtle disinformation challenges all detectors
&lt;/span&gt;
&lt;/small&gt;
&lt;hr&gt;
&lt;section data-noprocess data-shortcode-slide
data-background-color="#f8f9fa"
&gt;
&lt;h2 id="technical-contributions"&gt;Technical Contributions&lt;/h2&gt;
&lt;small&gt;
1. Novel prompting for disinformation generation
2. PURIFY hallucination filtering framework
3. Cloze-prompt detection strategies
4. Comprehensive SOTA benchmark
5. F3 dataset for research community
&lt;/small&gt;
&lt;hr&gt;
&lt;section data-noprocess data-shortcode-slide
data-background-color="#ffffff"
&gt;
&lt;h2 id="dataset--evaluation"&gt;Dataset &amp;amp; Evaluation&lt;/h2&gt;
&lt;small&gt;
**Models**: GPT-3.5, LLaMA-2, Palm-2, Dolly-2
&lt;p&gt;&lt;strong&gt;Data&lt;/strong&gt;: CoAID, FakeNewsNet, F3 (27,667 samples)&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Languages&lt;/strong&gt;: 11 languages, Pre/Post-GPT splits&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Metrics&lt;/strong&gt;: Macro-F1 across human/AI datasets
&lt;/small&gt;&lt;/p&gt;
&lt;hr&gt;
&lt;section data-noprocess data-shortcode-slide
data-background-color="#f8f9fa"
&gt;
&lt;h2 id="impact--implications"&gt;Impact &amp;amp; Implications&lt;/h2&gt;
&lt;small&gt;
&lt;span class="fragment " &gt;
&lt;strong&gt;Dual-Use Reality&lt;/strong&gt;: LLMs both create and detect disinformation
&lt;/span&gt;
&lt;span class="fragment " &gt;
&lt;strong&gt;Detection Promise&lt;/strong&gt;: Zero-shot capabilities show potential
&lt;/span&gt;
&lt;span class="fragment " &gt;
&lt;strong&gt;Security Concern&lt;/strong&gt;: Easy alignment bypass requires safeguards
&lt;/span&gt;
&lt;span class="fragment " &gt;
&lt;strong&gt;Research Direction&lt;/strong&gt;: Focus on subtle disinformation detection
&lt;/span&gt;
&lt;/small&gt;
&lt;hr&gt;
&lt;section data-noprocess data-shortcode-slide
data-background-color="#1e3a8a"
&gt;
&lt;h2 id="fighting-fire-with-fire-1"&gt;&amp;ldquo;Fighting Fire with Fire&amp;rdquo;&lt;/h2&gt;
&lt;p&gt;&lt;small&gt;&lt;em&gt;Re-purposing LLMs as countermeasures against disinformation&lt;/em&gt;&lt;/small&gt;&lt;/p&gt;
&lt;hr&gt;
&lt;section data-noprocess data-shortcode-slide
data-background-color="#ffffff"
&gt;
&lt;h2 id="questions--resources"&gt;Questions &amp;amp; Resources&lt;/h2&gt;
&lt;small&gt;
**Code**: https://github.com/mickeymst/F3
**Paper**: EMNLP 2023 Main Conference
**Contact**: jsl5710@psu.edu
&lt;p&gt;Penn State University | PIKE Research Lab
&lt;/small&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Thank You!&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>Fighting Fire with Fire: The Dual Role of LLMs in Crafting and Detecting Elusive Disinformation</title><link>https://jsl5710.github.io/publication/conference-paper-f3/</link><pubDate>Tue, 05 Dec 2023 00:00:00 +0000</pubDate><guid>https://jsl5710.github.io/publication/conference-paper-f3/</guid><description>&lt;div class="callout flex px-4 py-3 mb-6 rounded-md border-l-4 bg-blue-100 dark:bg-blue-900 border-blue-500"
data-callout="note"
data-callout-metadata=""&gt;
&lt;span class="callout-icon pr-3 pt-1 text-blue-600 dark:text-blue-300"&gt;
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m16.862 4.487l1.687-1.688a1.875 1.875 0 1 1 2.652 2.652L6.832 19.82a4.5 4.5 0 0 1-1.897 1.13l-2.685.8l.8-2.685a4.5 4.5 0 0 1 1.13-1.897zm0 0L19.5 7.125"/&gt;&lt;/svg&gt;
&lt;/span&gt;
&lt;div class="callout-content dark:text-neutral-300"&gt;
&lt;div class="callout-title font-semibold mb-1"&gt;Note&lt;/div&gt;
&lt;div class="callout-body"&gt;Click the &lt;em&gt;Cite&lt;/em&gt; button above to demo the feature to enable visitors to import publication metadata into their reference management software.&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;div class="callout flex px-4 py-3 mb-6 rounded-md border-l-4 bg-blue-100 dark:bg-blue-900 border-blue-500"
data-callout="note"
data-callout-metadata=""&gt;
&lt;span class="callout-icon pr-3 pt-1 text-blue-600 dark:text-blue-300"&gt;
&lt;svg height="24" xmlns="http://www.w3.org/2000/svg" viewBox="0 0 24 24"&gt;&lt;path fill="none" stroke="currentColor" stroke-linecap="round" stroke-linejoin="round" stroke-width="1.5" d="m16.862 4.487l1.687-1.688a1.875 1.875 0 1 1 2.652 2.652L6.832 19.82a4.5 4.5 0 0 1-1.897 1.13l-2.685.8l.8-2.685a4.5 4.5 0 0 1 1.13-1.897zm0 0L19.5 7.125"/&gt;&lt;/svg&gt;
&lt;/span&gt;
&lt;div class="callout-content dark:text-neutral-300"&gt;
&lt;div class="callout-title font-semibold mb-1"&gt;Note&lt;/div&gt;
&lt;div class="callout-body"&gt;Create your slides in Markdown - click the &lt;em&gt;Slides&lt;/em&gt; button to check out the example.&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;Add the publication&amp;rsquo;s &lt;strong&gt;full text&lt;/strong&gt; or &lt;strong&gt;supplementary notes&lt;/strong&gt; here. You can use rich formatting such as including
.&lt;/p&gt;</description></item></channel></rss>