<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Audio Language Models | Jason S. Lucas</title><link>https://jsl5710.github.io/tag/audio-language-models/</link><atom:link href="https://jsl5710.github.io/tag/audio-language-models/index.xml" rel="self" type="application/rss+xml"/><description>Audio Language Models</description><generator>HugoBlox Kit (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Sun, 20 Sep 2026 00:00:00 +0000</lastBuildDate><image><url>https://jsl5710.github.io/media/icon_hu_1b2044c02ce09a43.png</url><title>Audio Language Models</title><link>https://jsl5710.github.io/tag/audio-language-models/</link></image><item><title>AOR-Bench: Do Large Audio Language Models Over-Refuse Pseudo-Harmful Queries?</title><link>https://jsl5710.github.io/publication/conference-paper-aor-bench/</link><pubDate>Sun, 20 Sep 2026 00:00:00 +0000</pubDate><guid>https://jsl5710.github.io/publication/conference-paper-aor-bench/</guid><description>&lt;p&gt;Accepted to &lt;strong&gt;EMNLP 2026&lt;/strong&gt;, Budapest, Hungary, 24–29 October 2026.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;AOR-Bench&lt;/strong&gt; asks a question that safety evaluation usually skips: not whether a
model refuses harmful requests, but whether it refuses &lt;em&gt;safe&lt;/em&gt; ones that happen to
look harmful.&lt;/p&gt;
&lt;p&gt;Over-refusal is a real cost of alignment. A model that declines benign queries is
less useful, and the burden does not fall evenly — speakers whose accent, dialect or
phrasing sits further from the training distribution are more likely to be refused.
In the audio setting that risk grows, because prosody and acoustic ambiguity give the
model more ways to misread intent than text alone provides.&lt;/p&gt;</description></item><item><title>Lost in Speech: Trilingual Spoken Hallucination Detection Across Audio and Transcripts</title><link>https://jsl5710.github.io/publication/workshop-paper-lost-in-speech/</link><pubDate>Sun, 20 Sep 2026 00:00:00 +0000</pubDate><guid>https://jsl5710.github.io/publication/workshop-paper-lost-in-speech/</guid><description>&lt;p&gt;Accepted to the &lt;strong&gt;2nd Workshop on Speech and Audio Language Models (SALMA 2026)&lt;/strong&gt;,
co-located with EMNLP 2026 in Budapest, Hungary, 24–29 October 2026.&lt;/p&gt;
&lt;p&gt;Most hallucination benchmarks assume text. &lt;strong&gt;Lost in Speech&lt;/strong&gt; asks what changes when
the claim arrives as audio — and works across three languages rather than one, since
the answer depends on how well speech recognition serves each of them.&lt;/p&gt;
&lt;p&gt;The comparison between audio and transcript is the point. Transcription throws away
prosody, hesitation and disfluency that may carry signal about whether a speaker is
fabricating, while adding recognition errors of its own. A detector evaluated only on
transcripts can therefore look better or worse than it really is, and that gap is
widest exactly where ASR is weakest — the low-resource languages that most need the
detector to work.&lt;/p&gt;</description></item></channel></rss>