<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>NeurIPS |</title><link>http://xzhou98.github.io/tags/neurips/</link><atom:link href="http://xzhou98.github.io/tags/neurips/index.xml" rel="self" type="application/rss+xml"/><description>NeurIPS</description><generator>HugoBlox Kit (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Mon, 05 Oct 2026 00:00:00 +0000</lastBuildDate><image><url>http://xzhou98.github.io/media/icon_hu_da05098ef60dc2e7.png</url><title>NeurIPS</title><link>http://xzhou98.github.io/tags/neurips/</link></image><item><title>🎉 Paper accepted by NeurIPS-26</title><link>http://xzhou98.github.io/blog/neurips-26/</link><pubDate>Mon, 05 Oct 2026 00:00:00 +0000</pubDate><guid>http://xzhou98.github.io/blog/neurips-26/</guid><description>&lt;p&gt;Our paper &lt;strong&gt;&amp;ldquo;Towards Mitigating Deceptive Safety Alignment in Large Reasoning Models&amp;rdquo;&lt;/strong&gt; has been accepted to
&lt;strong&gt;NeurIPS 2026&lt;/strong&gt;. This is joint work with Saleh Zare Zade, Rafi Ibn Sultan, Alexander Kotov, and my advisor
Dongxiao Zhu.&lt;/p&gt;
&lt;p&gt;Large reasoning models are typically trained with rewards on their final answers, leaving the reasoning
trace largely unsupervised. We show this produces &lt;em&gt;deceptive safety alignment&lt;/em&gt;: a safe-looking final answer
that sits on top of unsafe reasoning. We introduce &lt;strong&gt;DSAR&lt;/strong&gt;, a metric that quantifies this inconsistency, and
&lt;strong&gt;SARA&lt;/strong&gt;, an RL method that rewards safety-aware reasoning as well as safe answers — reducing deceptive
alignment under both standard prompting and prefilling attacks while preserving helpfulness.&lt;/p&gt;
&lt;p&gt;
·
·
&lt;/p&gt;</description></item></channel></rss>