Towards Mitigating Deceptive Safety Alignment in Large Reasoning Models

October 5, 2026·
Xiangyu Zhou
Xiangyu Zhou
,
Saleh Zare Zade
,
Rafi Ibn Sultan
,
Alexander Kotov
,
Dongxiao Zhu
· 0 min read
Safe final answers can mask unsafe reasoning. SARA encourages safety across both.
Abstract
Large Reasoning Models (LRMs) are commonly trained with reinforcement learning (RL) to improve their generation of chain-of-thought (CoT) reasoning before producing final answers. However, RL rewards are typically assigned based on final answers, providing little or no direct supervision over intermediate reasoning. This can lead to deceptive safety alignment, where the reasoning trace and final answer convey inconsistent safety signals. To systematically investigate this phenomenon, we introduce DSAR (Deceptive Safety Alignment Rate), a metric that jointly assesses reasoning traces and final answers to quantify their safety inconsistency. Across multiple LRMs and benchmarks, we find that deceptive safety alignment is pervasive under standard prompting conditions and is substantially amplified under prefilling attacks. We further provide a hidden representation analysis showing that models exhibit stronger safety discrimination at the final-answer stage than during intermediate reasoning. To close this gap, we propose SARA (Safety-Aware Reasoning Alignment), an RL-based method that rewards both safety-aware reasoning and safe final answers, encouraging early harmful intent recognition and enforcing reasoning-answer consistency. Experiments show that SARA significantly mitigates deceptive safety alignment under both standard and adversarial settings while preserving helpfulness and utility.
Type
Publication
The Fortieth Annual Conference on Neural Information Processing Systems
publications
Xiangyu Zhou
Authors
Ph.D. Candidate in Computer Science

I am a Ph.D. candidate in Computer Science at Wayne State University, advised by Prof. Dongxiao Zhu in the Trustworthy AI Lab. My research focuses on trustworthy AI: making AI systems robust and aligned so they can be deployed safely.

I work across large language models (LLMs), large reasoning models (LRMs), and agentic AI, spanning safety alignment, fine-tuning, reinforcement learning, in-context learning, and machine unlearning. My work has been published at venues including NeurIPS, ICLR, CVPR, and AAAI (oral). In industry, I worked on generative retrieval as an AI/ML engineer intern at LinkedIn.

I am always happy to connect with others working on trustworthy AI. Feel free to reach out!