π Paper accepted by NeurIPS-26

Our paper “Towards Mitigating Deceptive Safety Alignment in Large Reasoning Models” has been accepted to NeurIPS 2026. This is joint work with Saleh Zare Zade, Rafi Ibn Sultan, Alexander Kotov, and my advisor Dongxiao Zhu.
Large reasoning models are typically trained with rewards on their final answers, leaving the reasoning trace largely unsupervised. We show this produces deceptive safety alignment: a safe-looking final answer that sits on top of unsafe reasoning. We introduce DSAR, a metric that quantifies this inconsistency, and SARA, an RL method that rewards safety-aware reasoning as well as safe answers β reducing deceptive alignment under both standard prompting and prefilling attacks while preserving helpfulness.
Paper Β· Code Β· Publication page

I am a Ph.D. candidate in Computer Science at Wayne State University, advised by Prof. Dongxiao Zhu in the Trustworthy AI Lab. My research focuses on trustworthy AI: making AI systems robust and aligned so they can be deployed safely.
I work across large language models (LLMs), large reasoning models (LRMs), and agentic AI, spanning safety alignment, fine-tuning, reinforcement learning, in-context learning, and machine unlearning. My work has been published at venues including NeurIPS, ICLR, CVPR, and AAAI (oral). In industry, I worked on generative retrieval as an AI/ML engineer intern at LinkedIn.
I am always happy to connect with others working on trustworthy AI. Feel free to reach out!