Accepted by NeurIPS-2026

Towards Mitigating Deceptive Safety Alignment in Large Reasoning Models featured image

Towards Mitigating Deceptive Safety Alignment in Large Reasoning Models

Reasoning models hide unsafe thoughts behind safe answers; we propose a metric DSAR to measure it and safety alignment method SARA to mitigate it.

avatar
Xiangyu Zhou
•