Deceptive Safety Alignment

Towards Mitigating Deceptive Safety Alignment in Large Reasoning Models featured image

Towards Mitigating Deceptive Safety Alignment in Large Reasoning Models

Reasoning models hide unsafe thoughts behind safe answers; we propose a metric DSAR to measure it and safety alignment method SARA to mitigate it.

avatar
Xiangyu Zhou
•