Public articles linked to the same research event.
arXiv The work introduces Suan, a new preference optimization algorithm whose objective is formulated directly at the gradient level, bypassing the standard variational derivation and yielding more interpretable and robust training dynamics; extensive evaluations across a diverse suite of competitive baselines and benchmarks show Suan achieves superior safety alignment while fully preserving response utility.
The work introduces Suan, a new preference optimization algorithm whose objective is formulated directly at the gradient level, bypassing the standard variational derivation and yielding more interpretable and robust training dynamics; extensive evaluations across a diverse suite of competitive baselines and benchmarks show Suan achieves superior safety alignment while fully preserving response utility.
The work introduces Suan, a new preference optimization algorithm whose objective is formulated directly at the gradient level, bypassing the standard variational derivation and yielding more interpretable and robust training dynamics; extensive evaluations across a diverse suite of competitive baselines and benchmarks show Suan achieves superior safety alignment while fully preserving response utility.
The work introduces Suan, a new preference optimization algorithm whose objective is formulated directly at the gradient level, bypassing the standard variational derivation and yielding more interpretable and robust training dynamics; extensive evaluations across a diverse suite of competitive baselines and benchmarks show Suan achieves superior safety alignment while fully preserving response utility.