Skip to main content

Research timeline

Related research and updates

Public articles linked to the same research event.

arXiv

DNAlign confines safety perturbations to a harmful subspace via null-space projection, cutting harmful outputs across LLM backbones while preserving fluency and factual utility

The work proposes DNAlign, a lightweight alignment framework combining control-theoretic optimization with null-space projection: it treats the LLM as a dynamic system and applies controllable perturbations to steer generation toward safe behavior, while a projection module restricts those perturbations to a harmful-related subspace derived from neutral hidden states so general knowledge and response quality are preserved, and a value function trained on human preference data adaptively optimizes the control signals; evaluations across multiple LLM backbones show it consistently reduces harmful outputs while maintaining fluency, coherence, and factual utility, achieving superior overall performance over prior alignment baselines without sacrificing generation diversity.