Public articles linked to the same research event.
arXiv The work proposes DNAlign, a lightweight alignment framework combining control-theoretic optimization with null-space projection: it treats the LLM as a dynamic system and applies controllable perturbations to steer generation toward safe behavior, while a projection module restricts those perturbations to a harmful-related subspace derived from neutral hidden states so general knowledge and response quality are preserved, and a value function trained on human preference data adaptively optimizes the control signals; evaluations across multiple LLM backbones show it consistently reduces harmful outputs while maintaining fluency, coherence, and factual utility, achieving superior overall performance over prior alignment baselines without sacrificing generation diversity.
The work proposes DNAlign, a lightweight alignment framework combining control-theoretic optimization with null-space projection: it treats the LLM as a dynamic system and applies controllable perturbations to steer generation toward safe behavior, while a projection module restricts those perturbations to a harmful-related subspace derived from neutral hidden states so general knowledge and response quality are preserved, and a value function trained on human preference data adaptively optimizes the control signals; evaluations across multiple LLM backbones show it consistently reduces harmful outputs while maintaining fluency, coherence, and factual utility, achieving superior overall performance over prior alignment baselines without sacrificing generation diversity.
The work proposes DNAlign, a lightweight alignment framework combining control-theoretic optimization with null-space projection: it treats the LLM as a dynamic system and applies controllable perturbations to steer generation toward safe behavior, while a projection module restricts those perturbations to a harmful-related subspace derived from neutral hidden states so general knowledge and response quality are preserved, and a value function trained on human preference data adaptively optimizes the control signals; evaluations across multiple LLM backbones show it consistently reduces harmful outputs while maintaining fluency, coherence, and factual utility, achieving superior overall performance over prior alignment baselines without sacrificing generation diversity.
The work proposes DNAlign, a lightweight alignment framework combining control-theoretic optimization with null-space projection: it treats the LLM as a dynamic system and applies controllable perturbations to steer generation toward safe behavior, while a projection module restricts those perturbations to a harmful-related subspace derived from neutral hidden states so general knowledge and response quality are preserved, and a value function trained on human preference data adaptively optimizes the control signals; evaluations across multiple LLM backbones show it consistently reduces harmful outputs while maintaining fluency, coherence, and factual utility, achieving superior overall performance over prior alignment baselines without sacrificing generation diversity.