Skip to main content
Back to timeline
arXivSource publication:

PersonaDose turns persona intensity into a requestable score: 17.8–33.2-point core-trait gains across three model families and 4.7–6.2-point targeting error on 14–22 of 28 reachable targets

Synopsis

The work introduces persona dosing, controlling a language model through a trait description plus a requested mean intensity, by specializing a shared, description-conditioned FLAS controller on persona responses and then calibrating its flow time against measured trait expression; across Llama-3.1-8B, Qwen3-8B, and Gemma-3-4B, PersonaDose raises core-trait expression at the Persona Vectors coherence floor of 75 by 33.2, 18.3, and 17.8 points over contrastive activation addition, calibration-selected settings retain an expression advantage on held-out questions, and calibrated requests over seven trained traits yield mean targeting errors of 4.7–6.2 points over 14–22 reachable targets out of 28 per model.

AI-generated editorial illustration: Persona Dosing: Calibrated Activation Steering for Graded Trait Control

Interpretation

It translates an activation intervention's strength parameter into behavioral score units: the trait description selects what to express and flow time selects how strongly, training does not pair responses with requested target intensities, and calibration after training inverts a measured dose–response curve to map a requested score to a flow time. Earlier activation interventions (CAA, RepE, representation engineering) expose a strength coefficient with no intrinsic behavioral meaning, and the same coefficient induces different expression levels across traits and models; this work separates the behavioral range a controller can reach from the precision with which a researcher can request a level inside it, and exposes the interface in behavioral units. The paper formalizes response curves, a coherence floor, and intensity targeting, and evaluates across three model families under one scoring protocol (GPT-4.1-mini scoring trait expression and coherence independently on a 0–100 scale); calibration uses isotonic regression with piecewise-linear inversion, and targets outside the fitted range are marked unreachable before test evaluation.

Persona-response specialization substantially expands expression under high coherence: at mean coherence 75, PersonaDose reaches core-trait expression of 75.1, 82.3, and 86.0 versus CAA's 41.9, 64.0, and 68.2, gains of 33.2, 18.3, and 17.8 points, while generic FLAS reaches 13.3, 46.2, and 42.7. The gains concentrate where direction baselines achieve little expression inside the high-coherence range: on Llama, sycophantic rises from 9.7 to 90.5 and impolite from 1.6 to 72.3; on Qwen and Gemma the largest improvement is on evil, from 3.3 to 71.5 and from 21.2 to 68.9. Tables 1 and 2 report peak expression across three models and five methods, with Appendix D giving all seven per-trait results; strength sweeps use the full 20-question set with ten responses per question, and intervals use 2,000 question-bootstrap resamples.

The expression advantage persists on held-out questions, but the calibration coherence floor does not transfer for every trait: settings selected on ten calibration questions give held-out mean expression of 75.6, 80.3, and 84.8 (CAA: 43.5, 63.3, 68.0), paired gains of 32.1 [26.4, 37.8], 17.0 [10.4, 23.7], and 16.8 [12.4, 21.5], with mean coherence 75.5, 83.2, and 79.4 and only 1/3, 2/3, and 3/3 traits retaining the floor on test. Separating calibration selection from test evaluation shows that expression gains transfer while the coherence constraint must be rechecked per trait on held-out questions, two things earlier work reporting sweep peaks did not separate. Held-out evaluation holds calibration policies fixed and uses paired bootstrap intervals sharing question draws; the paper notes that sweep peaks select and evaluate strengths on the same questions, whereas the held-out test separates these steps.

Calibrated intermediate-intensity requests are realizable on held-out questions, but coverage and accuracy are different questions: over seven trained traits mean targeting MAE is 6.1, 6.2, and 4.7 points with 22/28, 20/28, and 14/28 reachable requests, of which 19/28, 20/28, and 13/28 also meet the test coherence floor; after continued training of the Qwen controller, held-out expression rises from 78.7 to 93.7 (paired gain 15.0 [7.3, 23.3]) and its recalibrated map supports 11/12 coherent requests versus 8/12 for calibrated CAA. It reports the behavioral range a controller learns and the accuracy of requests within that range as separate metrics, and shows that a controller update changes the calibration map (the same sycophancy 60 selects different strengths for the initial and continued checkpoints), motivating remeasurement after an update. Tables 4, 5, and 6 with Appendices F and G give per-cell flow time, expression, absolute error, and coherence; the paper also states that the two MAE figures cover different request sets and therefore do not establish a paired targeting-accuracy advantage, and that secondary-judge MAE intervals include zero.

Perspective

The result is aimed at researchers who need to study persona traits at controlled intensities in a controlled research setting: within a model the seven traits share one controller, so switching traits does not require loading a different adapter, while different base models have separate controllers; a requested score specifies mean expression across questions and sampled responses, calibration is valid only within the fitted range, and out-of-range targets are marked unreachable before test evaluation. It makes score-based intensity requests operational on Llama-3.1-8B, Qwen3-8B, and Gemma-3-4B, and indicates that the calibration map should be remeasured after a controller update.

The evaluation covers seven trained traits, short responses, and PV judge scores, and does not establish generalization to unseen traits or long conversations; the number of reachable intensity levels varies by trait, with optimistic having only one reachable target on each model and Gemma hallucinating having none; evaluated strengths are nonnegative, so targets below the fitted unsteered expression level can fall outside the range; coherence measures intelligibility, not factual correctness or safety; sweep peaks select and evaluate strengths on the same questions, and although the held-out test separates these steps, a calibration coherence floor need not hold for every test trait; bootstrap intervals condition on fixed controllers and do not capture training or judge variability; the two MAE figures cover different request sets and do not establish a paired advantage; and controller continuation changes several training choices together, so its effects cannot be attributed to any single change.

Sources