Skip to main content
Back to timeline
arXivSource publication:

Socio-Foundation lifts an 8B model 11 points over Qwen3-8B on individual behavior simulation via three-stage hierarchical distillation

Synopsis

The work introduces the FONTS taxonomy of five individual-simulation capabilities, curates a roughly 10M-instance corpus from 14 datasets, and builds Socio-Foundation through a three-stage post-training pipeline: DAPO-trained task experts, capability experts consolidated by off-policy distillation, and a unified student via multi-teacher on-policy distillation; it scores 61.7 on the authors' IndiEval, 11.0 points above the Qwen3-8B base, and beats that base on three out-of-distribution benchmarks.

Source-provided article image: Socio-Foundation: A Model for Generalizable Individual Behavior Simulation via Hierarchical Capability Distillation
Figure 1 ·

Figure 1: Three approaches to individual simulation. (a) General-purpose models may flatten distinct personas. (b) Task-driven post-training yields fragmented skills. (c) Individual simulation Foundation model integrates complementary capabilities.

arXiv

Interpretation

It proposes the FONTS taxonomy, splitting individual simulation into five complementary capability dimensions (persona fidelity, outcome realization, behavioral naturalness, trajectory coherence, social grounding), and uses it to select 14 datasets from more than 40 surveyed, yielding a training corpus library of about 10 million instances. Individual simulation had been scattered across role-play, user simulation, and cognition-modeling settings without a shared capability structure; FONTS formalizes their common behavioral requirements so heterogeneous tasks can be organized by capability rather than by dataset. The taxonomy comes from two steps: bottom-up annotation and clustering (19 labels, eight preliminary clusters) and top-down consolidation guided by McAdams and Pals' New Big Five. The appendix lists 44 surveyed dataset/task entries plus 10 supplementary benchmark entries with per-entry annotations, and notes that one dataset can map to several dimensions under different task interpretations.

It develops a three-stage post-training pipeline: DAPO-trained task experts in isolated environments, consolidation into five capability experts by off-policy distillation, and unification into a single student by multi-teacher on-policy distillation (MOPD) on the student's own generated prefixes. Compared with pooling multi-task data or merging expert weights directly, the pipeline decouples specialization from integration: conflicting behavioral objectives mature separately, are aggregated at the capability level, and are then matched on on-policy prefixes to reduce the mismatch between distillation and autoregressive inference. Ablations show the full pipeline attains the highest average (61.7); replacing final-stage on-policy distillation with capability-expert off-policy distillation lowers it to 59.3, capability-expert averaging to 56.8, task-expert averaging to 54.5, task-expert SFT and task-expert OPD to 59.5 and 54.2, pooled-data SFT to 43.0, and reverse-KL OPD to 50.9. Removing any single capability expert lowers the overall average, with the largest drop when the N expert is removed.

It builds IndiEval, mapping 29 native metrics across 13 benchmark entries onto the five FONTS dimensions while preserving each benchmark's original evaluation protocol; Socio-Foundation improves on 24 of 29 metrics over Qwen3-8B, with an average of 61.7 versus 50.7. Individual-simulation evaluation had been fragmented across task-specific metrics that were hard to compare; IndiEval offers a cross-task capability view without altering native protocols and separates in-domain from out-of-distribution components. The main table covers 13 evaluation entries totaling 6,297 cases, with AgentSense and UserLM-LiC forming the OOD component. Three further benchmarks (Mistakes, HiToM, τ-USI) serve as additional OOD evaluation, where Socio-Foundation scores 57.00, 62.00, and 69.31, above both OSim 8B and Qwen3-8B.

In a repeated-evaluation stability study, five repeated runs of a fixed checkpoint give run-to-run standard deviations below 0.6 percentage points for LifeChoices accuracy, HUMANUAL response/state alignment, and MirrorBench GTEval, while other metrics in the same benchmark, such as PI and RNR, vary more. The analysis organizes evaluation pipelines into automatic scoring, LLM-based scoring, and interactive evaluation, and uses fixed human reference conversations as a control to observe variability in the auxiliary scoring component, giving context for interpreting small score differences. The five runs comprise 6,500 case evaluations over 1,300 distinct test cases, with checkpoint, evaluation seed, prompts, test cases, and scoring configuration held fixed; human-reference RNR scores across runs are 90.0, 92.0, 90.0, 90.5, and 91.0 (mean 90.700, standard deviation 0.837). The authors state explicitly that this study does not establish statistical significance for model differences.

Perspective

The result targets research and applications that use LLMs to reproduce how a specific individual communicates, decides, and responds in social contexts, including persona-grounded dialogue, user-behavior prediction, social-science experimentation, and multi-agent role-play. Methodologically, the pipeline is implemented on a Qwen3-8B backbone with rank-32 LoRA adapters, bfloat16 training, a 16,384-token context window, and disabled thinking mode; task experts are trained with DAPO for 200 steps, capability experts receive 200 updates each, and the final student has a 200-step budget, with DeepSeek-V4-Flash serving as both auxiliary dialogue model and reward/judge model. For evaluation, IndiEval preserves each benchmark's native protocol, the average weights F/O/N/T/S equally, and within a dimension metrics are averaged within a dataset before datasets are weighted equally; in-domain and OOD components (AgentSense, UserLM-LiC) both enter the average, while Mistakes, HiToM, and τ-USI do not. The repeated-evaluation stability study concerns a fixed checkpoint and is meant to contextualize small score differences.

The mapping between FONTS dimensions and datasets comes from the authors' annotation and clustering, and the appendix notes that one dataset can fall under different dimensions under different task interpretations and that these assignments are the authors' capability organization rather than category labels from the original dataset authors, so dimension boundaries may be redrawn in other uses. The average weights the five dimensions equally while the number of dataset entries per dimension is 8/3/3/2/3, so differences in dataset counts within a dimension shape the average. The repeated-evaluation stability study covers only the LifeChoices, HUMANUAL, and MirrorBench pathways; the authors state it is not a representative sample of all IndiEval tasks and is not used to establish statistical significance for model differences, and isolating judge variability would require rescoring frozen responses or dialogue trajectories. In addition, some metrics in the main and ablation tables take extreme values such as 0.0 for particular models, and the text does not unpack each such case individually.

Sources