Skip to main content
Back to timeline
arXivSource publication:

CogAdapt uses human EEG and gaze to pick which code-model blocks to adapt, gaining 10.86 points on LiveCodeBench over all-block fine-tuning while cutting about 87% of gradient-eligible adaptation parameters

Related research and updates

Synopsis

CogAdapt derives program-level and token-level priors from human EEG and eye-tracking data recorded during code reading, combines them with a frozen MoE model's response to each coding task to dynamically select a small set of transformer blocks for sparse fine-tuning, and reports consistent correspondence between human reading behavior and MoE computation across Qwen and GLM, best pass@1 on LiveCodeBench and BigCodeBench, gains of 10.86 and 6.29 percentage points over matched regular all-block fine-tuning on LiveCodeBench, and an 86.21–87.21% reduction in gradient-eligible adaptation parameters.

Source-provided article image: CogAdapt: Cognition-informed Sparse Adaptation of Code LLMs
Figure 1 ·

Figure 1. Human attentional hotspots coincide with MoE computation at the same code regions

arXiv

Interpretation

Human attention and EEG theta responses during code reading are positively correlated with MoE code-model computation at the same semantic code regions. Prior work largely used cognitive signals for token-level loss weighting or scanpath simulation; this work systematically tests correspondence between human responses and MoE routing and representation changes at the same code locations. On the NoviceVsExpert corpus (37 participants, 32 Java programs, 81 semantic code regions), 75.0–83.3% of analyzed regions show positive Spearman correlations across two MoE architectures and four code categories.

Human processing demand corresponds to MoE computation in a structured way across model depth, with different internal signals carrying this correspondence in different architectures. This indicates there is no single universally aligned layer, motivating task-dependent block selection rather than a fixed layer or all layers. Depth–time heatmaps compare five model-computation signals with EEG theta over 12 reading stages while controlling for program length; Qwen shows clearer correspondence in hidden-state shift and residual integration, GLM in expert-choice confidence and MoE write.

Using cognitive priors for task-dependent sparse adaptation improves code-generation accuracy while updating only about six blocks. Unlike all-block fine-tuning, fixed CKA-based selection, or random selection, CogAdapt selects different blocks per task and additionally weights training examples by an EEG-derived demand prior and code tokens by an attention prior. On Qwen3-Coder-30B-A3B-Instruct and GLM-4.7-Flash, LiveCodeBench pass@1 rises from 18.86% to 29.71% (Qwen) and from 17.71% to 24.00% (GLM); BigCodeBench reaches 37.73% and 31.82%, best or tied-best in all four settings.

Cognition-informed sparse adaptation reduces gradient-eligible parameters while also modestly lowering training time and GPU energy. Ablations show the gains are not explained by sparse fine-tuning alone: all-block, fixed random blocks, and model-only block selection all trail the full method, and removing EEG or gaze guidance lowers performance in most settings. About 6.1–6.6 blocks are selected per task versus 47–48, reducing gradient-eligible adaptation parameters from 25.95–27.06M to 3.32–3.73M (86.21–87.21% reduction), with 2.53–5.87% lower wall-clock training time and 2.12–4.55% lower estimated GPU energy.

Perspective

The work targets adaptation of MoE-based code-generation models in settings where a transferable human reading corpus exists and downstream tasks are Python code generation; priors are built once before training, and inference requires no EEG, eye-tracking, or reference solution, so it fits standard deployment. Its conclusions pertain to the two studied backbones (Qwen and GLM) and the two benchmarks (LiveCodeBench and BigCodeBench), and the approach can be extended by follow-up work to other models and tasks.

The human–model correspondence is carried by different internal signals across architectures, and how this varies in further model families remains open; transferring priors from Java reading data to Python tasks relies on syntactic correspondence, whose coverage deserves further study; and although adaptation parameters drop sharply, the full backbone still runs in the forward pass, so time and energy reductions are much smaller than the parameter reduction, leaving it an open direction to convert adaptation sparsity into computational sparsity such as conditional block activation or early exiting.

Sources