Skip to main content
Back to timeline
bioRxivSource publication:

A dedicated foundation model for fully human heavy-chain-only antibodies: HCAbLM learns sequence and functional grammar

Synopsis

The work first characterized fully human heavy-chain-only antibodies (HCAbs) independently of any HCAb-trained model, finding a reproducible distributional shift relative to conventional human VH domains localized predominantly to CDR1/2, CDR3 architecture and, where supported, a restricted framework region rather than widespread framework remodeling; on this basis the authors developed HCAbLM, described as the first foundation model pretrained specifically on a large-scale fully human HCAb repertoire using 31.

Source-provided article image: A foundation model learns the sequence and functional grammar of fully human heavy-chain-only antibodies
Fig. 1 ·

Fig. 1 | Molecular premise and evidence chain. a, A conventional IgG pairs a heavy-chain variable domain (VH) with a light-chain variable domain (VL), whereas a fully human HCAb lacks CH1 and a light chain, leaving the surface normally contacted by VL solvent-exposed. b, Schematic topology of the compared variable domains in IMGT framework (FR1–FR4) and complementarity-determining (CDR1– CDR3) regions. Dots indicate positions differing from the assigned V germline; the

bioRxiv · Page 4

Interpretation

The study first characterized fully human HCAbs in a model-independent way relative to conventional human VH domains, finding that the distributional shift is localized predominantly to CDR1/2, CDR3 architecture and, where supported, a restricted framework region rather than widespread framework remodeling. Prior discussion of antibody language models largely rested on generic repertoires; this work treats the sequence features of the HCAb format as an analysis object independent of HCAb-trained models, providing source-aware comparative evidence. Based on source-aware comparisons, reported as a reproducible distributional shift with an explicitly localized scope; the text does not give specific statistics or sample-size details.

The authors developed HCAbLM, described as the first foundation model pretrained specifically on a large-scale fully human HCAb repertoire, using 31.8 million sequences from 73 independently immunized HCAb mice. Unlike generic antibody language models learned from large natural repertoires, HCAbLM adopts a repertoire-specific pretraining strategy, addressing whether generic representations capture the constraints of specialized antibody formats. Grounded in an explicit sequence scale (31.8 million sequences) and number of independently immunized mice (73), constituting a methodological contribution.

HCAbLM learned a region-selective sequence-compatibility prior distinct from conventional antibody language models, and its frozen representations transferred to experimentally measured SEC purity, HIC behavior and thermal stability. The work links pretrained representations to experimentally measurable functional properties (purity, hydrophobic interaction behavior, thermal stability), not merely sequence-level fitting. Frozen representations were examined in grouped internal validation and retrospective cross-project evaluation, which is retrospective rather than prospective experimental validation.

The study proposes repertoire composition as an important biological design variable for foundation models of specialized antibody formats. It elevates the biological context of data provenance from a mere training resource to a design choice that shapes model representations. This conclusion is supported jointly by the model-independent distributional shift and the repertoire-specific pretraining results, and is an inference drawn from this work.

Perspective

The work targets the specific format of fully human heavy-chain-only antibodies, and its conclusions apply to settings that use HCAb repertoires as the training source and focus on experimentally measurable properties such as SEC purity, HIC behavior and thermal stability; whether the same holds for conventional antibodies or other specialized formats is not addressed in the text.

The loaded text is abstract-level and provides no figures, specific statistics, sample-split details or evaluation metric values; therefore the actual effect size of HCAbLM representation transfer, the concrete design of the grouped internal validation and retrospective cross-project evaluation, and the biological interpretation of the region-selective prior remain open questions for a careful reader.

Sources