Skip to main content
Back to timeline
arXivSource publication:

Learning Latent Protein Languages: PLLM Raises Compute-Scaling Exponent to 0.038 and SLL Cuts Validation Perplexity by 34%

Synopsis

This work introduces two learned discrete latent protein languages: PLL, a 4,096-state contextual alphabet built on a frozen ESM-2 encoder on the sequence side, and SLL, which adapts GCP-VQVAE Lite with auxiliary sequence and confidence supervision while remaining decodable to backbone coordinates; under matched downstream training, PLLM has a fitted compute-scaling exponent of 0.038 versus 0.020 for the amino acid autoregressive model and reduces the fraction of samples below a 1.5-bit residue-composition entropy threshold by 54%, while replacing the original tokenizer with SLL reduces best validation perplexity by 34% for sequence-to-structure prediction.

Source-provided article image: Learning Latent Protein Languages for Autoregressive Generation
Figure 1 ·

Figure 1: PLL construction. Stage 1 trains a vector-quantized variational autoencoder (VQ-VAE)-style latent reconstructor for frozen ESM-2 representations. Stage 2 finalizes PLL as a VQ-RAE by freezing the latent encoder and codebook and training a reinitialized transformer decoder with a linear residue-classification head.

arXiv

Interpretation

PLL re-alphabetizes protein sequences into a contextual language with one discrete code per residue, reconstructing the frozen ESM-2 representation closely and recovering amino acid identities at high accuracy. Unlike direct amino acid symbols, PLL codes can distinguish contextual states while preserving sequence length and residue alignment, and decoding does not require ESM-2. Stage 1 reaches median NMSE 0.044, reconstruction KL 0.003, and cosine similarity 0.970; Stage 2 achieves 99.81% micro token-level residue recovery on the UniRef50 validation split, with all standard amino acids above 99.17% class-wise recovery.

Under matched downstream sequence training, PLLM has a steeper fitted compute-scaling exponent than the amino acid autoregressive model and substantially reduces low-complexity drift in unconditional generation. The comparison uses the same UniRef50 sequences, token budget, and optimization settings, changing only the target representation, so the scaling and drift differences are attributed to the complete PLL target representation. PLLM has a fitted exponent of 0.038 versus 0.020 for the amino acid model, roughly 1.9 steeper; across the temperature sweep, the fraction of samples below the 1.5-bit entropy threshold drops from 43.0% for the amino acid model to 19.9% for PLLM.

SLL improves the semantic utility of structure codes while remaining decodable to backbones, and improves autoregressive target predictability for sequence-to-structure prediction. Keeping the GCP-VQVAE Lite architecture and training corpus fixed, only the tokenizer objective changes with auxiliary heads, so the structure codes carry sequence and confidence information. The combined recipe improves over the base tokenizer by 24.2% in functional-site AUROC and 35.0% in physicochemical Spearman on supervised PST; in a matched Prot2Token-style setup, changing only the target tokenizer reduces best validation perplexity by 34% and reaches a comparable training-loss regime 49% faster when measured by training epoch.

SLLM can be used for sequence-to-structure prediction and can be scaled at inference time using latent-token confidence to improve candidate quality. A DeepConf-style confidence selector combined with token-space consensus selects candidates before coordinate decoding, without an external structure-quality metric. On the CASP14/15/16 tuning set, DeepConf selection improves average TM-score from 0.564 for a single sample to 0.589, while oracle best@64 reaches 0.667; on CAMEO-2024 it improves from 0.725 to 0.754, while oracle best@256 reaches 0.825.

Perspective

The work targets single-chain monomers, using UniRef50 on the sequence side and AFDB representative structures plus the AFDBUniRef50 structure-token corpus on the structure side; PLL and SLL are pretrained separately, and sequence-to-structure prediction is conditioned on E1 sequence representations. For researchers aiming to apply the autoregressive recipe to protein generation, PLL provides a decodable sequence target language and SLL provides a decodable backbone target language, with demonstrated feasibility of candidate sampling and confidence-based selection in latent token space. The results apply to the reported single-chain, matched-training settings and evaluated benchmarks.

The scaling exponents summarize fitted lower-loss frontiers under a zero-offset power law, and fit confidence intervals or sensitivity to an additive irreducible-loss term are not reported; the amino acid-PLL comparison evaluates the complete PLL target representation and does not isolate the effects of ESM-2 semantics, vector quantization, vocabulary expansion, or learned tokenizer transformations. Generation benchmarks compare across pretrained systems whose training data, model sizes, and training modalities are not matched. SLLM remains below specialized backbone generators on self-consistency quality, with its strengths in diversity and novelty. Multichain complexes, joint PLL-SLL co-generation, and alternatives such as continuous-target generation or ESM-2 fine-tuning and distillation remain untested.

Sources