Skip to main content
Back to timeline
Nature NewsSource publication:

Merriam-Webster adds agentic and vibe coding as Nature reports Google DeepMind's SynthIDBio watermarking AI-generated proteins

Synopsis

Nature reports that Merriam-Webster added AI and computing terms such as agentic, AGI, natural language processing, prompt engineering and vibe coding, plus technical terms including biohack, direct air capture, geoengineering and uncanny valley, with lexicographer Peter Sokolowski noting the advanced knowledge these entries require and cognitive scientist Lera Boroditsky linking rapid technological change to language change; the same evidence bundle carries a Nature paper introducing SynthIDBio, which adapts SynthID-text's tournament sampling into ProteinMPNN to watermark protein sequences and fine-tunes AlphaFold 3's diffusion module to watermark structures, with in vitro validation showing no significant population-level difference in binding affinity for SARS-CoV-2 RBD, VEGF-A and PD-L

AI-generated editorial illustration: Sign of the times: AI lingo muscles into esteemed dictionary

Interpretation

Merriam-Webster's online dictionary added a batch of AI and technology entries in September, including agentic, AGI, compute as a noun, natural language processing, prompt engineering and vibe coding, alongside biohack, direct air capture, geoengineering and uncanny valley. Lexicographer Peter Sokolowski describes this batch as requiring advanced knowledge to explain, saying you cannot explain vibe coding without knowing what programming is, and attributes it to non-specialists' expanded access to highly technical terms through digital communication. Based on the word list and Sokolowski's quoted remarks in the Nature news story, which is reportorial evidence about language use rather than experimental data.

Cognitive scientist Lera Boroditsky argues that technological and scientific innovation often drives language change, that today's rate of technological advancement is comparable to or greater than the industrial revolution's, and that more complex societies need more tightly packed units of meaning. This supplies a language-and-cognition framing that links dictionary expansion to the pace of technological change. Commentary from a researcher on language and cognition; the story reports no specific experiment, sample or statistic.

A Google DeepMind team introduces SynthIDBio: SynthIDBio-sequence integrates SynthID-text's tournament sampling into ProteinMPNN, embedding a watermark during sampling and filtering on a g-value threshold calibrated to a target false positive rate, while SynthIDBio-structure fine-tunes AlphaFold 3's diffusion module jointly with a PointNet-inspired watermark detector. The paper states that concurrent protein-sequence watermarking work was limited to in silico experiments on wild-type monomer structures from the PDB, and that structure watermarking methods caused a noticeable accuracy decrease, so function-preserving watermarking was not yet available. Method details appear in the main text and Methods, including H=4, L=25, temperatures 0.1/0.3/0.5, loss weight w_wm=0.1, noise scales s in {0.001, 0.01, 0.1} Å, and a false-positive-rate calibration set of nearly 3.5 million sequences.

In vitro validation indicates the watermark did not disrupt binding: low nanomolar binders for SC2RBD and subnanomolar binders for VEGF-A and PD-L1, with no significant population-level differences in binding affinity distributions or hit rates between watermarked and non-watermarked designs and no detectable binding for all nine negative controls; structure watermarking exceeds 99.8% true positive rate at 0.1% false positive rate, and at s=0.001 LDDT and template modelling scores are not lower than the AlphaFold 3 baseline. This moves watermarking from in silico analysis toward wet-lab demonstration of preserved function, and reports structure watermark detection across proteins, RNA and DNA. Sequence results use surface plasmon resonance dissociation constants with two-tailed Wilcoxon rank-sum and Fisher's exact tests, 15 backbones per target and six designs per setting; structure results report true positive rates plus LDDT and template modelling scores on the AlphaFold 3 evaluation set.

Perspective

The work targets provenance for AI-generated protein sequences and structures: SynthIDBio-sequence applies to de novo design pipelines resequenced through autoregressive models such as ProteinMPNN, and SynthIDBio-structure applies to pipelines predicting structures with AlphaFold 3-style diffusion models, with s=0.001 recommended. The paper envisions nucleic acid synthesis providers, institutions managing databases such as the PDB, UniProt and GenBank, and developers of hosted design tools as potential users; the authors stress this is a technical proof-of-concept whose operationalization requires secure key sharing, tuning of false-positive-rate thresholds, and industry-wide coordination and standardization. The dictionary portion speaks to general readers and language researchers about technical vocabulary entering mainstream usage.

The paper states that SynthIDBio-sequence is a zero-bit watermark that cannot encode more information and is susceptible to removal by ProteinMPNN resequencing; a resequencing attack on 38,396 binders reduced estimated hit rates to 97% (SC2RBD), 70% (PD-L1) and 66% (VEGF-A) with structure-based filters, and lower without them. The structure watermark is robust to rigid transformations and noise but is destroyed by constrained relaxation using OpenMM with the Amber99sb force field, and differentiability from other structure watermarks has not been studied. The paper also notes that punitive use of the watermark is limited by the wide availability of open-source tools, since a malicious actor could simply use a non-watermarking tool. In addition, the news story in this bundle is an excerpt that breaks off at 'Rapid spread', so the full list of new dictionary entries and the remainder of the discussion cannot be confirmed from the loaded text; the paper portion is an excerpt of abstract, main text and methods, with full thresholds and attack results in the Supplementary Information, so the figures above reflect what the paper reports rather than an independent check.

Sources