Skip to main content
Back to timeline
bioRxivSource publication:

nanorepertoire: an end-to-end Nextflow pipeline for nanobody repertoire analysis

Synopsis

This work presents nanorepertoire, an end-to-end Nextflow DSL2 pipeline dedicated to camelid VHH nanobody repertoires that takes paired-end AIRR-seq FASTQ files through quality control, adapter trimming, read merging, in-silico translation, CD-HIT clonotyping and deep-learning CDR3 annotation with nanoCDR-X, and returns an interactive HTML report covering clonal architecture, CDR3 length and amino-acid composition, intra-clonal homogeneity, repertoire diversity and the computational carbon footprint of the run; applied to two publicly available SARS-CoV-2 RBD-selected llama libraries (4.

AI-generated editorial illustration: nanorepertoire: an end-to-end Nextflow pipeline for nanobody repertoire analysis

Interpretation

It introduces an end-to-end Nextflow DSL2 pipeline dedicated to camelid VHH nanobody repertoires, consolidating analyses that were previously assembled from standalone scripts into a single standardised, reproducible workflow. The text states that VHH analyses are "typically assembled ad hoc from standalone scripts, which limits standardisation and reproducibility across laboratories"; the pipeline offers a unified implementation in response. Evidence rests on the pipeline design and its open release (MIT licence, Zenodo archive, unchanged operation on local, HPC and cloud infrastructures), i.e. a tooling and engineering contribution.

The pipeline chains quality control, adapter trimming, read merging, in-silico translation, CD-HIT clonotyping and deep-learning CDR3 annotation with nanoCDR-X, and produces an interactive HTML report. The report describes clonal architecture, CDR3 length and amino-acid composition, intra-clonal homogeneity and repertoire diversity, and additionally reports the run's "computational carbon footprint", presenting analytical results alongside the computational cost of the run. Evidence is the explicit enumeration of pipeline steps and the described report contents; nanoCDR-X is cited as Bagordo et al., 2026.

Run on two publicly available SARS-CoV-2 RBD-selected llama libraries sampled before and after phage-display enrichment (4.8 million paired-end reads in total), it completed in 59 min on a 16-vCPU cloud instance, recovered 41,363 distinct CDR3 paratopes, and reproduced the expected contraction of clonal diversity upon selection. It demonstrates end-to-end usability on real public data and reproduces the expected biological phenomenon of diversity decline under selection pressure. Evidence comes from an actual run on two public libraries, with concrete figures for total reads, runtime, instance specification and CDR3 count; this is demonstrative application evidence.

Perspective

The pipeline targets AIRR-seq data analysis of camelid VHH nanobody repertoires and is intended for local, HPC and cloud settings; its demonstrated scenario is two publicly available SARS-CoV-2 RBD-selected llama libraries sampled before and after phage-display enrichment. For other species, other antibody formats or other sequencing designs, the text provides no validation, so separate assessment under those settings would be needed.

The loaded text is at the abstract level and provides no figures, parameter details or statistical testing, so the specific thresholds of each step, how report metrics are computed, and the accounting basis for the carbon footprint would still require the full text and code repository; how the pipeline performs on more libraries, at different sequencing depths and in other species, and how stable nanoCDR-X annotation is across datasets, remain open questions worth watching.

Sources