CodeGraph annotates 145 million source files into a knowledge graph with about 1 billion typed edges and grounds its concepts in Wikidata
Synopsis
The work presents a pipeline that uses a code-specialised LLM (based on Qwen3-Coder-30B-A3B-Instruct) to annotate source files under an open taxonomy, extracting algorithms, paradigms, design patterns, and application domains, then grounds these in Wikidata through a three-stage procedure (deterministic SPARQL, a Deep Research Agent for the long tail, and parent-of hierarchy rollup), with a calibrated quality-assurance protocol combining a human gold set and an LLM-as-a-judge filter; applied to the Stack-Edu corpus it yields CodeGraph with roughly 158 million nodes (about 145 million file nodes, about 63,000 extracted concept entities, and roughly 19,800 grounded Wikidata entities) and about 1 billion typed edges across 14 programming languages.
Interpretation
It builds an open-taxonomy semantic annotation pipeline for source code that lifts file-level metadata from lexical and syntactic signals to four high-level semantic axes: algorithms, paradigms, design patterns, and application domains. Existing code resources are either very large corpora exposing only the raw token stream and a language label (such as The Stack v2) or smaller resources with task-restricted supervision (such as CodeSearchNet function-level docstrings or Project CodeNet problem IDs and execution metadata); this work instead has the model generate free-form labels and reconciles them a posteriori, adding a semantic layer at pre-training scale. It was actually run over 145 million files of Stack-Edu (14 programming languages, with the Markdown subset excluded), consuming approximately 400,000 GPU-hours on 256 NVIDIA A100 GPUs within the LEONARDO supercomputer; prompts are decomposed along four orthogonal axes and enforce strict JSON schemas that require empty arrays when a file offers no support, turning silent omissions into explicit negatives.
It proposes a three-stage Wikidata grounding procedure that gives the locally extracted vocabulary a shared semantic frame of reference and a multi-level taxonomy. Rather than delegating all disambiguation to a model, it first uses deterministic SPARQL with a class-hierarchy filter for unambiguous concepts, hands only ambiguous or empty residuals to a Deep Research Agent backed by a separate open-weight model (Qwen3.6-27B), and finally imports the parent-of closure via a two-hop wdt:P279 traversal, producing a dual conceptual layer of a local corpus-defined vocabulary and a global Wikidata-grounded taxonomy. Grounding covers all four concept types; after consolidation, the linked shares across the domain, algorithm, design-pattern, and paradigm axes are reported, and grounded concepts collapse further along Wikidata's identifier space, indicating that a sizeable share of the consolidated long tail consists of lexical variants sharing an encyclopaedic referent; residual ungrounded concepts concentrate on the algorithm axis.
It provides a calibrated quality-assurance protocol that turns an LLM verifier into a per-annotation reliability filter and preserves both extractor and verifier decisions so downstream users can choose an operating point. Compared with narrower settings that use fixed taxonomies or lack a per-annotation reliability layer, this work uses a small human gold set to characterise task difficulty, compares six code-oriented models under a uniform protocol to select a production verifier, and then re-evaluates the whole corpus at silver tier. The gold set comprises 127 stratified source files, 17 annotators, and 292 judgements across 6 languages and 4 semantic categories; overall 68.4% of judgements were Okay, 17.7% NotOkay, and 13.9% Unsure, with modest overall Fleiss' kappa and marked variation by category (higher for domain and algorithm, much weaker for design pattern and paradigm); across 1,536 model-human comparison pairs the six candidate verifiers reach F1 between 84.9% and 88.5%, and Devstral-2 2512 is retained with 80.6% accuracy, 82.1% precision, 96.0% recall, and 88.5% F1; on a 10,000-file silver-tier subset acceptance is 93% for paradigm, 88% for algorithm and domain, and 82% for design pattern, with cross-language acceptance in an 87-90% band.
It reports queryable graph statistics and two use cases on a real corpus, illustrating algorithmic primitives recurring across domains and the possibilities for semantic code search and automated algorithm selection. Compared with the closest existing code knowledge graph, GraphGen4Code (1.3 million Python files limited to syntactic artifacts), CodeGraph is about two orders of magnitude larger and adds rich semantic metadata; its queries target what a file is about rather than intra-program calls and data flow. The graph has about 158 million nodes and about 1 billion typed edges with a low average degree, consistent with a large but sparse typed network; the corpus is skewed, with Java 30.6%, Python 16.9%, and C++ 11.2% together accounting for 58.7%; the domain distribution is also skewed, with web development leading at 62.1 million files, more than twice database (27.1 million), followed by data science (24.9 million) and game (16.4 million); the algorithm head consists of computational primitives rather than textbook algorithms (linear search 6.0 million, string operation 4.0 million, input validation 2.9 million, JSON 2.9 million, regular expression 2.3 million, string-searching algorithm 2.2 million); roughly 75.1% of domains and 38.8% of algorithms associate with fewer than 1,000 files.
Perspective
The result targets developers and researchers who need to retrieve or filter code by the algorithms, paradigms, or design patterns it embodies, and it applies to quality-filtered public code corpora such as Stack-Edu with the file as the annotation unit; the dual-layer structure lets fine-grained concepts be rolled up along the Wikidata hierarchy to broader semantic families and lets concepts be compared across languages under a common identifier space. The authors treat domain and algorithm as axes suitable for quantitative claims and release paradigm and design pattern as exploratory signals, recommending that downstream quantitative analyses prioritise the former two. Methodologically, deterministic SPARQL handles the unambiguous bulk while the Deep Research Agent is invoked only on residuals, keeping grounding deterministic where possible; both extractor and verifier decisions are preserved in the graph so users can pick an operating point matching their precision/recall preference.
The human gold set covers only 127 files and 292 judgements, and the authors note that annotating the full corpus at a conservative two minutes per file-axis pair would require roughly 9,500 person-years, so this scale reflects the intrinsic difficulty of the task; overall Fleiss' kappa is modest and markedly weaker for design pattern and paradigm, which the authors therefore release as exploratory signals. The silver-tier acceptance rate (about 88% overall) is higher than the gold-set Okay rate (68.4%); the authors treat the two as bracketing the range of plausible precision estimates and attribute the gap to false-positive behaviour and to converting human Unsure judgements into Okay verifier decisions. On the subset restricted to Wikidata-grounded annotations, acceptance drops markedly for all axes except paradigm, and the authors offer as a working hypothesis that algorithm and design-pattern labels typically describe constructs at the level of a single function or class, creating a granularity mismatch with whole-file judgement; a more precise characterisation, possibly with function-level evaluation, is left to future work. Roughly 75.1% of domains and 38.8% of algorithms associate with fewer than 1,000 files, so robust statistical claims are limited to high-frequency labels. Residual ungrounded concepts concentrate on the algorithm axis, and the authors see the Deep Research Agent as architecturally positioned to absorb this tail as SPARQL recall and candidate generation improve. The full corpus release, an extended cross-domain analysis, and the implications for AI model training are all left to future work.
