Skip to main content
Back to timeline
arXivSource publication:

Project Greenhouse pre-trained a 3B model from scratch on 8 H100s and fine-tuned it into a pointwise reranker that beats comparable open-weight baselines on TREC DL and BEIR, trailing only GPT-6.1 Sol

Synopsis

Project Greenhouse models LLM training as a directed property hypergraph and uses it to define what makes a model fully open and sovereign; as a first milestone, the authors pre-train a 3.29B-parameter causal language model from scratch on ClimbMix using only publicly available datasets, then fine-tune it on RLHN-250K with localized contrastive estimation to obtain the pointwise reranker Gaggle (base reranker). Its mean nDCG@10 exceeds all compared fine-tuned pointwise and listwise rerankers on TREC DL 2019–2023 and seven BEIR collections, with only GPT-6.1 Sol scoring higher on TREC DL.

Source-provided article image: Project Greenhouse: Progress Toward Fully Open and Sovereign Agentic Search
Figure 1 ·

Figure 1 : Validation loss on held-out ClimbMix during pre-training.

arXiv

Interpretation

The authors give a reusable criterion for "fully open and sovereign": LLM training is represented as a directed property hypergraph whose vertices are artifacts such as corpora and weights and whose hyperedges are training recipes and configurations; a model is fully open and sovereign when all upstream vertices and edges are public and free of proprietary dependencies. Prior discussion largely stopped at "open weights"; this formalization separates openness from reproducibility and lets a data ablation be expressed as a hyperedge with fewer tail vertices and a code ablation as a different hyperedge over the same tails. This is a conceptual contribution; the authors take commonly accepted datasets such as ClimbMix and MS MARCO as provenance boundaries and explicitly acknowledge that drawing those boundaries involves judgment.

Public data alone suffices to pre-train a backbone good enough for a competitive reranker: a 3.29B-parameter causal model trained for one pass over ClimbMix (287B tokens) on 8 H100s in roughly 1,600 GPU-hours reaches CORE 0.458, MMLU 0.493, and MMLU-Pro 0.178. Unlike the dominant approach, this backbone uses no third-party open-weight initialization and requires no instruction tuning, alignment, mid- or post-training, teacher supervision, or synthetic data. The evaluation code was validated against seven reference models with published CORE scores, agreeing within 0.011; MMLU and MMLU-Pro agree with the stock harness within 0.003. The authors also note that the models in their comparison table differ in data, architecture, and recipe, so the table positions their checkpoint rather than isolating a single factor.

The two-step recipe yields Gaggle (base reranker) with mean nDCG@10 of 0.641 on TREC DL and 0.547 on BEIR, above every compared fine-tuned pointwise and listwise reranker; on TREC DL only GPT-6.1 Sol (0.653) is higher, and on BEIR it is comparable to Gemma-4 (0.548). The result indicates that for pointwise reranking a from-scratch backbone is "good enough", so starting from an open-weight backbone is not necessary; at 3B parameters the model outperforms several 7B–27B comparison conditions. Evaluation uses top-100 BM25 candidates and nDCG@10 throughout; four same-recipe fine-tuning trials show a sample standard deviation of about 0.001 for benchmark means and about 0.003 averaged over individual collections, which the authors treat as the noise scale. The soup checkpoint averages weights from those four trials.

Controlled comparisons yield practical training lessons: the larger RLHN-680K data raises the TREC DL mean from 0.641 to 0.646; a query–passage–query prompt raises TREC DL from 0.641 to 0.647 and BEIR from 0.544 to 0.551; and LCE beats pointwise cross-entropy (0.641 vs. 0.628 on TREC DL, 0.544 vs. 0.528 on BEIR). These comparisons span backbones, data, attention pattern, prompt, pooling, self-filtering, and loss under a shared recipe, and they flag that a shared learning rate does not transfer across backbones with different pre-trained weight scales. Differences are read against the run-to-run noise scale measured in the variability study; the authors also state that the RLHN-250K versus RLHN-680K comparison changes dataset size, source distribution, and optimization steps together, so it cannot isolate a single cause.

Perspective

The result is aimed at research groups and organizations that need to control the whole training pipeline under limited compute: the reranker can be plugged in as a standalone component after BM25, dense retrieval, or an external search API, used inside agentic search to prioritize what enters the model context, or used as a relevance judge for rewards and verification. The authors scope the work to information seeking and explicitly exclude general reasoning and coding, and they expect initial systems to be hybrids of proprietary and open components. Provenance boundaries are drawn at commonly accepted datasets such as ClimbMix and RLHN-250K, which the authors acknowledge involves judgment.

The authors list several open questions: whether reranking is simply too easy for modern LLMs, which would explain how a simple two-step recipe reaches competitive effectiveness; whether the full suite of agentic capabilities can be achieved in a fully open and sovereign manner under limited compute, since a competitive reranker does not establish a path to a competitive agent; whether datasets themselves contain undisclosed filtering, synthetic generation, and annotation decisions, which argues for explicitly declaring provenance boundaries; and out-of-domain generalization, which has not been thoroughly tested. In addition, this is a fast parse of the text, so the equations, Figures 1–5, and some table values are not fully rendered here, and details of the exact objective forms and training curves should be checked against the original.

Sources