Evolved LLM features lift credit-risk classification from chance to 69.0% and expose zero-shot single-class bias
Related research and updatesSynopsis
The authors propose an evolutionary framework in which an LLM iteratively generates and refines natural-language binary features (rubrics), an LLM scores each sample against them, and the resulting vectors train a transparent classifier such as logistic regression; across AG News, GoEmotions, and HELOC, evolved features beat single-shot rubrics by +2.9 pp on average and outperform zero-shot LLM classification on two of three tasks, with the zero-shot LLM reaching only 50.7% on HELOC (near chance with a strong single-class bias) while evolved features reach 69.0% with balanced, auditable predictions.
Figure 3 : Cross-validation accuracy over evolutionary iterations. Solid lines show best-accepted CV score; shaded regions show the range across 3 seeds. Evolution makes decisive improvements early in the search on tasks where the initial rubric is suboptimal.
arXivInterpretation
An evolutionary feature-discovery loop in which the LLM acts as both mutation operator and fitness evaluator: each iteration extracts binary features under the current rubric, runs cross-validation, then feeds classification errors, per-class activation rates, and feature ablation scores back to the generator, which retains essential features, replaces low-value ones, and proposes new features targeting confused class pairs. Prior LLM feature-generation work (e.g., Singh et al., Balek et al.) produces static feature sets fixed at the start and never adapted to training-time error patterns; here the feature set itself is the object of iterative optimization, and the optimization targets the features of a transparent classifier rather than the prompt of an LLM classifier. The method is compared against TF-IDF, single-shot rubrics, and zero-shot LLM classification on three benchmarks under identical splits, with mean and standard deviation over 3 independent stratified splits; ablation scores and per-class activation rates are explicit feedback signals written into the evolution prompt.
Evolved features achieve the best average accuracy across the three benchmarks (63.8%), improving over single-shot rubrics by +2.9 pp on average, with the largest gains on GoEmotions (+6.3 pp) and HELOC (+3.7 pp). Single-shot rubrics average 60.9% and zero-shot LLM classification averages 59.7%; evolved features exceed the zero-shot LLM on two of three tasks while remaining fully interpretable. Results are reported as mean ± standard deviation, and evolved features show substantially lower variance on GoEmotions (1.6 vs 6.2), suggesting the search yields more robust feature sets even when initial rubric quality varies across splits.
On the HELOC credit-risk task the zero-shot LLM reaches only 50.7% accuracy, near chance, assigning 77–85% of predictions to the at_risk class across seeds, which yields at_risk recall of 0.80–0.84 but not_at_risk recall of only 0.14–0.26; evolved features reach 69.0% with per-class F1 of 0.64–0.71 for both classes. This failure is invisible in aggregate accuracy and surfaces only under per-class auditing; the interpretable feature layer makes the systematic single-class bias an inspectable object, whereas the black-box baseline offers no comparable audit entry point. The comparison uses identical data splits and multiple random seeds and reports per-class recall and F1 ranges; the evolved HELOC rubric consists of 10 natural-language threshold rules with logistic-regression coefficients.
Evolution helps most when label boundaries cannot be inferred from category names (the 7-class Ekman emotion mapping in GoEmotions) or when the LLM lacks domain reasoning (structured numerical credit data in HELOC); when categories are self-explanatory and well separated (AG News), the zero-shot LLM leads at 89.7% and evolution does not improve on the initial rubric (81.3% vs 82.7%). This gives an empirical characterization of when the method applies rather than a blanket claim that evolution always helps; it also notes that convergence mostly occurs within the first 5 accepted rubrics, so 10–15 iterations are recommended. Convergence trajectories show GoEmotions gaining +10 to +12 pp from the first accepted rubric before plateauing, and HELOC cross-validation rising from 0.65–0.69 to 0.69–0.74 before flattening; on cost, training needs roughly 5,000 LLM calls but inference needs exactly one call per sample, the same as the zero-shot feature method.
Perspective
The framework applies to classification tasks that need auditable decision logic, especially where label boundaries cannot be inferred from category names or where zero-shot LLM domain reasoning is unreliable; inputs may be natural-language text or structured data serialized to text. For practitioners who want per-item, checkable decision rationales (for example credit reviewers or model evaluators in regulated industries), it offers an end-to-end inspection path from natural-language feature definitions through binary feature values to logistic-regression coefficients. The authors recommend 10–15 iterations as sufficient for most tasks, and inference requires exactly one LLM call per sample, the same operational cost as the zero-shot feature method.
Feature values come from an LLM judge, so extraction errors propagate directly into the feature matrix used to train the classifier, meaning extraction accuracy needs validating separately from end-task accuracy; only a single LLM is used, leaving cross-model generalization open; the binary representation bottlenecks fine-grained tasks; and broader evaluation on financial and medical data, plus feature-pool architectures that decouple generation from L1-regularized selection, are directions the authors propose. In addition, the HELOC explanations are structured to support adverse-action notices required by the Equal Credit Opportunity Act, but establishing their legal validity would require fairness, proxy-discrimination, and domain-expert auditing beyond the paper's scope.
