Public articles linked to the same research event.
arXiv The authors propose an evolutionary framework in which an LLM iteratively generates and refines natural-language binary features (rubrics), an LLM scores each sample against them, and the resulting vectors train a transparent classifier such as logistic regression; across AG News, GoEmotions, and HELOC, evolved features beat single-shot rubrics by +2.9 pp on average and outperform zero-shot LLM classification on two of three tasks, with the zero-shot LLM reaching only 50.7% on HELOC (near chance with a strong single-class bias) while evolved features reach 69.0% with balanced, auditable predictions.
The authors propose an evolutionary framework in which an LLM iteratively generates and refines natural-language binary features (rubrics), an LLM scores each sample against them, and the resulting vectors train a transparent classifier such as logistic regression; across AG News, GoEmotions, and HELOC, evolved features beat single-shot rubrics by +2.9 pp on average and outperform zero-shot LLM classification on two of three tasks, with the zero-shot LLM reaching only 50.7% on HELOC (near chance with a strong single-class bias) while evolved features reach 69.0% with balanced, auditable predictions.
The authors propose an evolutionary framework in which an LLM iteratively generates and refines natural-language binary features (rubrics), an LLM scores each sample against them, and the resulting vectors train a transparent classifier such as logistic regression; across AG News, GoEmotions, and HELOC, evolved features beat single-shot rubrics by +2.9 pp on average and outperform zero-shot LLM classification on two of three tasks, with the zero-shot LLM reaching only 50.7% on HELOC (near chance with a strong single-class bias) while evolved features reach 69.0% with balanced, auditable predictions.
The authors propose an evolutionary framework in which an LLM iteratively generates and refines natural-language binary features (rubrics), an LLM scores each sample against them, and the resulting vectors train a transparent classifier such as logistic regression; across AG News, GoEmotions, and HELOC, evolved features beat single-shot rubrics by +2.9 pp on average and outperform zero-shot LLM classification on two of three tasks, with the zero-shot LLM reaching only 50.7% on HELOC (near chance with a strong single-class bias) while evolved features reach 69.0% with balanced, auditable predictions.