Skip to main content
Back to timeline
arXivSource publication:

Design Creativity Bench: 15 frontier models' UI designs fall 0.12 below humans in originality and 0.27 in cross-brief variation, despite exceeding 90% appropriateness

Synopsis

The work introduces Design Creativity Bench, which uses 80 tasks with two open-ended UI briefs per task sharing a goal but differing in product domain, has 15 frontier models each generate 160 HTML designs, and measures originality via embedding similarity (0.592 across models vs 0.764 for human-model pairs) and creative range (0.581 across models vs 0.902 for humans), alongside appropriateness against per-brief acceptance criteria (above 90% for every model, up to 99.2%, vs 98.0% for humans).

Source-provided article image: Design Creativity Bench: Measuring creativity in LLM-Generated UI
Figure 1 ·

Figure 1: Diversity in LLM UI generation compared with human designs. The examples are drawn from the benchmark 𝒟 B \mathcal{D}_{B} . The model-generated designs are markedly more similar across domain and content variants than the corresponding human reference designs. Each row pairs two prompts that share the goal shown (paraphrased) and differ in product domain and content; similarity is the pair score, from 0 to 1, where higher means more alike. We ask models to generate desktop designs for a 1280 × 720 1280\times 720 target viewport.

arXiv

Interpretation

The benchmark decomposes UI design creativity into three separately measurable quantities: originality (mean dissimilarity between designs from different models on the same prompt), creative range (dissimilarity between one model's designs across two briefs that share a goal but differ in product domain), and appropriateness (percentage of a brief's acceptance criteria a design meets), plus a combined overall score. Prior evidence on LLM output homogeneity came mainly from text and code generation; this work extends the measurement to UI design and covers both cross-model repetition and repetition across related briefs. 15 models each generated 160 designs (2,400 primary model designs) across 80 tasks and 80 differentiated brief pairs per model; originality rests on 15,840 unique model-model pairs and the human reference on 2,400 human-model pairs; intervals use 20,000 bootstrap resamples of whole tasks (seed 7).

Model-generated designs resemble one another more than they resemble the human reference: originality is 0.592 (95% CI [0.582, 0.602]) versus 0.764 (95% CI [0.751, 0.778]) for human-model pairs; the most original model, GPT-6 Astra, scores 0.646 and the least original, Kimi K3, 0.547, and no model interval reaches the human value. The result carries the output-homogeneity finding from text and code tasks into UI design and quantifies a within-model-population similarity baseline. Same-prompt cross-model pairings with same-lab models excluded; the human reference was produced by multiple human designers working from the same briefs with access to public inspiration libraries such as Behance and Kombai Gallery.

Changing the prompt does not change the design: within one model, designs for two briefs sharing a goal but differing in product domain differ by only 0.581 on average (95% CI [0.567, 0.597]), against 0.902 (95% CI [0.884, 0.919]) for human pairs; when the same model designs two unrelated tasks the mean distinctiveness is 0.811, indicating the brief change should be enough to motivate variation. The work treats cross-brief variation as a second diversity dimension independent of cross-model difference, and shows the two rankings are essentially unrelated (Spearman correlation; exploratory 95% model-bootstrap interval, 20,000 resamples, seed 7). 80 two-prompt pairs per model; model range spans GLM-5.3 at 0.634 to GPT-6 Astra at 0.472, yet no model interval reaches the human value; the least distinct archetype is settings (0.554) and the most distinct is listing (0.590).

Appropriateness is generally high: every model meets more than 90% of its briefs' criteria and more than 96% of critical criteria; Claude Opus 5.5 leads at 99.2% and GPT-6 Astra at 99.1%, both slightly above the 98.0% human reference. The most common appropriateness failures are missing controls for a workflow the brief names (33%) and sample data whose counts and totals do not match (26%). High appropriateness and high diversity are not in tension: the human baseline reaches 0.902 creative range at 98.0% appropriateness, so the diversity gap is not forced by meeting requirements. Each brief carries yes/no acceptance criteria, with critical criteria tagged as failures that mean the design does not do the job; the overall score weights appropriateness and diversity equally, with diversity combining originality at 0.9 and creative range at 0.1 (Appendix B8 gives the weight derivation and sensitivity: rank correlation at least 0.99 when the creative-range weight is 0 or 0.15).

Perspective

The benchmark targets UI design and is meant for comparing model outputs against a human reference under the same generation settings; it measures variation among acceptable designs rather than design quality or user preference. It can be used to evaluate methods that raise diversity while preserving brief satisfaction, and to place a later model on the frozen 15-model panel scale (Appendix B9 gives the placement rules for 18 further models). The human reference was created by multiple designers adapting public inspiration references, supporting model-human comparisons and human cross-brief comparisons.

Decoding parameters are uncontrolled: hosted APIs are called with default temperature and sampling, so results characterize default outputs rather than sampling strategies. Repetition within one context is not measured, since each design is generated in a fresh context. No diversity-seeking prompting was attempted. The human reference does not measure same-designer repetition or variation between different human designers responding to the same prompt. The 0.1 creative-range weight in the overall score comes from a proxy estimate of model usage shares (OpenRouter token rankings for its twenty most-used models, retrieved 30 September 2026), which overstates true shares. Appendix A1's enumeration of options for a single prompt weights decisions by how conventional an LLM judges them, and the authors state these weights are assumptions rather than estimates from real designers. In addition, several formulas and some numeric values appear as placeholders in the provided text, so specific calibration constants should be checked in the original appendices before reproduction.

Sources