StyleSeed
BENCH-V1 · complete · Jul 25, 2026

The gate moved both agents +5.3.

BENCH-V1 tested Codex and Claude Code on the same 20 UI fixtures under three conditions. Raw rules alone were inconsistent. The enforced render → score → revise loop improved both agents by the same 5.3 points.

cells scored
120
failed
0
judge spread
2 pt

The actual result

Rules are guidance. The gate is the mechanism.

A markdown rule set helped Codex, but hurt Claude Code in this run. Quality became consistent only when the agent had to render, receive a category score, fix named failures, and submit the best verified iteration.

Codex

+5.3 with gate
Bare
74.8
+ Rules
76.4
+1.6
+ Gate
80.1
+5.3

Claude Code

+5.3 with gate
Bare
74.1
+ Rules
70.4
-3.7
+ Gate
79.4
+5.3

120 scored cells = 20 fixtures × 2 agents × 3 conditions. Every condition completed 20/20.

The 80 line is StyleSeed’s shipping threshold, not a claim that either agent or judge is objectively “best.”

Gate-condition breakdown

Coherence held. Distinctiveness is still the frontier.

The gate produced its strongest results in brand fit, coherence, hierarchy, and color. Distinctiveness remained the lowest category, so BENCH-V1 is evidence of improvement—not a claim that the design problem is solved.

CategoryCodexClaude
Color81.881.8
Hierarchy82.882.0
Typography80.880.8
Spacing80.078.0
Cards79.978.9
Coherence82.083.9
Distinctiveness70.969.0
Accessibility78.676.4
Brand fit84.984.8

Protocol

A benchmark with receipts.

  • Same work. 20 neutral prompts across dashboard, marketing, app, commerce, and data-heavy domains.
  • Three conditions. Bare agent, StyleSeed rules, then rules plus the enforced score-and-revise gate.
  • Neutral judge. A fixed-seed Gemini 3.1 Pro judge scored nine published categories. It passed five repeats at 72, 72, 72, 74, 74—a 2-point spread.
  • Preserved evidence. Prompts, outputs, rendered screenshots, verdicts, fix prompts, and the selected best iteration are retained in the release archive.

What we did not hide

15 app-builder cells remain blocked. v0, Bolt, and Lovable require first-party access that was not available. No scores were estimated or imputed.

The old-vs-new judge correlation is not a clean agreement test. The earlier screenshots were not retained, so the reported Spearman ρ 0.394 on n=39 also contains generation variance.

One benchmark is not universal truth. BENCH-V1 supports the narrower claim that an enforced revision loop was more reliable than distributing rules alone in this protocol.

StyleSeed aims to make expert decisions repeatable by coding agents. This benchmark tests a render-score-revise loop; it does not prove expert-level quality or replace human design acceptance.

See the engine architecture