StyleSeed
BENCH-V1 · complete · Jul 25, 2026

The gate moved both agents +5.3.

BENCH-V1 tested Codex and Claude Code on the same 20 UI fixtures under three conditions. Raw rules alone were inconsistent. The enforced render → score → revise loop improved both agents by the same 5.3 points.

cells scored
120
failed
0
judge spread
2 pt

The actual result

Rules are guidance. The gate is the mechanism.

A markdown rule set helped Codex, but hurt Claude Code in this run. Quality became consistent only when the agent had to render, receive a category score, fix named failures, and submit the best verified iteration.

Codex

+5.3 with gate
Bare
74.8
+ Rules
76.4
+1.6
+ Gate
80.1
+5.3

Claude Code

+5.3 with gate
Bare
74.1
+ Rules
70.4
-3.7
+ Gate
79.4
+5.3

120 scored cells = 20 fixtures × 2 agents × 3 conditions. Every condition completed 20/20.

The 80 line is StyleSeed’s shipping threshold, not a claim that either agent or judge is objectively “best.”

Gate-condition breakdown

Coherence held. Distinctiveness is still the frontier.

The gate produced its strongest results in brand fit, coherence, hierarchy, and color. Distinctiveness remained the lowest category, so BENCH-V1 is evidence of improvement—not a claim that the design problem is solved.

CategoryCodexClaude
Color81.881.8
Hierarchy82.882.0
Typography80.880.8
Spacing80.078.0
Cards79.978.9
Coherence82.083.9
Distinctiveness70.969.0
Accessibility78.676.4
Brand fit84.984.8

Protocol

A benchmark with receipts.

  • Same work. 20 neutral prompts across dashboard, marketing, app, commerce, and data-heavy domains.
  • Three conditions. Bare agent, StyleSeed rules, then rules plus the enforced score-and-revise gate.
  • Neutral judge. A fixed-seed Gemini 3.1 Pro judge scored nine published categories. It passed five repeats at 72, 72, 72, 74, 74—a 2-point spread.
  • Preserved evidence. Prompts, outputs, rendered screenshots, verdicts, fix prompts, and the selected best iteration are retained in the release archive.

What we did not hide

15 app-builder cells remain blocked. v0, Bolt, and Lovable require first-party access that was not available. No scores were estimated or imputed.

The old-vs-new judge correlation is not a clean agreement test. The earlier screenshots were not retained, so the reported Spearman ρ 0.394 on n=39 also contains generation variance.

One benchmark is not universal truth. BENCH-V1 supports the narrower claim that an enforced revision loop was more reliable than distributing rules alone in this protocol.

StyleSeed fixes the design method, not one aesthetic: lock bounded decisions, choose the right output grammar, render, score, revise, and visually verify.

See the engine architecture