Codex
+5.3 with gate- Bare
- 74.8
- + Rules
- 76.4
- + Gate
- 80.1
BENCH-V1 tested Codex and Claude Code on the same 20 UI fixtures under three conditions. Raw rules alone were inconsistent. The enforced render → score → revise loop improved both agents by the same 5.3 points.
The actual result
A markdown rule set helped Codex, but hurt Claude Code in this run. Quality became consistent only when the agent had to render, receive a category score, fix named failures, and submit the best verified iteration.
120 scored cells = 20 fixtures × 2 agents × 3 conditions. Every condition completed 20/20.
The 80 line is StyleSeed’s shipping threshold, not a claim that either agent or judge is objectively “best.”
Gate-condition breakdown
The gate produced its strongest results in brand fit, coherence, hierarchy, and color. Distinctiveness remained the lowest category, so BENCH-V1 is evidence of improvement—not a claim that the design problem is solved.
| Category | Codex | Claude |
|---|---|---|
| Color | 81.8 | 81.8 |
| Hierarchy | 82.8 | 82.0 |
| Typography | 80.8 | 80.8 |
| Spacing | 80.0 | 78.0 |
| Cards | 79.9 | 78.9 |
| Coherence | 82.0 | 83.9 |
| Distinctiveness | 70.9 | 69.0 |
| Accessibility | 78.6 | 76.4 |
| Brand fit | 84.9 | 84.8 |
Protocol
What we did not hide
15 app-builder cells remain blocked. v0, Bolt, and Lovable require first-party access that was not available. No scores were estimated or imputed.
The old-vs-new judge correlation is not a clean agreement test. The earlier screenshots were not retained, so the reported Spearman ρ 0.394 on n=39 also contains generation variance.
One benchmark is not universal truth. BENCH-V1 supports the narrower claim that an enforced revision loop was more reliable than distributing rules alone in this protocol.
StyleSeed aims to make expert decisions repeatable by coding agents. This benchmark tests a render-score-revise loop; it does not prove expert-level quality or replace human design acceptance.
See the engine architecture