Codex
+5.3 with gate- Bare
- 74.8
- + Rules
- 76.4
- + Gate
- 80.1
BENCH-V1 tested Codex and Claude Code on the same 20 UI fixtures under three conditions. Raw rules alone were inconsistent. The enforced render → score → revise loop improved both agents by the same 5.3 points.
The actual result
A markdown rule set helped Codex, but hurt Claude Code in this run. Quality became consistent only when the agent had to render, receive a category score, fix named failures, and submit the best verified iteration.
120 scored cells = 20 fixtures × 2 agents × 3 conditions. Every condition completed 20/20.
The 80 line is StyleSeed’s shipping threshold, not a claim that either agent or judge is objectively “best.”
Gate-condition breakdown
The gate produced its strongest results in brand fit, coherence, hierarchy, and color. Distinctiveness remained the lowest category, so BENCH-V1 is evidence of improvement—not a claim that the design problem is solved.
| Category | Codex | Claude |
|---|---|---|
| Color | 81.8 | 81.8 |
| Hierarchy | 82.8 | 82.0 |
| Typography | 80.8 | 80.8 |
| Spacing | 80.0 | 78.0 |
| Cards | 79.9 | 78.9 |
| Coherence | 82.0 | 83.9 |
| Distinctiveness | 70.9 | 69.0 |
| Accessibility | 78.6 | 76.4 |
| Brand fit | 84.9 | 84.8 |
Protocol
What we did not hide
15 app-builder cells remain blocked. v0, Bolt, and Lovable require first-party access that was not available. No scores were estimated or imputed.
The old-vs-new judge correlation is not a clean agreement test. The earlier screenshots were not retained, so the reported Spearman ρ 0.394 on n=39 also contains generation variance.
One benchmark is not universal truth. BENCH-V1 supports the narrower claim that an enforced revision loop was more reliable than distributing rules alone in this protocol.
StyleSeed fixes the design method, not one aesthetic: lock bounded decisions, choose the right output grammar, render, score, revise, and visually verify.
See the engine architecture