Chart 01
How Every Model Performed
Pass rate on the a11y-critic eval suite (33 graded fixtures). Bars differentiated by platform via border style: solid = Claude API, dashed = Ollama local, dotted = OpenAI, double = Gemini.
View data table
| Model | Platform | Fixtures Tested | Pass Rate | Must-find Rate |
|---|---|---|---|---|
| qwen3:32b | Ollama (local) | 33 | 100% | 96% |
| Claude Sonnet 4.6 + thinking | Claude API | 33 (via escalation) | 100% | 100% |
| llama3.3:70b | Ollama (local) | 7 | 100% | 86% |
| qwen3.5:latest | Ollama (local) | 7 | 100% | 86% |
| GPT-5.2 | OpenAI | 33 | 91% | 100% |
| GPT-5.2 → GPT-5.5-low escalation | OpenAI | 33 (via escalation) | 100% | 100% |
| qwen3.5:27b | Ollama (local) | 17 (partial run) | 94% | 100% |
| Claude Haiku 4.5 | Claude API | 33 | 85% | 100% |
| Gemini 2.5 Flash | Gemini | 33 | 94% | 98% |
Chart 02
Bottom-Up: Finding the Cheapest Model That Works
Start with the cheapest tier. Only escalate failures. Total cost for all 33 fixtures: $0.65.
View data table
| Tier | Fixtures Run | Passed | Failed / Escalated | Cumulative Cost |
|---|---|---|---|---|
| Haiku 4.5 | 33 | 28 | 5 | $0.45 |
| Sonnet 4.6 | 5 | 4 | 1 | $0.63 |
| Sonnet 4.6 + thinking | 1 | 1 | 0 | $0.65 |
Chart 03
83 Graded Fixtures Across 3 Suites
Every fixture has a component, metadata, and scoring rubric. Difficulty tiers shown as stacked segments.
View data table
| Suite | Total | Tier 1 | Tier 2 | Tier 3 | Tier 4 |
|---|---|---|---|---|---|
| a11y-critic | 33 | 21 HAS-BUGS | 4 CLEAN | 5 FLAWED | 3 ADVERSARIAL |
| a11y-planner | 25 | 15 Domain-specific | 5 ARIA patterns | 5 Keyboard nav | — |
| perspective-audit | 25 | 16 HAS-BUGS | 5 CLEAN | — | 4 ADVERSARIAL |
Chart 04
Detection Is Solved. Judgment Isn't.
Pass rate by fixture difficulty tier — Claude Haiku 4.5, 33 fixtures. Every real bug was detected. Failures concentrate in adversarial verdict calibration.
FLAWED (100%): Subtle, partially-correct implementations — all 5 detected.
CLEAN (50%): 2 of 4 clean-code fixtures passed. 1 false positive on
button-skip-link-clean, 1 WARN on modal-complete-clean.ADVERSARIAL (0%): Haiku found the tradeoffs in all 3 fixtures (1/1 must-articulate) but could not calibrate the verdict for genuinely ambiguous design patterns.
View data table
| Difficulty Tier | Fixtures | Passed | Failed | Pass Rate | Pattern |
|---|---|---|---|---|---|
| HAS-BUGS | 21 | 21 | 0 | 100% | All bugs detected |
| FLAWED | 5 | 5 | 0 | 100% | Subtle bugs detected |
| CLEAN | 4 | 2 | 2 | 50% | 1 false positive, 1 WARN |
| ADVERSARIAL | 3 | 0 | 3 | 0% | Right analysis, wrong verdict |
Chart 05
What Does Accessibility Review Actually Cost?
Cost per 33-fixture run vs. pass rate. Bubble size reflects speed — smaller = faster. X-axis uses logarithmic scale.
View data table
| Platform / Strategy | Cost (33 fixtures) | Pass Rate | Avg Speed (sec/fixture) |
|---|---|---|---|
| Ollama qwen3:32b (local) | $0.00 | 100% | ~240s |
| Claude Haiku 4.5 | $0.45 | 85% | ~67s |
| Claude Escalation (Haiku+Sonnet+think) | $0.65 | 100% | ~91s blended |
| Claude Sonnet 4.6 (all 33) | ~$2.50 | ~97% | ~129s |
| Claude Opus 4.6 (all 33) | ~$15.00 | ~100% | ~200s |
Planner Suite
Planner Benchmark Results
25-fixture planner suite, section-presence scoring (post-002 instrument). Codex planner lane pending (plan 010); Gemini planner lane out of scope.
| Lane | PASS | Must-have criteria | Date |
|---|---|---|---|
| qwen3:32b (local) | 25/25 | 227/235 (96.6%) | 2026-06-11 |
| Claude Opus subagents | 25/25 | 234/235 (99.6%) | 2026-06-12 |