Evaluation Report — accessibility-skills

Accessibility Skill
Benchmark Dashboard

Results from 83 graded fixtures testing 11 models across 4 platforms — proving that structured accessibility review protocols work reliably at scale, with and without a cloud API.

83
Graded Fixtures
11
Models Tested
4
Platforms
$0.65
Total Cloud Cost (33 fixtures)

Chart 01

How Every Model Performed

Pass rate on the a11y-critic eval suite (33 graded fixtures). Bars differentiated by platform via border style: solid = Claude API, dashed = Ollama local, dotted = OpenAI, double = Gemini.

Claude API — labeled "(API)" on y-axis
Ollama / local — labeled "(Ollama)" on y-axis
OpenAI — labeled "(OpenAI)" on y-axis
Gemini — labeled "(Gemini)" on y-axis
Key finding: All tested models achieve ≥85% pass rate. The failures that exist are in ADVERSARIAL and CLEAN fixture verdict calibration — not in bug detection. Every model tested found real accessibility bugs when they were present.
View data table
Model Platform Fixtures Tested Pass Rate Must-find Rate
qwen3:32b Ollama (local) 33 100% 96%
Claude Sonnet 4.6 + thinking Claude API 33 (via escalation) 100% 100%
llama3.3:70b Ollama (local) 7 100% 86%
qwen3.5:latest Ollama (local) 7 100% 86%
GPT-5.2 OpenAI 33 91% 100%
GPT-5.2 → GPT-5.5-low escalation OpenAI 33 (via escalation) 100% 100%
qwen3.5:27b Ollama (local) 17 (partial run) 94% 100%
Claude Haiku 4.5 Claude API 33 85% 100%
Gemini 2.5 Flash Gemini 33 94% 98%

Chart 02

Bottom-Up: Finding the Cheapest Model That Works

Start with the cheapest tier. Only escalate failures. Total cost for all 33 fixtures: $0.65.

Resolved at this tier
Escalated to next tier
Cost at tier: Haiku $0.45 → Sonnet $0.18 → Sonnet+think $0.02
Strategy: Haiku is production-viable for bug detection (85% pass, 100% must-find on HAS-BUGS). Failures concentrate in verdict calibration on ADVERSARIAL and CLEAN fixtures. Escalating only those 5 failures brings total cost from $0.45 to $0.65 — a 44% premium for a 15-point pass-rate gain.
View data table
Tier Fixtures Run Passed Failed / Escalated Cumulative Cost
Haiku 4.5 33 28 5 $0.45
Sonnet 4.6 5 4 1 $0.63
Sonnet 4.6 + thinking 1 1 0 $0.65

Chart 03

83 Graded Fixtures Across 3 Suites

Every fixture has a component, metadata, and scoring rubric. Difficulty tiers shown as stacked segments.

HAS-BUGS / Domain-specific
CLEAN / ARIA patterns
FLAWED / Keyboard nav
ADVERSARIAL
View data table
Suite Total Tier 1 Tier 2 Tier 3 Tier 4
a11y-critic 33 21 HAS-BUGS 4 CLEAN 5 FLAWED 3 ADVERSARIAL
a11y-planner 25 15 Domain-specific 5 ARIA patterns 5 Keyboard nav —
perspective-audit 25 16 HAS-BUGS 5 CLEAN — 4 ADVERSARIAL

Chart 04

Detection Is Solved. Judgment Isn't.

Pass rate by fixture difficulty tier — Claude Haiku 4.5, 33 fixtures. Every real bug was detected. Failures concentrate in adversarial verdict calibration.

HAS-BUGS (100%): All 21 fixtures with planted bugs passed. Every must-find item detected.
FLAWED (100%): Subtle, partially-correct implementations — all 5 detected.
CLEAN (50%): 2 of 4 clean-code fixtures passed. 1 false positive on button-skip-link-clean, 1 WARN on modal-complete-clean.
ADVERSARIAL (0%): Haiku found the tradeoffs in all 3 fixtures (1/1 must-articulate) but could not calibrate the verdict for genuinely ambiguous design patterns.
View data table
Difficulty Tier Fixtures Passed Failed Pass Rate Pattern
HAS-BUGS 21 21 0 100% All bugs detected
FLAWED 5 5 0 100% Subtle bugs detected
CLEAN 4 2 2 50% 1 false positive, 1 WARN
ADVERSARIAL 3 0 3 0% Right analysis, wrong verdict

Chart 05

What Does Accessibility Review Actually Cost?

Cost per 33-fixture run vs. pass rate. Bubble size reflects speed — smaller = faster. X-axis uses logarithmic scale.

The escalation sweet spot: At $0.65 for 33 fixtures, the Haiku→Sonnet escalation strategy achieves 100% pass rate for 74% less cost than running Sonnet across all fixtures (~$2.50) and 96% less than running Opus (~$15). The local Ollama option costs nothing per-run once hardware is available, but requires a 20 GB model download.
View data table
Platform / Strategy Cost (33 fixtures) Pass Rate Avg Speed (sec/fixture)
Ollama qwen3:32b (local) $0.00 100% ~240s
Claude Haiku 4.5 $0.45 85% ~67s
Claude Escalation (Haiku+Sonnet+think) $0.65 100% ~91s blended
Claude Sonnet 4.6 (all 33) ~$2.50 ~97% ~129s
Claude Opus 4.6 (all 33) ~$15.00 ~100% ~200s

Planner Suite

Planner Benchmark Results

25-fixture planner suite, section-presence scoring (post-002 instrument). Codex planner lane pending (plan 010); Gemini planner lane out of scope.

Lane PASS Must-have criteria Date
qwen3:32b (local) 25/25 227/235 (96.6%) 2026-06-11
Claude Opus subagents 25/25 234/235 (99.6%) 2026-06-12