2026-05-19 · updated 2026-07-13 — accessibility-skills

Accessibility Skill Bundle

Planner, critic, tester, and perspective auditor — with 83 graded fixtures, 11 benchmarked models, and a team of 6 agents.

4
Skills
83
Fixtures
11
Models Tested
6
Agents
1 — Skills
Critic

a11y-critic

11-phase investigation protocol. Reviews ARIA patterns, focus management, state communication. Read-only.

Planner

a11y-planner

9-phase planning protocol. Maps to WAI-ARIA APG patterns, designs keyboard models, focus management.

Tester

a11y-test

Five modes: Playwright keyboard tests, axe-core scans, agent-browser recon, keyboard-a11y-tester journey audits, virtual-screen-reader component SR assertions. Real browser events, not synthetic.

Auditor

perspective-audit

7 access perspectives: screen reader, keyboard, low vision, cognitive, vestibular, auditory, environmental.

2 — Lifecycle
Plan
→
Critique Plan
→
[Perspective]
→
Revise
→
Implement
→
Test
→
Critique Impl
→
[Perspective]
→
Fix
→
Re-test

The critic serves at two checkpoints. Perspective-audit is escalated only when the planner or critic flags MEDIUM/HIGH alarm on specific access perspectives.

3 — Eval Suite
33
Critic Fixtures
4 CLEAN 21 HAS-BUGS 5 FLAWED 3 ADVERSARIAL
25
Planner Fixtures
5 ARIA 5 Keyboard 15 Domain
25
Perspective Fixtures
16 HAS-BUGS 5 CLEAN 4 ADVERSARIAL
4 — Benchmark Results

Critic: All Models Compared (33 fixtures)

Platform Model Fixtures Pass Must-find Status
Claude Haiku 4.5 33 28/33 100% 85% PASS
Claude Sonnet 4.6 5 (esc) 4/5 n/a 80% PASS
Claude Sonnet 4.6 + think 1 (esc) 1/1 n/a 100% PASS
Ollama qwen3:32b (BLIND 2026-07-13) 33 33/33 97% 100% PASS
Ollama qwen3:32b (non-blind) 33 33/33 96% 100% PASS
Ollama qwen3.5:27b 17 16/17 100% 94% PASS
Ollama llama3.3:70b (BLIND 2026-07-13) 33 33/33 92.6% 100% PASS
Ollama qwen3.5:latest (BLIND 2026-07-13) 33 33/33 98.5% 100% PASS
Ollama llama3.3:70b (non-blind) 7 7/7 86% 100% PASS
Ollama qwen3.5:latest (non-blind) 7 7/7 86% 100% PASS
OpenAI GPT-5.2 33 30/33 100% 91% PASS
OpenAI GPT-5.2 low 3 (esc) 1/3 n/a 33% PASS
OpenAI GPT-5.5 2 (esc) 1/2 n/a 50% PASS
OpenAI GPT-5.5 low 1 (esc) 1/1 n/a 100% PASS
Gemini 2.5 Flash 33 31/33 98% 94% PASS — 1 CLEAN FP, 1 ADVERSARIAL miss; pro escalation quota-pending

Perspective-Audit: qwen3:32b (25 fixtures, blind re-run 2026-07-13)

Tier Fixtures Pass Warn Fail Must-find
HAS-BUGS 16 16 0 0 36/37
ADVERSARIAL 4 4 0 0 100%
CLEAN (blind) 5 0 1 4 n/a

Blind CLEAN collapse: 4/5 clean fixtures draw false REVISE/BLOCK verdicts without the answer key (historical non-blind CLEAN was 4 WARN / 1 FAIL — answer-key-assisted). Detection tiers are blind-confirmed; receipts in evals/results/ollama-blind/README.md. Hint caveat (2026-07-16): blind lanes still saw inline // BUG: fixture hint comments — detection numbers are hint-assisted upper bounds (BENCHMARK.md hint-comment disclosure).

5 — Claude Escalation Strategy

Bottom-up: Cheapest tier that passes all 33 fixtures

Haiku 4.5
85%
28/33 pass · ~$0.45
5 failures: 2 CLEAN + 3 ADV
→
Sonnet 4.6
97%
4/5 escalated pass · +$0.18
1 remaining: modal-complete-clean
→
Sonnet + think
100%
1/1 escalated pass · +$0.02
Total cost: ~$0.65

Key finding: Haiku catches every real bug (21/21 HAS-BUGS, 5/5 FLAWED). Failures are judgment calls — false positives on CLEAN code, wrong verdicts on ADVERSARIAL ambiguity. Detection is not the issue; calibration is.

6 — Cross-Platform

Ollama (Local)

qwen3:32b recommended

Blind-confirmed: critic 33/33 PASS, 97% must-find, 0 FP. Caveat: perspective CLEAN verdicts inflate blind (4/5 false REVISE/BLOCK); detection numbers are hint-assisted upper bounds until a de-hinted re-run (fixtures carried inline BUG hint comments before 2026-07-16). 18.8 GB, ~2.6 min/fixture uncontended.

PRODUCTION

Claude API

Haiku + Sonnet escalation

100% pass via escalation. ~$0.65 for full suite. 67s avg per fixture.

PRODUCTION

Codex / OpenAI

GPT-5.2 → 5.5-low

GPT-5.2 passed 30/33. Escalation cleared the remaining CLEAN fixtures at GPT-5.5-low.

PRODUCTION

Gemini

2.5 Flash via gemini CLI

31/33 (94%) via gemini CLI; raw artifacts committed; pro escalation pending quota.

BASELINE
7 — Team Structure
R

bench-runner

Executes benchmarks across Claude API, Codex/OpenAI, local Ollama, and other hosted adapters when present; runs scoring scripts and reports raw results.

sonnet · Bash, Read
W

bench-reporter

Updates BENCHMARK.md, README.md, and EVAL-GAPS-PLAN.md with results and analysis.

sonnet · all tools
Q

bench-reviewer

Read-only quality auditor. Checks fixture/rubric consistency, scoring accuracy, result validity.

sonnet · read-only
B

fixture-builder

Creates and enriches fixture triplets (.md, .metadata.yaml, .rubric.yaml) across all 3 suites.

sonnet · all tools
8 — What's Next
  • ✓
    Eval suite complete — 83 fixtures across 3 suites, all with rubrics
  • ✓
    Ollama benchmarks — qwen3:32b 33/33 PASS, perspective 20P/4W/1F (non-blind); blind re-run 2026-07-13: critic confirmed 33/33 + 97% must-find, perspective 20P/1W/4F — CLEAN false positives surface blind. Blind full-suite critic lanes same day: qwen3.5:latest 33/33 + 98.5% (needs 32K ctx on 4 long fixtures), llama3.3:70b 33/33 + 92.6%
  • ✓
    Claude escalation — Haiku 85%, full pass at Sonnet+think, ~$0.65 total
  • ✓
    Team agents — 4 bench agents defined (runner, reporter, reviewer, builder)
  • ✓
    Codex/OpenAI benchmark — GPT-5.2 91%, full pass at GPT-5.5-low via bash ollama/codex-benchmark.sh
  • ✓
    Claude perspective lane (2026-07-13) — 25 fixtures via blind Opus subagents: post-003 scorer 20P/5W/0F, must-find 36/37; content-adjudicated 25/25 verdicts, 37/37 must-find, 0 CRITICAL/MAJOR on CLEAN fixtures. First blind lane — found the runners fed fixture answer keys to every prior lane; runners now strip them (guard test committed), scorer verdict/keyword bugs fixed (post-003), historical rows caveated as non-blind. Residuals: API claude-escalate cost-ladder variant blocked on a valid key. (Local blind re-runs: done 2026-07-13, see the Ollama benchmarks row above.)
  • ✓
    Gemini critic baseline — 2.5 Flash 31/33 (94%) via gemini CLI; raw artifacts committed to evals/results/gemini/; pro escalation pending quota
  • ✓
    Planner benchmarks — 25/25 on two lanes: qwen3:32b 227/235 (96.6%), Claude Opus subagents 234/235 (99.6%); Codex lane gated (plan 010)
  • ✓
    keyboard-a11y-tester adopted (2026-07-10) — 4th a11y-test mode: live-URL journey audits with emulated SR persona; cross-validated vs all 33 critic fixtures, zero false positives outside one pre-disclaimed class; our PR #7 merged upstream, pin bumped to release 0.5.0
  • ✓
    virtual-screen-reader adopted (2026-07-11/13) — 5th a11y-test mode: component-level SR assertions (names, reading order, live-region announcements) in Vitest/Jest/Storybook; fixture sweep 4/4 defects + 5/5 fixes + clean-fixture zero false-silent; Storybook 10 lane verified 12/12