Skip to main content
AI02 June 2026

Mostly testing

7 tasks. June 2, 2026 closed at 27.1x weighted leverage across 124.0 human-equivalent hours in 275 minutes of wall-clock time. Supervisory leverage came in at 354.3x.

That is 3.1 weeks of human-equivalent throughput in 4.6 hours. The ceiling was 80.0x; the floor was 7.9x. 7 of the 7 entries came from a single project.

Task Log

#TaskHuman Est.ClaudeSup.Factor
1Full accessibility audit + fix across all four avian-app clients (web/electron/android/ios): jsx-a11y errors, label associations, autofocus, tablist roles, jsx-a11y plugin + axe coverage 8->40, 100+ Compose…80.0h60m3m80.0x
2Full WCAG 2.1 AA accessibility audit across all 4 avian-app clients (web/electron source + 39-route axe sweep; 11 mechanical fixes, 3 false positives triaged, ledger reconciled, native heuristics)7.0h8m2m52.5x
3Run full accessibility audit across all four avian-app clients (web/electron/iOS/Android) — hand-created Android AVD, booted sim+emulator, ran axe-sweep/vitest/XCUITest/Espresso a11y suites4.0h9m1m26.7x
4Fix iOS + Android accessibility audit failures: 44pt hit-target on All Courses button + opaque-white hero subtitles (contrast) + accessibilityHidden on decorative SF Symbols; repair Android Compose a11y test…3.0h13m1m13.8x
5Resume avian-[engine subsystem] OMNISCIENT-ONLY cloud sweeps: diagnose OOM root cause, generate 42 omni profiles, write concurrency-capped batching runner, bring up engine+[engine subsystem] backend at hard…5.0h27m2m11.1x
6Finish AP Precalc math content: fix 3 [content generation] bugs (rep_pack schema-drop, judge-pool deadlock, id-collision), Sonnet regen of 38 goals, standalone re-judge of 381 items, prune to 0…20.0h120m6m10.0x
7Audit and review5.0h38m6m7.9x

Aggregate Statistics

MetricValue
Total tasks7
Total human-equivalent hours124.0
Total Claude minutes275
Total supervisory minutes21
Total tokens1,789,000
Weighted average leverage factor27.1x
Weighted average supervisory leverage factor354.3x
Human-equivalent weeks3.1

Analysis

The highest factor of the day came in at 80.0x and the lowest at 7.9x, a spread of 10.1 times between the two. That is a wide range for a single day, and it usually means the day mixed mechanical work with work that needed real judgement.

The largest single entry accounted for 80.0 of the 124.0 human-equivalent hours, or 65 percent of the day. One task carrying that much of the total is worth noting; the average is doing less work than it appears to.

Supervisory time was 21 minutes against 275 minutes of execution, a ratio of about 1 to 13. Supervisory leverage of 354.3x is the figure I find most honest, because it measures the hours I actually spent rather than the hours a machine spent on my behalf.

Every figure here is recorded at the time the work is done rather than reconstructed afterwards. The human estimate is my own judgement and carries the uncertainty that implies; the minutes and tokens are measured. The full dataset, including this day, is available for download.