22 tasks. July 9, 2026 closed at 24.8x weighted leverage across 593.5 human-equivalent hours in 1,435 minutes of wall-clock time. Supervisory leverage came in at 481.2x.
That is 14.8 weeks of human-equivalent throughput in 23.9 hours. The ceiling was 109.1x; the floor was 3.8x. 22 of the 22 entries came from a single project.
Task Log
| # | Task | Human Est. | Claude | Sup. | Factor |
|---|---|---|---|---|---|
| 1 | AVIAN simulation master list: cataloged 96 existing interactive components, orchestrated 22-agent workflow across all 261 academic domain specs, produced 2141-simulation master list + filterable gallery… | 160.0h | 88m | 4m | 109.1x |
| 2 | AVIAN simulations build-out: photoreal scriptable Three.js blood-flow reference (IBL/bloom/DoF/SSS + SimController narration-cue API, vendored offline three.js), upgraded gallery into live local test app with… | 48.0h | 55m | 4m | 52.4x |
| 3 | Documentation | 44.0h | 54m | 8m | 48.9x |
| 4 | WP13 platform-build closer: P0 boot-cache root-cause+fix (proven round-trip), both field guides rewritten w/ what-changed, doc-sync 7/7, spend_ledger writer, fleet hooks, final all-repo sweep | 44.0h | 56m | 1m | 47.1x |
| 5 | WP12 [engine subsystem] MCP catalog (66 tools, live stdio smoke) + CLI (13 groups, 2 live bugs caught) + docs set + backend punch list (batch-upsert 61-67x, spend/benchmarks/routing endpoints, severity-tile… | 44.0h | 82m | 1m | 32.2x |
| 6 | Option-4 round 3: voting engine + per-domain-class routing + fresh B-set protocol (checks-lib) | 16.0h | 36m | 4m | 26.4x |
| 7 | M1 status-vocab migration + ADR-0007 execution across 7 repos: 5-state lifecycle live (9 candidates), 70 CE [restricted], 16 excluded, private extinct; leak gates verified; 2 pre-existing bugs fixed;… | 32.0h | 82m | 1m | 23.4x |
| 8 | Option-4 round 2: grounded per-option judge + LLM-mode coverage calibration (checks-lib) | 12.0h | 33m | 3m | 22.0x |
| 9 | Coding | 40.0h | 111m | 5m | 21.6x |
| 10 | WP9a judge fixes + benchmark harness + real calibration (blind_compare confirm; fc check redesigned to clear bars; found the [engine subsystem] label-swap) + orchestrator [engine subsystem] server fix… | 28.0h | 95m | 1m | 17.7x |
| 11 | Updated both hosted field-manual artifacts ([content generation] Engine + AWS SAA-C03 worked example) to platform reality: new Registry section w/ SVG, blocking gates, 85-check framework + judge design,… | 6.0h | 25m | 2m | 14.4x |
| 12 | Option-4 round 7: mutator bug fixes + clean powered F-sets + pre-registered confirmation (checks-lib) | 10.0h | 45m | 3m | 13.4x |
| 13 | Post-[content generation] cleanup automation: [engine subsystem]-service build-dir cleanup after ingest + [engine subsystem] ingest-staging cleanup and real admin-prune delete mode | 14.0h | 65m | 8m | 12.9x |
| 14 | Post-[content generation] cleanup automation: [engine subsystem] build-dir cleanup + [engine subsystem] staging cleanup/real prune mode | 14.0h | 66m | 8m | 12.8x |
| 15 | G1 SAA-C03 golden-package operator pass: fixed reachable deterministic defects (v2-v4 lineage), discovered + precisely documented 7 platform bugs + 2 calibration classes; zero content harmed, live untouched | 7.0h | 37m | 1m | 11.4x |
| 16 | Option-4 round 7-EXT: pooled measurement extension → campaign GO + arm-ready package (checks-lib) | 4.0h | 21m | 2m | 11.3x |
| 17 | WP9b final: MCQ-judge escalation ladder (forced-CoT + ensemble, real calibration) — clean negative result establishing cheap-tier negation ceiling; JudgeClient hard-timeout fix; harvest economics quantified | 20.0h | 109m | 1m | 11.0x |
| 18 | Option-4 round 4: per-miss error analysis + negation-emphasis variant + C-set routing table (checks-lib) | 14.0h | 79m | 4m | 10.7x |
| 19 | Option-4 round 6: powered E-sets + Wilson bounds + mutator forensics + grok grounded probe (checks-lib) | 16.0h | 107m | 5m | 9.0x |
| 20 | Option-4 round 5: targeted fixes (retry/neg-scoped/derivable-support) + D-set close-out + tuning-pool bug catch (checks-lib) | 12.0h | 91m | 6m | 7.9x |
| 21 | FLEET-A operator pass: complete 407-domain dry-run ledger (2849 calls, 84s) + apply inventory + independent root-cause of action-apply and validate-noop blockers + 2 new bugs; [cost]spend | 5.0h | 43m | 1m | 7.0x |
| 22 | WP9a round 2: MCQ lever A/B re-benchmark (lever A reverted on evidence, lever B kept), [engine subsystem]-fix follow-through recalibration, honest no-winner verdict + escalation path | 3.5h | 55m | 1m | 3.8x |
Aggregate Statistics
| Metric | Value |
|---|---|
| Total tasks | 22 |
| Total human-equivalent hours | 593.5 |
| Total Claude minutes | 1,435 |
| Total supervisory minutes | 74 |
| Total tokens | 15,552,467 |
| Weighted average leverage factor | 24.8x |
| Weighted average supervisory leverage factor | 481.2x |
| Human-equivalent weeks | 14.8 |
Analysis
The highest factor of the day came in at 109.1x and the lowest at 3.8x, a spread of 28.6 times between the two. That is a wide range for a single day, and it usually means the day mixed mechanical work with work that needed real judgement.
The largest single entry accounted for 160.0 of the 593.5 human-equivalent hours, or 27 percent of the day. No single task dominated the total, so the weighted average is representative.
Supervisory time was 74 minutes against 1,435 minutes of execution, a ratio of about 1 to 19. Supervisory leverage of 481.2x is the figure I find most honest, because it measures the hours I actually spent rather than the hours a machine spent on my behalf.
Every figure here is recorded at the time the work is done rather than reconstructed afterwards. The human estimate is my own judgement and carries the uncertainty that implies; the minutes and tokens are measured. The full dataset, including this day, is available for download.