11 tasks. May 25, 2026 closed at 18.6x weighted leverage across 79.5 human-equivalent hours in 256 minutes of wall-clock time. Supervisory leverage came in at 136.3x.
That is 2.0 weeks of human-equivalent throughput in 4.3 hours. The ceiling was 33.6x; the floor was 5.3x. 11 of the 11 entries came from a single project.
Task Log
| # | Task | Human Est. | Claude | Sup. | Factor |
|---|---|---|---|---|---|
| 1 | [engine subsystem]: prompt caching + Anthropic Batches API integration across [content generation] scripts in core/avian-engine (1649 LOC, 5 files) | 28.0h | 50m | 5m | 33.6x |
| 2 | Invert leverage tracking policy: CSV first then cloud second both mandatory; patched global CLAUDE.md Rules block, /fix skill Step 6k, 16 tool-loader Step 4 blocks, and /leverage-post Phase 2 reconciliation | 4.0h | 8m | 2m | 30.0x |
| 3 | Infrastructure | 4.0h | 8m | 4m | 30.0x |
| 4 | Audit and review | 3.0h | 6m | 2m | 30.0x |
| 5 | Content production | 12.0h | 25m | 2m | 28.8x |
| 6 | /leverage-post reconciliation Phase 1+2: backfilled 139 CSV rows across 12 days (5/14-5/25), verified all in sync, 0 stragglers remaining | 1.5h | 4m | 1m | 22.5x |
| 7 | core/avian-[engine subsystem] rebuild: brain answerer switched to direct Anthropic SDK with prompt caching (257 LOC), all zero/pmp sweep profiles flipped to omniscient:false (46 files), headless runner… | 10.0h | 30m | 4m | 20.0x |
| 8 | Coding | 3.0h | 10m | 2m | 18.0x |
| 9 | avian-audits: content-audit P4.1 pair-density check with PGWA-class detection (61 LOC), canonical.json headline counts bumped to 2026-05-25 audit snapshot, per-activity-format trackers added (scenarios,… | 4.0h | 15m | 3m | 16.0x |
| 10 | core/avian-engine: autopilot_service legacy coverage-damping ceiling lifted + bulk_amplify_fleet and bulk_backfill_recall custom_id format fix with error logging | 2.0h | 10m | 2m | 12.0x |
| 11 | Design and build AVIAN staging environment: 4 Terraform stacks (valkey/engine/api/app-web) reusing prod ALB/RDS/S3 + /staging skill with up/down/status/extend and at-based 3h auto-teardown | 8.0h | 90m | 8m | 5.3x |
Aggregate Statistics
| Metric | Value |
|---|---|
| Total tasks | 11 |
| Total human-equivalent hours | 79.5 |
| Total Claude minutes | 256 |
| Total supervisory minutes | 35 |
| Total tokens | 1,063,000 |
| Weighted average leverage factor | 18.6x |
| Weighted average supervisory leverage factor | 136.3x |
| Human-equivalent weeks | 2.0 |
Analysis
The highest factor of the day came in at 33.6x and the lowest at 5.3x, a spread of 6.3 times between the two. That is a wide range for a single day, and it usually means the day mixed mechanical work with work that needed real judgement.
The largest single entry accounted for 28.0 of the 79.5 human-equivalent hours, or 35 percent of the day. No single task dominated the total, so the weighted average is representative.
Supervisory time was 35 minutes against 256 minutes of execution, a ratio of about 1 to 7. Supervisory leverage of 136.3x is the figure I find most honest, because it measures the hours I actually spent rather than the hours a machine spent on my behalf.
Every figure here is recorded at the time the work is done rather than reconstructed afterwards. The human estimate is my own judgement and carries the uncertainty that implies; the minutes and tokens are measured. The full dataset, including this day, is available for download.