Skip to main content
AI17 July 2026

Coding and testing

20 tasks. July 17, 2026 closed at 13.2x weighted leverage across 279.2 human-equivalent hours in 1,265 minutes of wall-clock time. Supervisory leverage came in at 197.1x.

That is 7.0 weeks of human-equivalent throughput in 21.1 hours. The ceiling was 22.7x; the floor was 3.3x. 16 of the 20 entries came from a single project.

Task Log

#TaskHuman Est.ClaudeSup.Factor
1Coverage agent A: api core +923 statements (328 tests; 2 production bugs fixed incl. never-working micro-challenge endpoint)28.0h74m6m22.7x
2Scenario E PASS: real-study advancement harness (1040 lines), 3 transitions verified live, 2 methodology fixes (ground-truth polling, order-0 gate), competence-floor chain traced9.0h28m1m19.3x
3Fix loop: 4 CONCERN packages root-caused (store/staging sync gap, NOT content) + fixed free via re-ingest/materialize/reload — re-sweep 4/4 at 100% (1363/1363)8.0h29m1m16.6x
4Fix local avatar upload: endpoint-aware S3 helper in avian-api routing avatars/resume uploads to the local MinIO emulator; auth-service parity; provisioned buckets + policies5.5h20m3m16.5x
5WP-3.1 engine focus store + dual-pool planner: 009 focus table + FocusConfig + frontier/maintenance budgets + CRUD (78 tests)36.0h135m6m16.0x
6Coverage agent C: manifold/ring/training/governance +634 statements (159 tests; hnswlib unlock)20.0h75m5m16.0x
7WP-2.1 engine placement service: 3-layer seeding + entity_placements + TT endpoints + posterior placement priors (85 tests)28.0h110m6m15.3x
8Fix profile avatar upload (local S3 emulator) + P0 gateway entitlement bug blocking all paid courses + remove fabricated-question fallback + honest error UI, then audit & fix all 17 web-app error-maskings…26.0h110m5m14.2x
9SAP-C02 scenario fix loop: 4 ceilings (pace/decay/seeding-trap/gate-deadlock) + 2 engine features + scenario C exam PASS32.0h150m12m12.8x
10Debugging6.0h29m1m12.4x
11WP-2.2 seeded-prior decay blend: two-gate root cause (honesty clamp dominant), patch-package application + fixture repairs, ADR-0011 — suite 6999 green16.0h79m2m12.2x
12Scenario B: 9-journey mid-course omniscient sweep + comparator — 18/18 readiness, 6/8 faster-assertion misses root-caused to unbuilt WP-2.2 seeded-prior cliff (telemetry archaeology + engine source trace)5.0h28m1m10.7x
13WP-2.4 web Position Fix flow + CHEM1B catalog surfacing + QA-status gate (351 tests green)20.0h115m10m10.4x
14Audit and review3.0h18m3m10.0x
15WP-3.3 gateway focus proxies + focus.updated SSE + completion-candidate seam (23 tests)5.0h30m5m10.0x
16WP-3.2 unit completion gate + consent-gated advancement (62 tests; effort-bounded amendment)16.0h105m7m9.1x
17WP-2.3 gateway placement proxies + SSE placement.completed + enrollment seed (29 tests + pre-existing coverage gate fix)6.0h40m4m9.0x
18WP-3.2 follow-up: unit.completion-candidate relay wired to real engine payload (4 tests)2.5h17m3m8.8x
19[engine subsystem] telemetry WP: RAG/omni telemetry parity + comparator snapshot fallback + snapshot-dir relocation + Temporal DB persistence (59 tests)6.0h50m3m7.2x
20Fleet sweep 100% closure: final 3 stale-staging misses fixed (signature-verified, push+materialize+reload, re-sweep 3/3 at 100%) — 406/406 final tally1.2h23m1m3.3x

Aggregate Statistics

MetricValue
Total tasks20
Total human-equivalent hours279.2
Total Claude minutes1,265
Total supervisory minutes85
Total tokens12,122,660
Weighted average leverage factor13.2x
Weighted average supervisory leverage factor197.1x
Human-equivalent weeks7.0

Analysis

The highest factor of the day came in at 22.7x and the lowest at 3.3x, a spread of 7.0 times between the two. That is a wide range for a single day, and it usually means the day mixed mechanical work with work that needed real judgement.

The largest single entry accounted for 28.0 of the 279.2 human-equivalent hours, or 10 percent of the day. No single task dominated the total, so the weighted average is representative.

Supervisory time was 85 minutes against 1,265 minutes of execution, a ratio of about 1 to 15. Supervisory leverage of 197.1x is the figure I find most honest, because it measures the hours I actually spent rather than the hours a machine spent on my behalf.

Every figure here is recorded at the time the work is done rather than reconstructed afterwards. The human estimate is my own judgement and carries the uncertainty that implies; the minutes and tokens are measured. The full dataset, including this day, is available for download.