Skip to main content
AI13 April 2026

Coding and testing

23 tasks. April 13, 2026 closed at 32.7x weighted leverage across 858.0 human-equivalent hours in 1,576 minutes of wall-clock time. Supervisory leverage came in at 287.6x.

That is 21.4 weeks of human-equivalent throughput in 26.3 hours. The ceiling was 200.0x; the floor was 6.7x. 15 of the 23 entries came from a single project.

Task Log

#TaskHuman Est.ClaudeSup.Factor
1Build 3-tier course catalog system for accelastudy.ai: data build script, courses.jinja catalog, category-page.jinja provider pages, course-page.jinja detail pages, subscription detection JS. 892 total pages.40.0h12m5m200.0x
2Full tools fleet consistency audit (auth, API, DB, theme, frontend) across 15 repos + canonical standards doc in tools/CLAUDE.md40.0h12m5m200.0x
3Testing160.0h55m3m174.5x
4Build @avian/app-shell package + migrate all 15 tool frontends (auth, theme, login, FOUC, fonts, favicons) — 244 test files, 0 failures120.0h45m8m160.0x
5Coding40.0h30m5m80.0x
6Analyze service gaps across 2048 labs, build service config registry for 322 slugs, AdaptiveServiceDashboard (all cloud providers), TerminalView (git/docker/k8s/cloud-shell), CodeEditorView…80.0h70m15m68.6x
7Build and run Playwright E2E tests for all 2048 console-sim labs: config, data-testid attributes, test manifest generator, guided smoke spec, watch mode spec, run full suite 2048/2048 pass24.0h25m3m57.6x
8Fix [engine subsystem] broken prompts, garbage RAG chunks, and busted UI — rewrote MCQ pipeline to use question bank loader, fix correctness detection, store factual chunks, add event IDs, fix calibration…16.0h18m5m53.3x
9Refactor console-sim to single generic executor: deploy all 2048 labs, create generic executor factory with UI automation, delete 290 hand-written executor files, simplify registry40.0h45m8m53.3x
10Fix 6 CSS defects in avian-app-web + avian-activities-react (font sizes, flashcard occlusion, timed recall instability, SVG overflow). Add 26 data-testids across 7 activity components. Update 9 page object…24.0h30m3m48.0x
11Audit and review40.0h55m5m43.6x
12Debugging16.0h25m3m38.4x
13AVIAN [engine subsystem] night session: ran 45-day simulation (ongoing), fixed datetime bugs, MCQ dual-variant selectors, verifier method/options bugs, tour dismissal, activity filtering. Added Valkey prompt…60.0h240m20m15.0x
14Testing40.0h180m15m13.3x
15Testing8.0h45m5m10.7x
16Debugging6.0h35m5m10.3x
17Design and frontend3.0h18m8m10.0x
18Autonomous [engine subsystem] simulation iteration — fixed 15+ issues across study pipeline, RAG ingestion, exam trigger, exam execution, and worker completion until Alex Chen ran full study→exam journey…80.0h480m30m10.0x
19Testing3.0h20m5m9.0x
20Testing4.0h28m5m8.6x
21Coding6.0h45m8m8.0x
22Deployment6.0h45m5m8.0x
23Coding2.0h18m5m6.7x

Aggregate Statistics

MetricValue
Total tasks23
Total human-equivalent hours858.0
Total Claude minutes1,576
Total supervisory minutes179
Total tokens6,464,000
Weighted average leverage factor32.7x
Weighted average supervisory leverage factor287.6x
Human-equivalent weeks21.4

Analysis

The highest factor of the day came in at 200.0x and the lowest at 6.7x, a spread of 30.0 times between the two. That is a wide range for a single day, and it usually means the day mixed mechanical work with work that needed real judgement.

The largest single entry accounted for 40.0 of the 858.0 human-equivalent hours, or 5 percent of the day. No single task dominated the total, so the weighted average is representative.

Supervisory time was 179 minutes against 1,576 minutes of execution, a ratio of about 1 to 9. Supervisory leverage of 287.6x is the figure I find most honest, because it measures the hours I actually spent rather than the hours a machine spent on my behalf.

Every figure here is recorded at the time the work is done rather than reconstructed afterwards. The human estimate is my own judgement and carries the uncertainty that implies; the minutes and tokens are measured. The full dataset, including this day, is available for download.