Skip to main content
AI08 July 2026

Coding and testing

34 tasks. July 8, 2026 closed at 41.8x weighted leverage across 1,916.5 human-equivalent hours in 2,749 minutes of wall-clock time. Supervisory leverage came in at 991.3x.

That is 47.9 weeks of human-equivalent throughput in 45.8 hours. The ceiling was 154.8x; the floor was 4.8x. 34 of the 34 entries came from a single project.

Task Log

#TaskHuman Est.ClaudeSup.Factor
1Content [content generation] Platform comprehensive design: [engine subsystem] registry + validation framework + audit-as-code + benchmarked model routing + MCP/API/console + 13-WP implementation plan80.0h31m6m154.8x
2AccelaStudy Classroom: research all 51 US homeschool jurisdictions (requirements/accreditation/lab-science/personal-finance) into state-requirements.json; author 30 world-language 102 specs + Personal Finance…180.0h90m6m120.0x
3Coding320.0h170m1m112.9x
4Review alpha K-12 domains; propose grades 5-12 curriculum for classroom.accelastudy.ai (scope-and-sequence + state-requirement mapping + published HTML artifact)18.0h13m3m83.1x
5Audit AccelaStudy activity system; brainstorm + design 34-concept gamified activities playbook artifact16.0h12m2m80.0x
6WP2 [engine subsystem] registry core: 13-table schema + hardlink version store + real 407-pkg backfill (33.6GB) + 54-endpoint API + ingest protocol + jobs/executor + lifecycle/leases/promotion + catalog…140.0h125m1m67.2x
7Documentation40.0h40m6m60.0x
8WP4 audit parity + adjudications: 299-item disposition map (zero silent losses), parity gate PASS, dup-options false-positive verdict, 34k dangling-pair confirmation, canonical duplicate-key root-cause; +…44.0h44m1m60.0x
9WP10 [engine subsystem] operations console: 52 components, 3 faces, 10-tab content drill-down, issue-centric flow live-verified on real fleet data, Playwright 12/12 incl mobile+dark, caught real API shape…88.0h92m1m57.4x
10WP0 content [content generation] platform: scaffold services/[engine subsystem] (FastAPI backend w/ API-key roles auth + async alembic + 63 tests/99% cov; React/Vite frontend w/ ThemeProvider + ApiKeyGate +…28.0h30m15m56.0x
11WP11 rubrics: backfill_rubrics action (plan/apply, schema-validated) + goal-linkage embed matcher (floors, fallback) + rubric-aware case generation; SAA-C03 pilot proven (linkage 8-to-0, coverage 100%)36.0h41m1m52.7x
12WP0 [engine subsystem] + avian-content-checks scaffold: 2 private repos + FastAPI backend (auth roles) + themed React shell + check-framework lib (170 tests; 99-100% cov; ports registered)28.0h33m1m50.9x
13Remediate notification-service + onboarding-service: ~58 verified security/correctness fixes across both (all tiers), ~218 new regression tests, suites green (NS 466 / OS 395)100.0h125m3m48.0x
14Coding56.0h72m1m46.7x
15WP5a: model-based checks ([engine subsystem]+LLM judges), campaign runner, WP4 diffs applied for avian-content-checks65.0h87m5m44.8x
16WP5a model-based checks: 13 [engine subsystem]/LLM judges (parse-failure-proof verdict channel) + campaign runner w/ hard budget stops + WP4 diffs applied + real calibration that caught a broken judge for…65.0h89m1m43.8x
17WP1 content [content generation] platform: CONTENT_TYPES registry + full Pydantic content schemas in avian-[engine subsystem]-runtime (35 content types, ~20 schema modules, fixture mini-corpus, real-data…36.0h50m8m43.2x
18WP8: actions catalog (13 [engine subsystem] worker-side runners incl. 3 LLM-gated + FakeLLM tests) + [engine subsystem] hub (catalog/scope/planning/API/2 hub handlers) + 12-endpoint content-browse API;…72.0h100m8m43.2x
19WP8 actions catalog (13 actions, new-version-via-ingest, parity on 6 w/ 2 real bugs caught) + content-browse API (10 endpoint groups) + 14 legacy scripts deleted across engine+[engine subsystem]72.0h100m1m43.2x
20WP1 CONTENT_TYPES registry (35 rows w/ requiredness predicates) + 18 content schema modules + fixture mini-package + writer-schema drift tests; 3 real drift bugs found+fixed; PACKAGE_FILES retired to derived…36.0h52m1m41.5x
21WP3 deterministic check catalog: 71 checks + 31 mutators + runner/CLI + first-ever 552,976-finding fleet scan (backfilled record)72.0h104m1m41.5x
22Critical-goal requirement in domain-spec generation (module+gate+authoring wiring+5-layer critical-field threading), enforce in runtime_artifacts+tests, apply to 360 specs+16 packages; Class1 canonical…44.0h65m4m40.6x
23WP7 containerized job execution: real [engine subsystem]-worker image (2.27GB, 2 build bugs fixed) + container-per-job executor (kill-cancel 0.29s, restart-reattach, resource caps) + /complete client wiring;…52.0h80m1m39.0x
24WP3b wire real 71-check catalog into [engine subsystem]: RealChecksBackend + findings dedup/resolve persistence + /validate + audit job kinds + issue-centric summary endpoints; 2 real bugs fixed (candidacy…60.0h114m1m31.6x
25Complete code review of notification-service and onboarding-service (8 parallel review passes + hand-verification; 62 findings)18.0h35m3m30.9x
26Extend Gap 8 outcome-aware credit to all clients: electron IPC options bridge + 6 graded screens + React TDZ fix; iOS EngineClient parity + pbxproj build-break fix; web+android reconciled (case_study…13.0h30m2m26.0x
27Testing8.0h20m3m24.0x
28Coding20.0h55m4m21.8x
29Debugging40.0h110m3m21.8x
30AVIAN cleanup buckets 1-2: log dumps + superseded scripts across 7 repos (grep-verified; engine suite 7294 green; ~290MB reclaimed)5.0h20m1m15.0x
31Class 4 P2.2: build goal_weights+goal_similarity runtime artifacts for 32 packages (via patched write_runtime_artifacts, criticals applied) + finish Class 3 P26.3 question-tier fix + P6/P19 forensic…7.0h35m2m12.0x
32P4.1 [engine subsystem] (11 pkgs +5016 pairs, 0 goals starved) + P24.2 question-quality repair (160 dup->0, 33 missing expl->0) + [engine subsystem] pair-gate (16.9% weak flagged) + MCQ correctness…7.0h60m3m7.0x
33Move 10 private packages (HP/ASOIAF/48-Laws ~1.1GB) out of active store to private-packages/ + durable private-status filter in content-audit pkg_dirs + canonical 417->407 + purge from both S3 backup regions…2.5h25m1m6.0x
34Finalize 14 held IB HL @0.93 domains via scope-aware adversarial-validation fix + critique persistence + blind-panel diagnosis + provenance ledger + crash/config fixes + 2 pipeline field-manuals48.0h600m15m4.8x

Aggregate Statistics

MetricValue
Total tasks34
Total human-equivalent hours1,916.5
Total Claude minutes2,749
Total supervisory minutes116
Total tokens44,625,000
Weighted average leverage factor41.8x
Weighted average supervisory leverage factor991.3x
Human-equivalent weeks47.9

Analysis

The highest factor of the day came in at 154.8x and the lowest at 4.8x, a spread of 32.3 times between the two. That is a wide range for a single day, and it usually means the day mixed mechanical work with work that needed real judgement.

The largest single entry accounted for 80.0 of the 1,916.5 human-equivalent hours, or 4 percent of the day. No single task dominated the total, so the weighted average is representative.

Supervisory time was 116 minutes against 2,749 minutes of execution, a ratio of about 1 to 24. Supervisory leverage of 991.3x is the figure I find most honest, because it measures the hours I actually spent rather than the hours a machine spent on my behalf.

Every figure here is recorded at the time the work is done rather than reconstructed afterwards. The human estimate is my own judgement and carries the uncertainty that implies; the minutes and tokens are measured. The full dataset, including this day, is available for download.