I haven't opened a video editor once. There's no timeline, no keyframe panel, no export dialog. Every one of these videos is a web page that Claude Code wrote, rendered frame by frame in a headless browser, checked against a long list of rules, and handed back to me as an MP4.
A few people have asked how that actually works. Here is the whole process, including the parts that went wrong. If you'd rather watch than read, here's the 96-second version, made the same way. It's narrated by an ElevenLabs clone of my voice, and the captions are burned in, so it works muted.
The Videos Are Web Pages
The engine underneath is HyperFrames, an open-source framework from HeyGen that renders HTML to video. A composition is an ordinary HTML file. Its layout is CSS, its motion is a GSAP timeline, and a handful of data attributes tell the renderer when each clip starts and how long it runs. To render, HyperFrames opens the page in headless Chrome, seeks the timeline to the instant of each frame, takes a screenshot, and hands the stack of frames and the mixed audio to FFmpeg.
That sounds like a strange way to make a video until you notice what it buys you. Every design decision is text. A headline that lands half a beat late is a one-line diff. A layout that sits 45 pixels off center (more on that below) is a constant in a JavaScript file. And because the timeline can be sought to any instant, frame 312 of the final render looks exactly like frame 312 of the preview, every single time.
Nothing on Screen Is Invented
This is the rule the whole project hangs on: every screen in these videos is the real product, captured live. Nothing is mocked up, redrawn, retyped, or retouched. A prospective customer who spots a fake screen in a product video has every reason to stop trusting the product.
So before a single frame gets built, a capture agent stands up the entire application on my machine: a production build of the web app, the API, and the adaptive engine, serving the same course content students see. It creates a fresh learner through the app's own sign-up, an illustrative student named Jordan Lee (the videos label the account "Illustrative learner"), and then drives the app with Playwright the way a student would. It enrolls, opens a lesson, answers questions, misses one on purpose, and opens the explanation. Answers come from the course packages' own answer keys, so when the script clicks the right answer, it's actually right.
Every screen is shot twice: on a 1,440 × 900 desktop at 2x, and on a 390 × 844 phone at 3x. The portrait videos should show the app the way it looks on a phone, not a desktop screenshot squeezed into a tall frame.
Each capture gets logged with:
- the route, the app build, the engine commit, and the content versions it was served
- a SHA-256 hash of the image
- the measured position of every element a video might circle or crop to
- a pass or a fail against a keep-out list: error states, empty states, anything unfinished, and anything that would make a claim the product doesn't fully back yet
Motion was the one surprise. Playwright can record video, but it records in CSS pixels, so a page rendered at 2x fills only the top-left quarter of the recording. The fix was to skip the recorder and take bursts of full-resolution screenshots instead. One round of Recall Sprint on desktop came back as 544 frames at 2,880 × 1,800, and the video plays them back like a flip book.
Every Number Comes From One File
AccelaStudy keeps every figure a customer might see in one file of canonical values: prices, course and question counts, the trial terms, the web address, and the exact form of every trademark. Nothing in a video types a number. A frame asks for the monthly price and gets whatever that file says today, formatted the way the file says to format it. The caption text for each post works the same way. It's written with placeholders and filled in at delivery.
The test suite enforces all of it. A frame that contains a typed four-digit figure fails. So does a frame that says "pass" (the SAT®, PSAT/NMSQT®, and ACT® have no passing score), promises a score gain, or says "Know your score before test day" without the caveat that practice scores are estimates. By the end of the first batch the suite held 386 tests, and it runs before every commit.
The Storyboard Is the Clock
Each post starts as a storyboard: a Markdown file with a little YAML at the top. It lists the scenes in order, what each one shows, and how long each one lasts. Those lengths are whole bars of music, not arbitrary seconds.
The music is ours. Claude generated four short beds with ElevenLabs Music, two at 120 beats per minute and two at 96, all built in the same shape: a hook, a groove, a lift, and an ending with a real final hit. A post takes bars from the start of a bed and splices to that bed's ending exactly on a downbeat, with a 12-millisecond crossfade, so every video ends on a musical button instead of a fade-out. Because every scene is measured in bars, every cut lands on the beat without anyone nudging it there.
Sound effects are cued in the storyboard by scene and second. Every render is normalized to −14 LUFS, the loudness the streaming platforms expect, with the true peak held at −1 dB or lower. The four beds came to 116 seconds of generated music from the credits my plan already included, so the soundtrack cost nothing extra.
A Kit, Not a Template
The videos share a small JavaScript kit the agents built in the first session. It draws the cover card, the app window and phone screen that hold the captures, the rings and cursor that point at things, the offer, and the end card.
The cover matters more than it looks. Frame one of every video is its cover, pixel for pixel, so the thumbnail a platform shows and the first frame of playback never jump. The render compares the two and fails if they don't match.
Each post is built twice. The portrait version isn't the landscape version squashed. It's laid out again for a tall screen, and everything that matters has to stay inside a safe area that TikTok's and Instagram's buttons and captions won't cover.
The Render Is a Gate
There is exactly one way to render: npm run render. It takes a machine-wide lock so two renders never fight over the processor, renders both layouts at 30 frames per second, and refuses to deliver anything that misses:
- the planned length, to within 50 milliseconds
- H.264 video with AAC audio at 48 kHz
- −14 LUFS integrated loudness, with the true peak at −1 dB or lower
- a first frame that matches the cover
- every line of text inside the portrait safe area
- the index at the front of the file, so the video starts playing before it finishes downloading
A 20-second post renders in about two minutes on my MacBook Pro. Delivery writes one folder per post with both videos, both covers, the caption text, alt text for the cover, and a checksum file.
Who Does What
The work is split into lanes, and each lane is its own Claude Code agent. One captures screens, one built the kit, and every post gets its own author. The main session, the one I actually talk to, orchestrates. It writes each lane's brief, watches their progress, and reviews every video before I ever see it.
That review is not a skim. It pulls stills at every cut, lays them out in contact sheets, draws guide lines over them, and reads the measurements for each file. Each agent also reports its progress out loud through my speakers, in a British butler's voice I picked, and to a chat channel I can read from my phone.
My part is smaller than you'd think and more important than it looks. I picked the 30 topics. I make the calls that should belong to a person: what the product is allowed to claim, and what gets fixed tonight instead of cropped around. And I watch the output the way a customer will, on a phone, which is how I caught the first real mistake.
What Went Wrong
This is the part I find most useful, so here it is unvarnished.
Every portrait video was off center. I opened the first five portrait cuts and everything sat slightly left. It was 45 pixels, in all five. The safe area was defined to keep clear of the button column on the right edge of TikTok and Reels, but it reserved that space on the right side only, so every layout centered on x = 495 instead of 540. The agents had followed the rule perfectly. The rule was wrong. The fix was a symmetric safe area, a 45-pixel shift in every portrait layout, fresh renders of all five, and a new test that fails if the kit's portrait layouts ever drift off center again. It took 38 minutes from my message to five re-delivered videos.
An agent sat for an hour waiting for permission. My Claude Code setup runs with permission prompts turned off. One command still asked for approval anyway: deleting a temporary folder the agent had created itself. It was late, nobody was at the keyboard, and the capture lane just sat there. The fix is now baked into every agent definition: don't delete anything, make a new folder instead.
Too many agents, one laptop. To go faster, I let four post authors and the capture agent run at the same time. Each one launches headless Chrome to check its own work. Along with everything else running on the machine, that pushed the load average past 60 on a 10-core laptop, and the capture of the in-app assistant started timing out. The rule now is one video at a time.
The agents found bugs in their own kit. In real renders, the cover vanished from frame one, because the compiler strips a timing attribute that the preview kept. Rings and the cursor painted underneath a new screen after a swap. A click's ripple showed up before the click. Screens that came to rest at a slight two-degree tilt broke thin one-pixel borders into dashes. Each fix shipped with a test that fails without it.
Filming found real problems in the product. Four course descriptions promised score gains, which we don't do. The capture agent flagged them, and they were rewritten at the source that same evening instead of being cropped out of frame.
And one shot is still waiting. When the assistant capture timed out, the agent had a choice between faking an answer and waiting for a real one. The rule made that choice for it: never stage, type, or paste a product's response. That shot waits until the real assistant answers.
My Voice, Disclosed
Five of the 30 posts are narrated, and the narrator is me, more or less. Post 7 is an excerpt from a film I made about how AccelaStudy AI was built. It's narrated by an ElevenLabs clone of my voice, reading lines from my own article. The video discloses that three times: on the end card ("Narrated in Charles Sieg's voice, synthesized with ElevenLabs."), in the post text, and in the MP4's metadata. It also carries burned-in captions for muted viewing, plus a caption file for the platforms that accept one.
By the Numbers
| Videos in the series | 60 (30 posts, 2 layouts each) |
| First batch | 5 posts, 10 videos, 15 to 31 seconds each |
| Live captures in the first batch | 75 (36 desktop, 39 phone) |
| Wall clock, from "do the first 5" to delivery | 2 hours 34 minutes |
| Agent time, across six parallel lanes | about 8.5 hours |
| Estimated hours for a senior engineer to do it by hand | 33.5 |
| My time at the keyboard | about 2 minutes |
| Tests guarding the kit and the copy | 386 |
| Render time for a 20-second post | about 2 minutes |
What the Job Became
In my article about building AccelaStudy AI, I wrote that my job was deciding what to build and making sure what came back was right. Making videos turned out to be the same job. I don't move keyframes. I decide what each post should say, I say no to the claims we can't back, and I look at the result on a phone before anyone else does. The agents do everything in between, and they do all of it in HTML.
The posts are going out now. You can see what they're about at test-prep.accelastudy.ai.
SAT® is a registered trademark of the College Board. PSAT/NMSQT® is a registered trademark of the College Board and the National Merit Scholarship Corporation. ACT® is a registered trademark of ACT, Inc. None of them is affiliated with, or endorses, AccelaStudy AI.
