Making NYC Systems videos legible
I am going to see if I can automate a lot of the NYC Systems video production work with a remote Grok agent in a Railway Sandbox.
Wish me luck.
Okay. It worked. Kind of messily at first, then cleanly. This is the write-up.
Why this even matters
We run NYC Systems. Speakers come in, talk about real systems, and leave us with two awkward truth-sources:
- A room camera (Lumix S9, beautiful, also 35–50 GB of 4K HEVC that will cook your MacBook Air if you disrespect it)
- A Meet / screen recording of the deck (clean slides, often no useful audio, longer than the talk because the whole night is one recording)
What we want on YouTube is simple:
- Clean slides on the NY Systems template
- Speaker in frame
- Title + speaker in Times New Roman Italic
- Date under New York Systems Talks
- Trail of Bits as venue host, because they feed us
What we had was me scrubbing two timelines, aligning timecodes by eye, cropping people by hand, and waiting for renders while my laptop battery cried.
We must make the insights of our speakers legible to everyone. Sitting on raw masters for weeks is the opposite of that.
The constraint that shaped everything
The slide stream has no audio worth syncing on.
So the sync problem is not “cross-correlate the waveforms and go home.” It’s:
When the deck changes on Meet, the TV in the room changes. Watch the glass.
That sounds obvious. It is also where I spend a significant amount of time every edit. So I went on a tangent — YOLO-ish person crop, TV rectangle, fingerprint the Meet slides, align change events on the room TV ROI to Meet slide advances — then back to compositing.
No audio. Visual only. Offset lives in a sync.json.
What we actually built
I split the problem into two tools that match how I already think about the pipeline:
1. derive_pair — get two mezzanines that share a clock
room master + Meet download
↓
auto TV ROI + speaker crop (with overrides)
visual slide-change alignment
↓
slides_mezz.mp4 (silent, clean deck)
speaker_mezz.mp4 (person + room audio)
sync.json
Dogfooding this on Olivia Gallucci’s talk ($ ./ai-assisted -macOS-vr):
- Room mezz: ~36 min
- Meet file: ~96 min (whole night)
- Forced offset after change-event align: room t=0 ↔ Meet t=113s
- Verify by scrubbing: same slide on the TV and in the mezz at the same media time. Good enough to ship.
I lean on Grok 4.5 a lot here because I personally like fast, intelligent models that I drive and can verify, over launching 5–6 parallel agents and watching quality fall off a cliff. One agent, tight loop, check the frames.
2. compose — the webapp, but local Python
We also stood up a little Meet Composite webapp on Railway (bucket multipart, layout UI, render from S3). That was useful for the dogfood loop and the branded layout constants.
But the webapp is not magic. It’s orchestration around one FFmpeg graph:
background PNG
+ slides pane (cover into 1454×817 @ 408,90)
+ speaker pane (cover into 262×262 @ 38,344)
+ captions
→ 1920×1080 H.264
So I ported that graph into python -m derive_pair.compose. Same layout JSON the UI uses (NYST_LAYOUT). No S3 required when the files are already on disk. Laptop still does the heavy encode — M5 VideoToolbox for mezz prep, libx264 for the final when drawtext isn’t in the Homebrew build (we bake captions as a Times New Roman Italic PNG overlay instead).
Design language, finally correct:
- Top left: New York Systems Talks / June 20 (we paint out the baked-in “February 2024” on the template and put the real date in Times New Roman Italic)
- Bottom under the deck: title and speaker, Times New Roman Italic, no “Title:” chrome
That is the look. If the type is wrong, the whole thing feels like a student project.
Rough edges (I already found like 14)
Being honest, because the real ones are in the Slack Connect already:
- Auto TV detect will happily lock onto a chair if you let it. Venue priors + hand ROIs still win for this room.
- Room TV “slide changes” include people walking past glass. Strong-MAD filtering helps; it’s not perfect.
- Speaker crop is “good enough for PiP,” not a cinema reframe. Sometimes you get half a person and a lot of TV.
- Homebrew ffmpeg without freetype means captions are a PIL overlay, not drawtext. Fine. Document it.
- 50 GB masters still want a mezz first. In a perfect world the webapp would HandBrake them for mezz. Today:
encode_hw.shon the M5, or compress before upload. - Cloud agents are great for battery life. I totally forgot how much it helps until the Mac stopped trying to leave low-earth orbit.
None of that is a reason to go back to manual Final Cut for every talk. It’s a reason to keep dogfooding.
The flow I settled on
- Dump the S9 cards, grab the Meet export.
- Mezz the room cam on the M5 (VideoToolbox — don’t cook the chassis).
python -m derive_pairwith venue ROIs + checksync.json/ overlay stills.- Nudge
--offsetif the title slide is off by a breath. python -m derive_pair.compose --date "June 20" --title "…" --speaker-name "…".- Ship.
Renders aren’t burning my legs. Speakers get out faster. The archive stops being a graveyard of unreadable 4K.
What’s next
- Tighten auto speaker crop so PiP is actually a face, not a plant.
- Make offset confidence something I’d trust without scrubbing (not there yet).
- Maybe the webapp stays as the “share a link and tweak layout” surface; Python stays the production path.
- Expect to see more NYC Systems videos just… out. Near instantly.
We must make the insights of our speakers legible to everyone.
That’s the whole product.
Co-host of nycsystems.xyz. Work on computers and the clouds that hold them. If you spoke at a night and your video is still in purgatory, this is me fixing that.