Post · mixed
Eight Hours to Build a Video Agent. Then I Edited Myself Out.
My first video agent experiment at tokens& NYC: real footage, English subtitles, motion, sound, and a lesson in context—with an unfortunate submission-day plot twist.

Eight hours to build my first video agent. One very human mistake waiting at the finish line.
I already build agents at work, mostly around enterprise workflows. I have built more outside work, too, and lately I have been working on voice agents. So I signed up for a hackathon to make myself try something unfamiliar: an agent that edits video.
How hard could it be? Famous last words, now with subtitles.
Eight hours, a pile of footage, and an arcade
The event was Real-Time Video Agents Hack — NYC, presented by tokens& as part of the VAST Builders Challenge. The brief was to find new value in stored video. Doors opened at 8:30 a.m.; demos started at 4:30 p.m. An eight-hour window to turn an idea into something worth showing.
I brought footage from my visit to the Nbox AI data center. Servers, GPUs, the usual infrastructure story. And then, unexpectedly, arcade games.
That surprise became the brief:
An unexpected AI data center visit. An arcade hiding inside. Make it a fun vertical video.
The source footage included me speaking Chinese. The hackathon was in English. The edit needed English subtitles, animated text, sound effects, and enough personality to make the arcade reveal land. I wanted short cuts in both 15- and 30-second versions.
An ordinary instruction, until you try to teach an agent what “fun” means.
The work between the prompt and the MP4
I started with Media Spotlight and the story context already attached to the footage. Finding a file is useful. Finding the moment that belongs in the story is where editing starts.
Then came the decisions: which source clip to pull, where to trim it, how long to hold a shot, when to bring in a caption, and where an arcade sound would actually help the joke. Remotion gave me a way to put the cuts, motion, text, and sound into a renderable composition.
The demo shows the workflow and the result:
For a short hackathon presentation, I could show Media Spotlight and the final edit live. The recorded editing process fills in the middle: pulling real footage, selecting moments, and assembling the cut. There is a surprising amount of work between “make it fun” and an MP4.
The rough architecture: two agents, one story
This walkthrough shows the idea behind the prototype:
The flow has two agents, with a Story connecting them.
- The Context Agent collects clues. It reads scenes and subjects from the picture, words and timestamps from the audio, and file and camera information. That becomes a searchable index of source moments.
- My context gives the clues a direction. “Show the human side of the data center” changes which moments matter. The indexed footage and that intent become a Story: a hook, a discovery, a payoff, and the source clips to support them. The deck calls this a “Story cartridge.”
- The Video Agent works from that Story. I describe the edit in chat—length, mood, emphasis—and it uses the story clues to select moments and plan the cut. The same Story can support different versions.
- The edit becomes a rendered video. For these two shorts, the selected real footage, English captions, animated text, and sound are assembled through Remotion.
It is a rough architecture, and the deck illustrates the handoff. The finished shorts below show what I made with that workflow during the hackathon.
The biggest lesson was context
This was a deeper learning experience than I expected. Video brings together language, timing, image, sound, and taste. A caption can be correct and still arrive too late. An effect can look great and still distract from the moment.
Fast compute and good generation and editing tools help enormously. They shorten the feedback loop: try an idea, see it, change it, try again. GPU access makes that loop much more practical.
But the useful edit depends on the context you give the agent.
It needs to understand why this footage matters, who the audience is, and what the viewer should feel.
What amazed me was finding a surprising story inside the whole pile of footage, with my context guiding the choice. There were server rooms, a strange noise, my “is this place haunted?” reaction, and then the arcade reveal. Together, those moments became a little mystery with a funny answer.
A chronological talking-head recap would have flattened that into “I arrived, I looked around, I saw some machines.” I wanted the surprise: an AI data center with an unexpected human side. The agent needed enough context to recognize which moments carried that story, then shape the setup, reveal, and payoff.
Without that context, it is easy to get a technically polished pile of captions and transitions. Familiar talking-head AI slop, just rendered faster.
I came away wanting to get better at giving an agent the context to make an editing decision. The prompt, the selected footage, and the story behind it all matter.
The final cut was me
Then came submission.
I had forgotten to use the required sponsor VM for my project. I was disqualified.
That one is on me. I had spent the day thinking about the video's context and missed a rather important piece of the hackathon's context: the submission rules.
The agent edited the footage. I edited myself out of the competition.
Still, I left with a demo, my first video agent experiment, and a much clearer sense of how much judgment goes into editing. A very good learning day. A very questionable eligibility strategy.
Next time: read the brief, read the rules, then hit render.
The two finished cuts
Here is the story the edit found in the footage. Both cuts keep my original Chinese speech, with English subtitles, animated text, and sound.
15 seconds — the quick surprise
A short setup, the arcade reveal, and the reaction that brings the joke home.
30 seconds — the full side quest
A little more room for the strange noise, the mystery, and the games behind it.

