11 SEPT 2026AI · Process · Filmmaking

Building an Empty NYC

Go to project

I set out to build a small cinematic using 100% AI-generated media, with one hard goal: character consistency and photorealism blended with the surrealism that AI makes possible. I wanted to show that AI-generated media can be pushed through real editing and filmmaking decisions and come out the other side as a cinematic. Not a narrative, but something that holds the feeling behind the visuals. What I was after was the quiet contentment of moving through a surreal, isolated city, and how the character feels alone in a place that should be full of people.

Using Higgsfield, I locked the environment first: barren downtown Manhattan, mid-afternoon, 3pm, no people, no moving cars. Then the character, me. With Soul 2.0 I created a character ID of myself, generated a final character sheet for the models to reference in every shot, and planned a first shot to lock the environment and the character together. That first frame set the rules for everything that followed.

The full Higgsfield canvas: location references, the character sheet, and generated stills feeding into the shot list
Aerial reference of downtown ManhattanA lone figure at an empty midtown intersection, shot straight down
The six-panel character sheet used as the identity reference in every shot

Consistency in world building

The empty-city idea only works if it is airtight. No pedestrians, no distant figures, no parked or moving cars, anywhere, ever. Every shot locked to 3pm so nothing feels stitched together from different days. Haze dissolving the far background the way real atmospheric perspective softens a skyline a mile out. Small rules, but they are the difference between "AI video" and a place that feels like it genuinely exists and is simply, eerily empty. Writing the environment out in detail up front is what holds the world together across every generation.

Every generated frame had to come out flat. No baked colour grade, no film grain, no stylised look from the model. I did all of that myself afterwards in After Effects: one grayscale LUT and one grain pass across every shot, so the whole reel reads as a single continuous piece of footage. Higgsfield kept offering its own presets mid-generation, and I had to decline them every time or they would bake a grade in before I could touch it.

The images were locked in over multiple iterations with Nano Banana Pro. Despite being one of the flagship models out there, it stayed cost-efficient.

Graded still: a walk signal lit against a blown-out empty street
Higgsfield canvas: the character reference feeding a fan of prompts and generated stills

Vocabulary for streamlined communication

Getting a still to look right and getting it to hold up across a full video generation turned out to be two different disciplines. Stills settled into a five-part structure: composition, the character exactly as he appears on the sheet, the environment, the anchoring landmarks, and a closing pass describing the actual film-camera physics I wanted.

Video generation needed far more scaffolding. A ten-block prompt covering framing, identity lock, how figures relate across the frame, the movement itself, the final frame, environment, sound, and a realism pass tying it back to the same flat look as the stills. It is a lot of overhead for a few seconds of footage, but it is the only way I found to keep identity, wardrobe, and camera physics from drifting the moment things started to move.

Using Claude, I wrote the environment and character instructions out as reusable prompts. It also helped me organise and pull the right references quickly. With the filmmaking and prompting language in place, Seedance 2.5 and the other models produced usable generations while keeping the process quick and far less iterative.

The subway shot on the canvas: reference plate, prompt blocks, generated still, then the video generations

Walking through the city, then bending reality

The early shots just established the character: leaning on a coffee shop wall mid-sip, a brisk walk shot entirely from behind, climbing out of a subway stairwell into the empty street. Once that held, the landmarks came, and so did the surrealism. A perched crouch on a Chrysler Building eagle gargoyle, a walk down an empty Brooklyn Bridge, a body floating above the city, all of it carrying the same isolation.

The surreal shots are where the project actually says something. A few ideas got scrapped along the way: standing at mid-rib height against the Flatiron Building, scaled to the size of the block, or alone at centre court in Madison Square Garden. The duplicates that chase down the Brooklyn Bridge and through the subway added to the awe.

Graded still: the character walking away down an empty downtown street, shot from behind
Graded still: duplicates of the character moving through an empty subway car

What I proved to myself

Consistency and photorealism do not have to fight the surrealism. A fully AI-generated pipeline can hold one believable identity across a dozen locations and still make room for the impossible: a dystopian world, self-multiplied crowds, a body floating over the skyline.

For all of the SFX on the inserts and shots I used ElevenLabs. Iterating there was fast and cheap, and the turnaround was great.

ElevenLabs sound effect history: footsteps, doors and ambience generated per shot

Finally I brought everything together by cutting to the bass-heavy "FATHER" by Ye and Travis Scott. The audio laid the foundation for the shots I generated against it. I also built transitions that vary between shots using one-framers: invert, directional blur, overlay. At around 30 seconds, this was my first complete piece built this way.

The After Effects timeline: shot layers, one-framer transitions, and the generated SFX stack
The full Higgsfield canvas