How Do AI Creators Actually Make Their Videos?
← Back to Creator Stories

How Do AI Creators Actually Make Their Videos?

People keep typing the same question into search: how does this creator actually make their videos? We've been asking them directly. Here's what five of the processes really look like — scripts, storyboards, hundreds of discarded generations, and a lot of editing.

Art by The Archive In Between, from their film "The Starfall Initiative" — find them on YouTube at @thearchiveinbetween

Somebody watches a Gossip Goblin short or a cursejourney reel, and the first thing they type into Google is some version of the same question: how does he actually make these? I know because I typed it myself, more than once, before I started asking the creators directly.

The assumption behind the question is usually "there's a magic prompt." There never is.

Every process we've deconstructed starts somewhere old-fashioned — a script, a storyboard, a world that existed before the tools did — and then runs through a stack of tools handing off to each other, with a human throwing away most of what the machines produce. The tools change month to month. The shape doesn't.

Here are five of the processes we've pulled apart, each with a link to the full breakdown.

How does Gossip Goblin make his videos?

Zack London's answer is the bluntest one we've collected: "Every video starts with a script." Not a prompt, not a mood board — a script. The worlds he builds, from cyborg aristocrats to entire goblin civilizations, hang off the writing, and everything downstream exists to serve it.

Then comes the part that never shows up in the finished video. For one set of characters he ran "probably 400 prompts / 1600 images" just to get the faces right. Environments are worse — he's said flatly that "this part is not enjoyable," describing generating a mountain of background shots just to get something he could force his characters into. Animation brings its own grind: lip sync is "a colossal headache," camera movement is manual, and a single 90-second dialogue scene took "about 150 generations."

The stack is half a dozen specialists handing off: Midjourney for characters and costumes, Seedream to blend characters into environments, Veo 3 for dialogue-heavy scenes, Runway and HeyGen for performance and lip-sync passes, ElevenLabs for voices, and CapCut to cut it — a choice he jokes about, but it assembles a dozen moving parts fine. End to end, one piece takes him 12–14 hours.

The full breakdown is in the work hidden inside Gossip Goblin's worlds.

"Pomegranate," Gossip Goblin's newest short film — @Gossip.Goblin on YouTube

How does Neural Viz make the Monoverse?

Josh Kerrigan was a filmmaker for over a decade before any of this — film school, years of LA production work, a TV pilot he's said he sold before going full-time on Neural Viz in early 2025. That background shows up in the design of the world itself: he studied what the models are bad at and built a show that hides it. AI struggles with realistic humans, so his characters are bulbous cartoon aliens your brain never expects to look real. Clean 4K makes rendering artifacts scream, so everything is graded like a worn-out 20th-century broadcast — VHS noise, soft focus, a tape copied one too many times.

The workflow underneath is a writers' room, not a prompt box. He writes a full script — slug lines, blocking, the whole format — then storyboards every shot and generates a still per panel, mostly in Midjourney, holding lighting and sight lines consistent so the cuts don't fall apart. The signature move: he acts the scenes out in front of his webcam, and Runway's Act-One maps his real performance — the timing, the head turns, the delivery — onto the alien characters. Hedra handles lip-sync, ElevenLabs does voices (sometimes layered over his own), and Premiere cuts it like any other edit.

His own numbers: about twelve hours and roughly a hundred bucks a month in subscriptions for a two-to-three-minute episode. And he's clear about where the result actually comes from: "Everything I do within these tools is a skill set that's been built up over a decade plus."

The full breakdown is in The Showrunner: how Neural Viz makes an entire TV universe.

"Human Hunters," from the Monoverse — @NeuralViz on YouTube

How does Kelly Boesch make her AI films?

Boesch spent 17 years at IMAX — graphic design, film production, marketing, the development decks that convince studios to fund projects. When she picked up AI generation three years ago, she arrived already knowing what a cinematic image is supposed to feel like. That training is why her method runs opposite to how most people approach AI video.

Most people go text-first: describe the scene, generate, hope. Boesch works image-first. She generates a still in Midjourney, gets the composition exactly right, then animates from that image using Runway, Pika, Higgsfield, or Veo 3 depending on what the scene needs — each tool behaves differently, and she moves between them. Image-to-video keeps the art direction in her hands where text-to-video would hand it to chance.

Her line on where the human work lives: "You can't just let the AI do the work — the editing, color correction, and sound mixing are where the human touch turns a clip into a story." The same pipeline thinking extends past video — her debut album Fairytale came out on a real label last year, lyrics written by her, music produced through Suno under her direction.

The full breakdown is in how Kelly Boesch built a studio of one.

"Not Made For The Cage," a 4K AI film by Kelly Boesch — @kellyeld2323 on YouTube, kellyboesch.com

How does cursejourney get that found-footage look?

Mike Chhay has been making cursejourney since August 2023: AI horror styled to look like photographs you weren't supposed to come across — demons, gods, and monsters rendered sepia-toned and grainy, like documentation instead of generation. The insight is that the "close but wrong" quality most people fight to remove is exactly what horror can use. His flagship series is literally called "photos i found in the basement."

He publishes the pipeline on a public tools page, and it's seven tools deep for a single short. Midjourney generates the base image — "the AI image generator that has created the majority of my cursed old photo series," in his words. Before anything moves, Magnific upscales the still so the animation starts from a high-resolution frame — that's what keeps faces and details consistent. Photoshop handles cropping, color, and filters. Then Kling animates most shots — "the tool you want to use for the majority of the time" — with Seedance held back for "epic transformations and action scenes." Premiere assembles, Suno scores, ElevenLabs adds voice and sound, and Topaz upscales the finish to 4K — the same class of tool used to restore old blurry footage, which fits work meant to look dug up rather than made.

The lineup keeps shifting (his 2024 pieces ran on Luma before Kling took over), and the one thing he keeps loose is the exact recipe for the aged sepia-grain look. That lives in the edit and the concept. It's worked well beyond the horror crowd, too — his animated "cursejourney cat" was selected for AIGA Arizona's Best of 2025, with over 140 million views on Instagram alone.

The full breakdown is in cursejourney makes AI horror that looks like found footage.

"photos i found in the basement" — @cursejourney on YouTube, cursejourney.com

How does The Archive In Between publish four stories a week?

The Archive is a sci-fi, horror, and fantasy anthology — human-written stories paired with AI imagery, framed as transmissions from a library that sits between worlds. It runs across TikTok, Instagram, YouTube, Substack, Patreon, Discord, and Spotify, publishing around four stories a week, one of them long-form. That's a lot of finished narrative for a project run by one person — and it's the Curator's full-time income now, funded by readers through Patreon and book sales rather than ads. Their art is also the cover image on this piece.

The labor split is the sharpest we've documented. Every story is human-written — no model drafts a sentence. "ChatGPT is the best thesaurus in the entire world. That is the extent of my use of AI in the writing." The imagery is the opposite: generated in volume, then ruthlessly culled. They describe that side of the work as sifting — picking the passable out of a mountain of output — and they refuse the word "artist" for it entirely: "The artistry comes across in the writing part of it."

The perspective comes from experience on both sides. The Curator worked about four years as a professional illustrator and designer before generative tools arrived, and credits the project's traction to "50% luck in timing and 50% my perspective." The words are authored, the images are selected, and they keep that line deliberately unblurred.

The full breakdown is in the Curator behind The Archive In Between won't call the AI part art.

"The Starfall Initiative" — @thearchiveinbetween on YouTube

The pattern, if you're looking for one

Read these processes side by side and the same shape keeps showing up. The work starts before the AI does — a script, a storyboard, years of writing, a trained eye. The generation step is a volume game everyone describes with mild exhaustion: 400 prompts, 150 generations, a mountain sifted for the passable. And the finish is old-fashioned post-production — editing, color, sound — which more than one of them names as the place the human touch actually lives.

The tool names in these breakdowns will be stale within a year. I update them when the creators do.

These five are a sample, not the whole shelf — every process we've pulled apart lives in the Creator Stories collection, and there are more coming. If there's a creator whose process you want deconstructed, I want to hear about it.