What I've Learned About Gen AI in Video Production
- Amit Apte
- Jul 24
- 5 min read
Over the past couple months, I’ve been diving deep into Generative AI, LLMs, and creative tech. The AI landscape moves insanely fast—literally changing day by day, week to week—so keeping up feels like a full-time job. In this article, I’m looking at these tools purely through a creative lens, rather than a business or engineering one.
I’ll be exploring the following areas as they relate to content creation:
Choosing the right tools
Choosing the right models
Storyboarding & pre-production
Editing: Control vs. Automation
Costs & practical realities
When I first started this journey, I wasn’t starting from ground zero. Having spent 18+ years at Adobe, I’ve been around the Creative Cloud suite for a long time and had a solid foundation in traditional content creation workflows. Historically, producing video content—especially animated or high-concept work—meant hiring actors, production crews, and animators. Without an established go-to team, production costs quickly became prohibitive, leading to scaled-back ambitions or abandoned projects. Even as a jack-of-all-trades (music producer, videographer, editor), I often felt constrained in how much high-quality content I could realistically churn out solo.
One project I’ve always wanted to tackle is an animated series. But quotes from animators ranged from $1,000 to $25,000 per minute depending on style and complexity. Even on the low end, a 20-minute episode would run $20,000 minimum.
Enter Generative AI. Everything has changed.
Choosing The Right Tools
When building an AI workflow for animated or stylized video, you need to account for four core layers: script/storyboarding, visual imagery, music/score, and editing.
You essentially have two strategic choices:
The All-in-One Route: All-in-one platforms like RunwayML, Higgsfield, and Kling offer image/video generation, audio generation, voiceover, and light editing in a single web interface. They are cost-effective, frictionless, and ideal if you don't come from a post-production background.
The Piecemeal (Hybrid) Route: If you want higher fidelity and deeper creative control, you can pair AI generation tools with traditional editors like Adobe Premiere, DaVinci Resolve, or Final Cut (or accessible suites like Canva and Adobe Express). Going this route gives you fine-grained timeline control, though it requires bridging assets across multiple tools yourself.
You also have to decide on your interface preference: Do you prefer a visual, node- or layer-based UI for tactile editing, or do you prefer orchestrating your creative workflow through an LLM like Gemini, Claude, or ChatGPT via natural language prompting?
Choosing The Right Models
There is a huge variety of generative AI models available today. A new one seems to launch every few weeks, but a handful consistently stay at the forefront.
Image Models
Options like GPT Image, Midjourney, Flux, and NanoBanana lead the pack. Midjourney dominated the aesthetic and lifestyle space for a long time, but competitor models have caught up significantly.
All of them still occasionally hallucinate (extra fingers, asymmetrical anatomy), and they tend to gravitate toward default "archetypes." For example, if you prompt "good-looking Indian man," you’ll get very similar-looking faces across generations. Even with highly specific prompts ("man in his 40s, salt-and-pepper hair, fit with a subtle belly, mustache, no beard"), characters across different user accounts end up sharing a recognizable "AI aesthetic."
Takeaway: Generic prompting works fine for background b-roll or standard marketing assets. But if you are building unique IP or recognizable characters, start with actual reference photos or custom character illustrations, then feed those into image models for pose and style manipulation.
Video Models
Top performers include Google Veo, Kling, Seedance, and Runway Gen-4.
However, your video output is only as good as your reference image. Pure text-to-video prompts rarely deliver scene-to-scene coherence or character consistency. Image-to-video (using a strong reference image to lock in character, lighting, and composition) is essential.
Current Landscape (as of mid-2026): Google Veo and Omni offer some of the most cinematic and photorealistic motion models available, though their higher tier pricing yields fewer credits per tier. Kling and Seedance are powerhouses for larger projects requiring volume at scale. Runway Gen-4 remains a strong economical option, though it requires tighter prompting to preserve character stability.
Audio Models
AI audio spans speech, music, and sound effects:
ElevenLabs leads in voice cloning, text-to-speech, and custom character voice design.
Whisper remains a benchmark for speech recognition and audio-to-text processing.
Suno leads in full AI music generation (a topic that deserves its own standalone article).
Bottom Line: There is no single "correct" stack. When starting pre-production, consider signing up for an aggregator or platform suite (like Runway, Higgsfield, or Adobe Firefly) to test multiple models against your specific visual style. Once you find what clicks, commit for the duration of that project.
Storyboarding & Pre-Production
Once you’ve locked in your stack, storyboarding becomes your most critical phase. Pre-visualization saves massive amounts of rendering time and token spend down the road.
In AI video production, storyboarding serves two functions:
The Visual Roadmap: Structuring your narrative sequence.
The Prompt Generator: Each tile on your storyboard becomes the direct input for your Image-to-Video generation.
Uploading your storyboard frame as a visual anchor—combined with a detailed shot prompt specifying motion, camera angles, and action—dramatically increases model accuracy.
The same applies to audio storyboarding: Record rough voice notes or sound effect mimics directly on your phone, upload them, and use them as temporal or stylistic reference tracks to drive your AI audio generators.
For storyboarding tools, native environments inside Runway or Higgsfield work great, as do collaborative canvases like Miro, Figma, or Adobe Firefly Boards.
Editing: Control vs. Automation
Once your assets are generated, it's time to assemble the final edit.
Automated AI Editing: Platforms exist where you upload raw clips and an algorithm automatically cuts them to a template or music beat. This works well for short social reels, product promos, or visualizers where complex narrative pacing isn't required.
Timeline-Based Manual Editing: For storytelling with specific emotional beats, manual timeline editing remains essential. Prompting an LLM via chat to edit complex video clips frame-by-frame is still far too cumbersome and imprecise compared to standard video editing software.
Depending on your comfort level, consumer tools like Canva, Adobe Express, or iMovie offer speed, while professional NLEs like Adobe Premiere, DaVinci Resolve, or Final Cut offer uncompromised precision.
Cost Breakdown: A Real-World Example
To put this into perspective, here is the real-world breakdown from a recent project of mine: promoting my new song, "Dudes Night Out," with a 1-minute social media reel.
Traditional Production Estimate: A quick estimate for producing a 1-minute professional live-action or traditional animated reel (actors, crew, location fees, animated assets, editing) ranges from $3,000 on the low end to $20,000+.
My AI Workflow:
Designed core character models and environments using Gemini Pro and NanoBanana.
Used those visual anchors inside RunwayML running the Wan 2.6 video model for clip generation.
Generated ~40 five-second 1080p clips, selecting the best 25 for the final edit.
Assembled and cut the final track inside Adobe Premiere.
Total Generation Cost: Approximately 2,500 tokens / generation credits totaling ~$50 in model compute costs. Since I handled the edit in Premiere, editing costs were covered under my existing subscription.
Conclusion & Limitations
Generative AI offered massive savings in both time and budget for this project, allowing me to produce high-concept promo material that previously would have been cost-prohibitive.
That said, the technology comes with clear limitations you need to manage:
Homogenization (Lack of Originality): Because everyone has access to the same public models and default prompts, AI-generated content can easily start looking identical. Crafting a unique visual style requires extra intent in pre-production.
Loss of Granular Control: Models don't always follow every detail of your prompt. You often have to adapt your narrative or shot choices based on what the model outputs.
Character Consistency: Keeping a character's features identical across dozens of generated clips takes work. Without strict image prompts or custom model training (like LoRAs, which I’ll cover in a future piece), your character's likeness can drift over time.
Copyright & Rights Management: Public foundation models currently carry ambiguity around IP ownership, and digital watermarking / C2PA metadata clearly marks raw output files as synthetic media. (In an upcoming post, I'll dive into training local, custom AI models on proprietary IP to solve this).
As always, keep creating!

