The question I get from prospective clients more than any other in 2026: should we just use AI for this?

It’s a fair question. Sora, Veo, Runway, and the rest are demonstrably better than the same models were 18 months ago. The temptation to skip the production cycle entirely — and the cost — is real.

The honest answer requires separating two questions:

  1. Can generative AI produce the film I’m describing today?
  2. Will it be able to in 18 months?

The first answer is almost always no. The second is usually also no, but for different reasons. Let me unpack both.

What AI video models can actually do in 2026

The current state of the art (as of this writing) — Sora 2, Veo 3, Runway Gen-4 — produces:

  • Short clips, 8–15 seconds, in a single shot.
  • Convincing-looking environments and atmospheres. Cityscapes, weather, abstract motion.
  • Solid camera moves at moderate complexity (push-ins, dolly slides, slow zooms).
  • Reasonable lighting that doesn’t fight itself within a single shot.

What it doesn’t do well, today:

  • Hold a face consistent across more than 10 seconds. The protagonist of clip 1 has subtly different eyes/jawline in clip 2.
  • Render hands in close-ups. Still uncanny.
  • Match light and color across cuts. Two AI-generated shots edited together look like two different worlds.
  • Lip-sync to specific dialogue. Approximate, not synced.
  • Maintain product specificity. Generate “a coffee cup” — fine. Generate “this specific brand’s coffee cup with the logo legible” — still fails.
  • Operate from a treatment. You can prompt, but you can’t direct. The model will generate what it understands you to mean, not what you specifically want.

A 30-second brand film consists of 10–25 shots, each with consistent characters, lighting matched across cuts, brand-specific products, and dialogue. The current models can produce one shot of that. They cannot produce the whole film.

What this means for shoot-day production

For now, almost nothing. The work that happens on a film set — scripts, cinematography, talent direction, lighting design, location work — is exactly the work AI models cannot replicate.

This isn’t a “human creativity” argument; it’s a controllability argument. A director on set adjusts the talent’s expression by half a notch between takes. A model cannot be directed at that resolution because the model doesn’t actually know what “half a notch” means in the context of the specific film you’re making.

Generative AI in 2026 is most usefully thought of as a stock footage library that responds to natural language. Better than searching Getty for “city skyline at dusk.” Worse than shooting your own city skyline at dusk.

Where AI is actually disrupting production

Three places, all upstream or downstream of the camera:

1. Pre-production

Treatment writing, reference research, mood boarding, casting search, location lookups. These are bookkeeping tasks that don’t appear in the final film, and AI saves real time on them.

A treatment that used to take a director 8 hours to write now takes 3 — same quality, the model handles the prose-expansion work while the director makes the actual creative decisions. Reference research that used to take half a day takes 30 minutes.

2. Post-production

Transcription, scene logging, alternate-take generation (for patches, with talent consent), color matching as a starting point for human grades, automatic captioning.

This is the largest workflow shift of the last three years. It’s also boring — nobody’s making conference panels about transcription. But the cumulative time saved across an editorial pipeline is significant.

3. Stock and animation

Where a brief used to call for “animated lower-thirds” or “an abstract transition,” AI tools (Runway, Kling, native Resolve features) generate them faster than a motion designer working from scratch. The motion designer still finishes; the starting point is just better.

The real question isn’t “will AI replace filmmakers”

It’s: what does AI shift from human labor to human judgment?

Production work has always been a mix of:

  • Mechanical work (scheduling, transcription, editing rough cuts, basic color)
  • Craft work (cinematography, sound design, color finishing)
  • Judgment work (script, casting, edit-room storytelling, brand fit)

AI in 2026 is automating the mechanical layer and assisting the craft layer. It is not — yet, and probably not soon — replacing the judgment layer.

The filmmakers who are losing work to AI are the ones who were primarily doing mechanical work: stock footage shooters, basic explainer animators, social-cutdown editors. The filmmakers who are gaining work are the ones whose value was always in judgment.

What the AI-replaces-everything narrative gets wrong

Three errors:

1. Conflating “looks impressive in a demo” with “works in production.” Sora 2 demos are stunning. Sora 2 production work is rare because controllability is poor. The gap between demo and reliable workflow is years, not months.

2. Underestimating editorial. The hard part of filmmaking is the edit room — knowing what to cut, where to land a moment, when to break the rhythm. No current model does this. Most don’t even understand the question.

3. Pricing AI as if its output is a substitute for production. A 30-second AI-generated spot still requires a creative director, a script, sound design, music licensing, color, and approval cycles. The savings are in camera and crew (maybe 30–50% of a budget). The other half is the same.

What we tell brand-side clients in 2026

Three positions:

1. If your brief is “cheap and fast,” AI tools probably help. Ad-hoc social content, internal communication, low-stakes recurring video. The output is good enough; nobody’s evaluating the cinematography on a 6-second pre-roll.

2. If your brief is “memorable and brand-defining,” it doesn’t help. A hero brand film, a documentary, a customer success piece that closes deals — these depend on the judgment layer that AI doesn’t reach.

3. The middle is where it gets interesting. Hybrid productions where AI handles 10–20% of the deliverable (specific cutaways, transitions, pre-pro automation) and human filmmakers handle the rest. We do this on most brand projects in 2026 already; clients usually don’t notice unless we point it out.

The longer arc

In 2030 or 2032, generative video will be controllable enough to produce real production work end-to-end. When that happens, the brand films that compete will be the ones with the strongest scripts, the most considered art direction, and the clearest point of view — because the picture itself becomes commodity.

In other words, when AI can produce any image cheaply, the differentiator becomes which image is worth producing. Which has always been the actual hard part.

We’re betting the studios that focused on craft and judgment outlast the ones that competed on production cost. The technology cycle accelerates that bet, it doesn’t break it.

If you’re scoping a project and weighing AI options against traditional production, send us the brief. We’ll show you which parts of your project the current tools handle well, which parts they don’t, and where the line is for what you actually need.