The field guide to professional AI video.
Four short chapters on making footage that looks directed instead of generated — and on spending money like a producer while you do it. No sign-up needed to read; everything here works in any tool, and best in this one. Film work continues in the AI video for filmmakers and previz guide.
Seedance 2.5 is live in Saga.
Saga currently offers tested 4 to 30 second generation at 480p and 720p, with image, video and audio references. Advertised 4K output stays off until the provider path is verified.
Cast the model like a producer.
Every model in the studio has a personality and a price. The single most expensive habit in AI video is running every idea through the most expensive model. Cast per job instead:
| Image model | Use it for |
|---|---|
| Nano Bananacheapest in the roster | Exploration. Blocking compositions, testing ideas, burning through variations for pennies. |
| GPT Image Mini | Fast, cheap drafts with a different eye than the Banana family — good second opinion. |
| GPT-5.4 Image 21K · up to 6 references | Generation and source-aware editing when prompt fidelity and reference control matter more than final resolution. |
| Nano Banana 2 | The mid step: stronger prompt adherence when a draft is close but not holding detail. |
| GPT-5 Image | Character and mood. Strong on people, wardrobe and dramatic light. |
| Nano Banana Prothe only true 4K image model here | The finish. Final frames, text in image, and stills you intend to animate. |
| Video model | Use it for |
|---|---|
| Seedance 2.0 Fast480p · 720p — 4–15 s | Motion drafts. Test whether the shot works at all before paying finish rates. |
| Kling 3.0 Standard720p — 3–15 s | Reliable general-purpose motion at a working price. |
| Kling 3.0 Pro720p — 3–15 s | Cleaner physics and camera discipline when Standard almost gets there. |
| HappyHorse 1.1720p · 1080p | A different motion signature — worth an audition on stylized work. |
| Seedance 2.0480p → 4K — 4–15 s, audio | The finish: highest ceiling in the roster, longest durations, sound on tap. |
| Sora 2 Pro4 / 8 / 12 / 16 / 20 s | Long single takes and complex staging, in fixed clip lengths. |
Write light, not adjectives.
Quality words — cinematic, 4K, masterpiece, ultra-realistic — are the things these models learned to fake first. What they can't fake is a lighting plan. Describe the shot the way a director of photography would brief it:
- Name the light. Where the key comes from, what fills the shadows, whether a rim separates subject from background. Warm or cool, hard or soft.
- Pick two imperfections. Flyaway hair, uneven skin sheen, dust in a light shaft, a scuffed prop. Perfect frames read as renders; two deliberate flaws read as photography.
- Give eyes a catchlight on any prominent face, and keep the subject tonally separated from the background — that's most of what "professional" means in a portrait.
- Describe materials, not moods. "Salt-dulled chain mail" beats "epic armor". Models render nouns and physics better than vibes.
- Say what is there, not what isn't. Image models handle negation poorly — phrase everything positively.
epic cinematic 4K masterpiece of a warrior on a cliff, ultra realistic, dramatic lighting, best quality
A weathered warrior stands on a basalt cliff at golden hour. Low sun as the key light rakes across salt-dulled chain mail; cool sky fill softens the shadows; a rim of backlight separates his silhouette from the sea haze. Wind lifts loose strands of hair, and a faint catchlight sits in his eyes.
Same idea, radically different footage.
Stills first. Then motion.
Professional AI video is mostly still-image work. A video render can cost as much as twenty drafts of the frame it starts from — so make every creative decision in the cheap medium first:
- Explore in cheap stills. Run the idea through Nano Banana or GPT Image Mini until the composition, wardrobe and light are right. Each attempt costs cents.
- Finish the keeper. Re-render the winning frame in Nano Banana Pro at final resolution. This is the frame your video inherits everything from.
- Compose for motion before you animate. Layer the frame — foreground, midground, background — leave negative space in the direction things will move, and keep hands and faces away from the frame edges, where video models degrade first.
- Animate a draft, then the finish. Drive the still through Seedance Fast at 480p to prove the motion, then pay finish rates once, not five times.
Fetching live draft-vs-finish quotes…
Quoted live from the studio's own pricing call just now — the draft habit is the affordability model.
Inspect a real Seedance 2.0 source image, motion prompt and resulting clip →
Direct the motion. Don't describe a montage.
Video prompts fail differently from image prompts: the model has to spend your seconds. Give it a single take it can actually execute:
- One camera move, named. A slow push-in, a drift left, a rise. Two moves in one clip is how you get neither.
- Describe change over time. What enters, what settles, what the light does between the first frame and the last. A video prompt is a sentence with a beginning and an end, not a list of attributes.
- Keep clips short and cut like an editor. Two intentional 5-second shots beat one 15-second wander — and cost about the same. Sequences are made in the edit, not in one render.
- Add audio when it earns its cost. Sound is priced into the quote on models that support it — flip it on for the finish, not the drafts.
When you'd rather write one plain sentence, the studio's prompt enhancement expands it with this craft — light, motion grammar, composition — before it reaches the model, and stores the expanded prompt with the render so you can read exactly what was sent and learn from it.