Saga Studio

The field guide to professional AI video.

Four short chapters on making footage that looks directed instead of generated — and on spending money like a producer while you do it. No sign-up needed to read; everything here works in any tool, and best in this one.

Chapter 01 casting A model is a camera package, not a brand allegiance. Producers rent per job.

Cast the model like a producer.

Every model in the studio has a personality and a price. The single most expensive habit in AI video is running every idea through the most expensive model. Cast per job instead:

Image modelUse it for
Nano Bananacheapest in the rosterExploration. Blocking compositions, testing ideas, burning through variations for pennies.
GPT Image MiniFast, cheap drafts with a different eye than the Banana family — good second opinion.
Nano Banana 2The mid step: stronger prompt adherence when a draft is close but not holding detail.
GPT-5 ImageCharacter and mood. Strong on people, wardrobe and dramatic light.
Nano Banana Prothe only true 4K image model hereThe finish. Final frames, text in image, and stills you intend to animate.
Video modelUse it for
Seedance 2.0 Fast480p · 720p — 4–15 sMotion drafts. Test whether the shot works at all before paying finish rates.
Kling 3.0 Standard720p — 3–15 sReliable general-purpose motion at a working price.
Kling 3.0 Pro720p — 3–15 sCleaner physics and camera discipline when Standard almost gets there.
HappyHorse 1.1720p · 1080pA different motion signature — worth an audition on stylized work.
Seedance 2.0480p → 4K — 4–15 s, audioThe finish: highest ceiling in the roster, longest durations, sound on tap.
Sora 2 Pro4 / 8 / 12 / 16 / 20 sLong single takes and complex staging, in fixed clip lengths.
Exact per-render prices are quoted live on the rate board — the same call the studio makes before it charges you.
Chapter 02 the prompt Perfection is the strongest AI tell. Two flaws, placed on purpose, read as reality.

Write light, not adjectives.

Quality words — cinematic, 4K, masterpiece, ultra-realistic — are the things these models learned to fake first. What they can't fake is a lighting plan. Describe the shot the way a director of photography would brief it:

  • Name the light. Where the key comes from, what fills the shadows, whether a rim separates subject from background. Warm or cool, hard or soft.
  • Pick two imperfections. Flyaway hair, uneven skin sheen, dust in a light shaft, a scuffed prop. Perfect frames read as renders; two deliberate flaws read as photography.
  • Give eyes a catchlight on any prominent face, and keep the subject tonally separated from the background — that's most of what "professional" means in a portrait.
  • Describe materials, not moods. "Salt-dulled chain mail" beats "epic armor". Models render nouns and physics better than vibes.
  • Say what is there, not what isn't. Image models handle negation poorly — phrase everything positively.
Before — adjective spam

epic cinematic 4K masterpiece of a warrior on a cliff, ultra realistic, dramatic lighting, best quality

After — a lighting plan

A weathered warrior stands on a basalt cliff at golden hour. Low sun as the key light rakes across salt-dulled chain mail; cool sky fill softens the shadows; a rim of backlight separates his silhouette from the sea haze. Wind lifts loose strands of hair, and a faint catchlight sits in his eyes.

Same idea, radically different footage.

Chapter 03 the budget Stills cost cents. Motion costs dollars. Make every mistake in the cheap medium.

Stills first. Then motion.

Professional AI video is mostly still-image work. A video render can cost as much as twenty drafts of the frame it starts from — so make every creative decision in the cheap medium first:

  1. Explore in cheap stills. Run the idea through Nano Banana or GPT Image Mini until the composition, wardrobe and light are right. Each attempt costs cents.
  2. Finish the keeper. Re-render the winning frame in Nano Banana Pro at final resolution. This is the frame your video inherits everything from.
  3. Compose for motion before you animate. Layer the frame — foreground, midground, background — leave negative space in the direction things will move, and keep hands and faces away from the frame edges, where video models degrade first.
  4. Animate a draft, then the finish. Drive the still through Seedance Fast at 480p to prove the motion, then pay finish rates once, not five times.

Fetching live draft-vs-finish quotes…

Chapter 04 the take One shot, one camera move. Coverage comes from more shots, not busier ones.

Direct the motion. Don't describe a montage.

Video prompts fail differently from image prompts: the model has to spend your seconds. Give it a single take it can actually execute:

  • One camera move, named. A slow push-in, a drift left, a rise. Two moves in one clip is how you get neither.
  • Describe change over time. What enters, what settles, what the light does between the first frame and the last. A video prompt is a sentence with a beginning and an end, not a list of attributes.
  • Keep clips short and cut like an editor. Two intentional 5-second shots beat one 15-second wander — and cost about the same. Sequences are made in the edit, not in one render.
  • Add audio when it earns its cost. Sound is priced into the quote on models that support it — flip it on for the finish, not the drafts.

When you'd rather write one plain sentence, the studio's prompt enhancement expands it with this craft — light, motion grammar, composition — before it reaches the model, and stores the expanded prompt with the render so you can read exactly what was sent and learn from it.

Read enough. Go shoot something.