Let’s Make a Music Video ~ Step 1

Let us plan a music video built from Maroli Tango chapter illustrations. We will select start frames and task PAI Studio with animation and assembly.

ChatGPT replied:

The chapter illustrations are unusually well suited to this because they already function like cinematic keyframes. Rather than asking an image-to-video model to invent scenes, we can treat each illustration as a deliberately composed start frame and give PAI Studio narrowly constrained animation jobs.

I’d build the project in five passes.

1. Treat the song as the master timeline

Before choosing pictures, we map the finished music track into structural beats: intro, verse, lift, chorus, bridge, instrumental passages, final escalation, outro.

For each section we decide whether the visual should be:

  • narrative — recognizable events from Maroli Tango
  • character — Carmen, Brandon, Mason, Marie, Suban, etc.
  • world spectacle — ships, Virtuality, Dust Cloud Spiral, Washington, Jivada
  • emotional punctuation — a look, embrace, confrontation, isolation, revelation
  • kinetic transition — drone flight, spacecraft departure, city movement, camera travel

That keeps the video from becoming an illustrated chapter summary. It should feel like the emotional memory of the novel set to music.

2. Build a start-frame library

We go through the existing illustrations and classify them roughly like this:

TypeBest animation treatment
Close character shotbreathing, eyes, hair/fur, subtle expression, camera creep
Two-character dialoguerestrained body motion, glance changes, parallax
Interior scenecamera push/dolly, AR displays, environmental motion
Exterior citytraffic, atmospheric movement, signs/light, slow camera
Spacecraftcamera orbit, translation, atmospheric haze, distant traffic
Action imagestronger camera movement, debris, weapon effects, character movement
Virtuality/surrealtransformation, impossible camera travel, environmental morphing

Some images will make excellent 4–6 second shots. Others may support 8–10 seconds if there is enough depth.

A critical rule: don’t animate everything. Models get worse as the requested motion becomes more complicated. We want the existing illustration to remain recognizable.

For example:

Carmen doesn’t need to walk across a room.
She may only look toward Brandon while the camera moves slowly around her.

That can look much more expensive than attempting full-body performance.

3. Write animation directions as shots, not prose

For every selected start frame, we’ll create a compact production card.

Something like:

SHOT 17 — Carmen / White House

Duration: 5.5 sec
Start frame: Carmen chapter illustration
Camera: slow 15-degree arc right, slight push-in
Subject: Carmen remains nearly still; subtle breathing; eyes shift toward someone offscreen
Environment: curtains move faintly; distant city lights shimmer
Mood: controlled anxiety, private rather than presidential
Continuity constraints: preserve face, age, wardrobe, hairstyle, room architecture
Avoid: dialogue mouth motion, exaggerated head turn, new objects, costume changes

The important distinction is that we’re telling PAI what may move and what must not change.

For character continuity, I’d repeatedly specify:

  • preserve exact facial identity
  • preserve species anatomy
  • preserve costume
  • no added tails
  • no added fingers/limbs
  • no new characters
  • no text mutation
  • maintain illustrated/cinematic realism

Especially with the Anye and Maroli, anatomical restraint will matter.

4. Use the music video grammar to hide AI weaknesses

Rather than producing thirty-second generated sequences, I would favor short deliberate shots, probably averaging around 3–6 seconds.

A very rough rhythm might be:

Quiet verse: 5–8 sec shots
Build: 3–5 sec
Chorus: 2–4 sec
Instrumental/action passage: occasionally 1–2 sec
Emotional climax: return to longer shots

And not every transition needs generative animation. Hard cuts, dissolves, dip-to-black, match cuts, flash frames, and editorial reframing can do enormous work.

We can also reuse a strong illustration in multiple ways:

Shot A: wide composition
Shot B: crop on Carmen
Shot C: crop on Brandon
Shot D: animated full frame

Four visual beats from one piece of artwork without regenerating anything.

5. Give the video an arc independent of chapter order

For Maroli Tango, I would initially explore an arc something like:

I — Wonder
Earth, Anye civilization, Washington, ships, children, strange technology.

II — People
Carmen/Brandon. Mason/Marie. Domestic and human-scale images.

III — Something is wrong
Dust Cloud Spiral, military imagery, drones, meetings, increasingly uneasy expressions.

IV — Collision
Political confrontation, spacecraft, military action, rapid cutting.

V — Choice
Carmen becomes the visual center. The story stops being about whether somebody will rescue Earth and becomes about humanity deciding what it will do.

VI — Release
Large cosmic imagery followed by something surprisingly intimate rather than ending on the biggest explosion.

That would serve the central idea of Maroli Tango particularly well:

the fate of civilizations rendered through the decisions of individual people.


The practical workflow can therefore be:

song → timecoded beat sheet → illustration shortlist → shot order → individual PAI animation prompts → generated clips → assembly → identify weak transitions → create a handful of purpose-built bridge frames → final assembly.

I would resist generating new artwork until we’ve exhausted the existing chapter illustrations. Their consistency is an asset. Once the rough cut exposes an actual missing shot—we need Carmen alone here, we need an exterior of Bharamin here, we need Mason looking upward here—then we make that illustration specifically as a start frame.

The best next step is to give me the song/audio or lyrics. I can then turn the actual track into a shot-by-shot visual treatment and we can start choosing which existing Maroli Tango illustrations belong at each musical moment.

Leave a comment

Blog at WordPress.com.

Up ↑