AI Motion Control

Give it a character image and a video of the movement you want, and the motion from the reference is carried onto your character. You are showing the model the movement rather than writing an adjective and hoping. That is the difference between directing a shot and rolling the dice on a prompt.

How motion control works

  1. Upload your character

    A photograph, a still you generated earlier, or an illustration. This is who appears in the finished clip. Their face, their clothing, their proportions. Whatever the reference does, it will be this character doing it.

  2. Add the motion reference

    A short video of the movement itself: a dance, a walk, a gesture, a piece of choreography, someone demonstrating a technique. This clip supplies the timing and the body movement; nothing about how it looks needs to match your character.

  3. Choose where the framing comes from

    Character orientation set to "From Image" keeps the pose and framing of your still and takes only the movement, with references up to around ten seconds. Set to "From Video", the orientation follows the reference instead and you can use clips up to about thirty seconds.

  4. Set quality and decide about sound

    Output runs at 720p or 1080p. There is also a toggle for keeping the original sound from the reference clip, which matters when the movement is tied to music or speech and gets in the way when you intend to lay your own audio over the top.

  5. Add a prompt only if you need one

    The text box is optional and it is there for motion or style, not for describing the whole scene. The reference is already saying what the movement is; the prompt is where you nudge the look of it.

  6. Generate, then use it

    The clip lands in your library alongside everything else, so it can be cut against other footage or finished with voiceover and captions. Generation takes as long as it takes. Expect to wait through it rather than watch it appear.

What people use it for

Dance and choreography

Full-body movement is the hardest thing to get from a text prompt and the easiest thing to film. Record the routine once, or use a reference you have rights to, and put your character through it exactly as performed.

A brand character that performs

Mascots, illustrated hosts and recurring presenters need to do the same things week after week. Motion control lets you reuse one approved character image and change only what they are doing.

Consistent presenters across a series

A series falls apart when the person on screen is subtly different in every episode. Holding the character image fixed and varying the reference keeps the performer recognisably the same across an entire run.

Demonstrations and gestures

A stretch, a lift, a hand movement, a piece of equipment being used correctly. When the exact motion is the content, describing it in words is the wrong tool. Showing it is the whole job.

Social formats built on movement

Short-form video is full of formats where the movement is the format. Matching one precisely matters more than inventing something new, and a reference clip is a precise instruction in a way that "trendy dance" is not.

Bringing a still character to life

Concept art, a book character, a game design, a photograph of someone in costume. The picture already exists and is already approved; what it lacks is a performance, and that is exactly what the reference supplies.

What is motion control in AI video?

Motion control means deciding the movement in a generated clip rather than leaving it to the model. In this tool it works by transfer: you supply a character image and a separate reference video, and the movement in the reference is applied to your character. The pose over time, the timing, the rhythm of the performance.

This splits a generation into two decisions that are normally tangled together. Appearance comes from the image: face, clothing, proportions, style. Movement comes from the video: what the body does and when. Because each is fixed by something you supplied, neither is being invented, and neither can drift into something you did not ask for.

It is worth separating two things that both get called camera or motion control. Take camera language. A slow push in, a pan, a drift. A text prompt carries that reasonably well, because a camera move is simple and continuous. Body movement is not: a dance has hundreds of positions in a specific order, and there is no sentence that specifies them. That is the gap motion transfer fills.

The practical consequence is repeatability. Prompt-driven motion is probabilistic. Run it twice and you get two performances. A reference video is deterministic input: the same reference produces the same choreography, on any character you point it at. For anyone making a series, or matching a format, or approving work before it goes out, that predictability is worth more than variety.

Motion control vs describing the motion in a prompt

The alternative is image-to-video with a written description of the movement, and for a great many shots it is the right choice. A prompt is fast and it needs nothing but words. Models handle the vocabulary of cinema well enough that adding a reference clip would be pointless ceremony. A push in, a slow orbit, hair moving in wind.

Prompts fail in a specific place: complex, sustained, human movement. Direction and speed are one or two numbers, and a sentence can carry them. A dance routine, a martial arts form, a gesture sequence or a piece of physical comedy is a long list of body positions in a strict order, and no amount of adjectives will specify it. Ask for it in words and you get something in the general area. Plausible movement, but not the movement. If the movement was the point, plausible is a failure.

The second failure is repeatability. Because generation is probabilistic, the same prompt gives a different performance every time. That is useful when you are exploring and painful when you are matching. A second video in a series, a variant of an approved ad, the same routine performed by three different characters. A reference clip removes the variance: the movement is fixed input, so what changes between runs is only what you changed.

The third is describing what you cannot name. Some movement has a name a model knows, and some is just how a particular person moves. A way of walking, a habit of the hands, comic timing. You can film that in one take. You will not write it down.

The two techniques also sit at different points in the same process, and they combine. Generate or choose the character still first. Image tools return results faster than video, so that is the place to iterate on how the character looks. Once the picture is right, use it as the character input here and let a reference video decide what it does. If you want a camera move on top of a fixed performance, that is a prompt-driven job, and the two clips can meet in the edit rather than in a single generation.

Getting better results

  • Match the framing of image and reference

    This is the single biggest cause of poor results. A head-and-shoulders portrait paired with a full-body dance reference is asking the model to invent a body it has never seen. If the movement involves legs, the character image should show legs. Frame both the same way and most other problems go away.

  • One subject in the reference, clearly separated

    The reference should have one person, fully in frame, distinct from the background. Crowds, partial occlusion, someone walking behind furniture, or a subject that matches the wall behind them all make the movement harder to read, and anything the model cannot read it will guess at.

  • Choose the orientation setting deliberately

    "From Image" holds the pose and framing of your still and takes only the movement. That is the right choice when the character shot is already composed the way you want. "From Video" hands orientation to the reference and accepts much longer clips, which is what you want when the reference has its own camera work worth keeping.

  • Keep the reference camera steady

    A locked-off or gently moving camera gives clean, unambiguous body movement. Handheld footage mixes camera shake into the performance, and the result is a character that appears to move in ways the performer never did.

  • A sharp, well-lit character image

    The generation inherits whatever is in your still, including soft focus, heavy compression and awkward cropping. Even lighting and a clear silhouette give the model something to hold on to; a dim, cluttered image gives it less, and the shortfall shows up as instability once the character starts moving.

  • Use the prompt for style, not choreography

    The reference is already specifying the movement, so repeating it in words adds nothing and can pull against it. Keep the text for how it should look. The quality of light, the mood of the setting. Let the video do the directing.

  • Iterate at the lower quality

    Work out the pairing of character and reference at 720p, where you are testing whether the combination works at all. Move to 1080p once you have settled on the version you want, rather than rendering every experiment at the higher setting.

  • Decide about sound before you generate

    Keep the reference audio when the movement is tied to it. Music with a beat, or speech where the mouth has to match. Turn it off when you are building your own soundtrack, so you are not editing around audio you never intended to use.

  • Cut the reference to the part you need

    Trim to the movement itself before uploading. Lead-in, setup and a few seconds of standing still are all generated material you will only cut out later, and a shorter reference keeps the whole run tighter.

Questions

What do I need to supply?
Two things, both required: a character image and a reference video showing the movement. The text prompt is optional and exists for motion or style notes. The reference is already carrying the actual choreography.
How long can the reference video be?
It depends on the character orientation setting. With orientation taken from your image, references run up to around ten seconds. Switching orientation to come from the video allows much longer references, up to roughly thirty seconds. That is the setting to use when you have a full routine to transfer.
Does my character need to look like the person in the reference?
No. The reference supplies movement, not appearance. What does help is similar framing and a comparable body position at the start. A full-body reference works best with a full-body character image, because the model is mapping one onto the other.
What resolution does it output?
You choose 720p or 1080p before generating. Testing combinations at 720p and reserving 1080p for the version you intend to publish is the sensible order, since most of what you learn from an attempt is visible at either setting.
What happens to the audio from the reference?
There is a toggle for keeping the original sound from the reference clip. Keep it when the movement is tied to music or speech; turn it off when you plan to add your own voiceover, music or captions in the editor afterwards.
Is this the same as controlling the camera?
Not quite. This tool controls what the subject does, by transferring movement from a reference. Camera movement is something a written prompt handles well in the video generators, because a camera move is a simple continuous instruction in a way that a body performance is not. A push in, a pan, an orbit.
Can I use the same character across several clips?
Yes, and that is much of the point. Hold the character image constant and change the reference video, and you get a run of clips featuring recognisably the same character doing different things. Far more consistent than generating each shot from a prompt.
Why did the result come out distorted?
Most often the character image and the reference disagree about framing, or the reference has more than one person, an unsteady camera, or a subject that blends into the background. Re-crop so both clips show the same amount of the body, pick a cleaner reference, and try again.
Can I use what I make commercially?
What you make is yours. Check the terms for the specifics, and make sure you have the rights to both the character image and the reference footage you are transferring movement from.

AI Motion Control

Give it a character image and a video of the movement you want, and the motion from the reference is carried onto your character. You are showing the model the movement rather than writing an adjective and hoping. That is the difference between directing a shot and rolling the dice on a prompt.

Start creating