Most AI video is prompted into existence: you describe motion in words and hope the model choreographs it. Motion control inverts that. You give Kling a still image of a character and a real video of a person moving, and it renders the character performing that exact movement — your timing, your gestures, your pauses. It is the difference between directing an actor and describing one.
In Vidshift’s Character Swap, motion control is the final step: the still image is the merged portrait of you as the character, and the video is a clip you film of yourself. Because you control both inputs, you control almost everything about the output — which also means most quality problems are filming problems, and they are fixable.
What the reference video must be
- MP4 or MOV — phone cameras record in these already. WebM is rejected.
- 3 to 30 seconds long, up to 100MB.
- One person, fully in frame, doing the motion you want.
Length is worth deciding deliberately: animation bills per second of output (4 credits per second on the standard model, 5 on Pro), so a 30-second clip costs almost four times what an 8-second one does. Trim before you upload — the strongest few seconds of a performance nearly always beat all of it.
The one rule: match the still image
Motion control maps the body in your video onto the body in the image, frame by frame. Every difference between them is something the model must invent its way around, and invention is where warping comes from. So film the reference clip in the same spot, framing, and camera distance as the base photo you merged from. Feet in the photo but not the video (or vice versa) is the classic mistake — match your crop. Keep the camera still; a tripod or a propped phone beats a handheld shot.
Movements that transfer well
Deliberate, mid-speed movement with clear silhouettes transfers almost perfectly: talking with gestures, walking, dancing with defined moves, turning, pointing. What degrades: very fast motion (blurs the mapping), long spins (the model must hallucinate the character’s back), heavy occlusion like hands crossing the face for many frames, and interactions with props the portrait doesn’t show. None of these are forbidden — they just cost attempts. Expect one or two retakes when you push complexity; iterate on the short clip, not the 30-second epic.
Follow the video, or follow the image?
Vidshift exposes Kling’s two orientations as a choice. Follow my movement prioritizes the video: best for complex body motion, supported up to 30 seconds. Follow the portrait’s framing prioritizes the image’s composition: best when the shot itself matters — camera moves, tight framing — and runs up to 10 seconds. If in doubt, follow your movement; it is the option that makes the result feel performed.
Try it on your own footage
The full workflow — photo, character, merge, reference clip, animate — is laid out in the step-by-step guide, and the swap itself starts with a free merge, so you can see your character before spending a credit on motion.