“Putting yourself in an AI video” sounds like one operation, but the versions of it that actually look good are built from two: an image merge that swaps a character into a real photo of your scene, and motion transfer that drives the merged image with a short clip of you moving. Do them in that order and the character inherits your lighting, your framing, and your performance — which is exactly what makes the result read as filmed rather than generated.
This is the workflow Vidshift’s Character Swap walks you through in five steps. Here is each step, with the details that matter.
Step 1 — Shoot a base photo like it’s the first frame
The base photo is not a profile picture; it is the set of your video. Shoot yourself in the exact place, framing, and outfit your final video should have, because everything about the scene survives the swap. Three things pay off disproportionately:
- Even light on your face and body. The merge keeps your lighting, so harsh shadows or a blown-out window behind you end up on the character too.
- Room around you in the frame. Motion needs space — if your elbows touch the edges of the photo, they will clip once you move.
- A phone photo is fine. Uploads are compressed client-side, so even a large photo lands in under a second on mobile data.
Step 2 — Pick a character reference that shows what you mean
The character image tells the model who you become. A clear, front-lit image of the character’s face and upper body beats concept art with dramatic side lighting. One rule is not negotiable: the character must be fictional, AI-generated, your own persona, or a person who has consented. Swapping a real, identifiable person in without permission is a deepfake, and it is banned under the acceptable-use policy.
Step 3 — Merge into one portrait
The merge blends the character into your base photo while keeping the scene, pose, and lighting — you get a single photorealistic 2K portrait, usually in under a minute. Your first merge is free, no account needed. Look at it critically before moving on: this portrait is the single frame every second of your video will be built from, so a flaw here (mangled hands, a warped background edge) appears in every frame later. Regenerating with a tweaked prompt is cheap; re-rendering a video is not.
Step 4 — Record your reference video in the same spot
Now film yourself performing the motion you want the character to make: MP4 or MOV, 3 to 30 seconds, up to 100MB — a normal phone recording qualifies. The single biggest quality lever in the whole workflow is matching the base photo: same spot, same framing, same distance from camera. The animation maps your body in the video onto the body in the portrait; the closer they start, the cleaner the transfer. For the finer points, see the motion-control guide.
Step 5 — Animate and export
The final step transfers your recorded performance onto the merged portrait. It is the heavy computation — typically 4 to 7 minutes — and it bills per second of output, so trim your clip to the seconds you actually want before uploading. The free tier renders at 720p with a watermark; Creator and Pro plans unlock 1080p and 4K without one.
How long does all of this take?
About ten minutes the first time, including shooting the photo and the clip. The steps save progress per browser, so you can shoot on your phone, continue on desktop, or do the whole thing from the phone the footage lives on. Start with the free merge and judge the portrait before spending anything on animation.