Make a duo rap-style video from two photos
Turn a pair of portraits into a short orange-booth performance. This creator is for a new duo scene with generated gestures and possible audio. Start with the two people you want on screen; the performance setup is already prepared.
Is this the right AI rap video workflow?
Use this creator when you want two people sharing a short studio scene, with expressions, gestures and camera movement generated from their photos. A playful duo reveal or a friends' studio joke fits the format. The current output is 8 seconds, 720p and 16:9 landscape.
If you need a complete music video for a finished song, plan a different workflow: an audio track, multiple shots and an editing timeline. This preset has no song upload, lyric input or exact lip-sync control. It cannot guarantee that the people will deliver a particular verse or move to an existing beat.
You do not need to select a model or write a prompt here. The built-in scene gives both people the same performance setting while allowing different gestures. The template guide compares this approach with fixed-motion and editor templates.
From two portraits to a saved clip
- Choose the pair. Use one clear photo of each adult, with permission to create the clip. Avoid a crowd, a tiny face or an object covering the mouth.
- Check the previews. Open the creator, add both images and confirm that each slot contains the intended person. Photo preparation alone is not a completed video.
- Review the generation cost. Sign in and check your balance and the displayed credit requirement. Current packages appear in pricing.
- Generate once and follow the task. Wait for a final status. If the connection drops, check your recent videos before starting a separate task.
- Play and download. Watch the full result, listen to its sound, then save a useful version from your task history.
A short clip can still vary in how closely it preserves a face or how naturally the hands move. Clear inputs improve the starting point but are not a guarantee of a particular result. For upload and recovery questions, use the detailed tutorial.
What happens to the music and vocals?
The video model can generate sound alongside the visuals. You may hear a musical backing, vocal-like performance or other generated sound, but the audio is not a selectable track and may be absent. Listen before posting, especially if the output appears to put words in someone's mouth.
The preset requests original audio. It does not include the original Hotel Lobby recording or let you enter exact lyrics. Paying for a generation does not choose a particular song. See the audio guide for soundtrack options.
If the picture works but the sound does not, keep the picture. In an editor, save a copy, mute the original sound and add audio suitable for your intended use. Check the mouth movement after replacement: changing the track does not regenerate the performance or fix synchronization.
Prepare the landscape clip for a social post
Before exporting a vertical version, preview where both faces sit throughout the shot. Cropping a 16:9 scene to fill a 9:16 screen removes much of its width and can cut out one performer. An editor layout that keeps the whole landscape clip inside a vertical canvas can preserve the pair more reliably.
Keep captions away from faces and hands. An original line such as “Our imaginary studio debut” can introduce the joke. Check your final export on a phone, including the first frame, sound level and ending. A smooth preview inside an editor is not the same as checking the saved file.
Ask both participants to review a realistic clip before sharing it. Describe it as AI-generated so viewers understand that the apparent performance was created from photos. The site's content policy covers misleading impersonation and other prohibited uses.
Improve the next version without starting over blindly
Face hard to recognize: replace the affected portrait with a sharper, unobstructed image. Change that input first so you can tell whether it helped.
One person disappears in a vertical edit: return to the uncropped download and choose a layout that preserves both performers. This can be an editing problem rather than a generation problem.
Different movements from a trend reference: this preset creates a new performance. Keep a coherent version you like; use a reference-motion tool if a fixed routine is essential.
Unwanted sound: try editing the audio of the saved clip before generating new visuals.
Task still processing: check recent videos and its final status before submitting again. For a persistent issue, send the task identifier through support, without including signed media links in a public post.