Without sound
Visual motion only — your image and motion prompt animate the clip. No soundtrack, SFX, or spoken lines.
- Motion + lighting from your prompt
- No separate audio field
- Lower token cost per generation
Upload a still photo, write your own prompt for motion and mood, then generate a short clip in the Playbox app — with or without sound.
One image + your prompt drives the generation. Audio is optional and uses a separate sound prompt when enabled.
A clear photo of the subject you want to animate — face, full body, or scene.
Your own description: camera move, gesture, lighting, and atmosphere.
Add SFX, music, or dialogue — models can speak popular languages.
Playbox AI turns one still into a short, prompt-driven clip. You steer the shot in plain language; the app handles rendering after you continue from the site.
Drop a PNG, JPG, or WEBP. Clear subjects and good lighting give the strongest results.
Describe motion, camera, and mood in your own words — or start from a prepared preset.
Generate silent video, or enable sound and fill a separate field for audio direction.
Every image-to-video run uses tokens. Adding audio costs a little more and unlocks a dedicated sound prompt for SFX, music, and dialogue.
Explain the exact SFX, music bed, or spoken lines you want — keep it concrete so the model can match the scene.
People in the photo can speak in all popular languages. Write the dialogue in the language you want to hear.
The number of words does not change the generation cost. With sound is always 30 tokens; silent stays 25.
Short, specific phrases usually beat long vague paragraphs — for both motion and sound.
Name camera move, subject action, and mood (e.g. “slow push-in, soft smile, warm window light”). Leave audio out of this field when you use the sound option.
Separate layers: ambience, music, then dialogue in quotes. Example: “Rain on glass. Quiet guitar. She says: ‘We should go.’”
You can use a prepared preset or write a custom motion prompt — both work for image to video.
With-sound generations use 30 tokens instead of 25 because audio (SFX, music, or speech) is rendered too.
Yes. Describe dialogue in the sound prompt. Speech works across popular languages; longer scripts do not raise the token price.
Preview on the site if you like, then continue to the Playbox app to run the final generation.