Image to video

Image to video

Upload a still photo, write your own prompt for motion and mood, then generate a short clip in the Playbox app — with or without sound.

One image + your prompt drives the generation. Audio is optional and uses a separate sound prompt when enabled.

What you provide

Source image

A clear photo of the subject you want to animate — face, full body, or scene.

Motion prompt

Your own description: camera move, gesture, lighting, and atmosphere.

Optional sound

Add SFX, music, or dialogue — models can speak popular languages.

How image to video works

Playbox AI turns one still into a short, prompt-driven clip. You steer the shot in plain language; the app handles rendering after you continue from the site.

1. Upload a photo

Drop a PNG, JPG, or WEBP. Clear subjects and good lighting give the strongest results.

2. Write your prompt

Describe motion, camera, and mood in your own words — or start from a prepared preset.

3. Choose audio

Generate silent video, or enable sound and fill a separate field for audio direction.

Two options: without sound or with sound

Every image-to-video run uses tokens. Adding audio costs a little more and unlocks a dedicated sound prompt for SFX, music, and dialogue.

Option A 25 tokens

Without sound

Visual motion only — your image and motion prompt animate the clip. No soundtrack, SFX, or spoken lines.

  • Motion + lighting from your prompt
  • No separate audio field
  • Lower token cost per generation
Option B 30 tokens

With sound

Uses 30 tokens instead of 25. An extra prompt field appears so you can specify which sounds, music, or dialogue the video should include.

  • SFX, ambience, or score
  • Dialogue / lip-synced speech
  • Popular languages supported for speech

Sound prompt details

What to write

Explain the exact SFX, music bed, or spoken lines you want — keep it concrete so the model can match the scene.

Languages

People in the photo can speak in all popular languages. Write the dialogue in the language you want to hear.

Word count & cost

The number of words does not change the generation cost. With sound is always 30 tokens; silent stays 25.

Prompt tips for stronger clips

Short, specific phrases usually beat long vague paragraphs — for both motion and sound.

Motion prompt

Name camera move, subject action, and mood (e.g. “slow push-in, soft smile, warm window light”). Leave audio out of this field when you use the sound option.

Sound prompt

Separate layers: ambience, music, then dialogue in quotes. Example: “Rain on glass. Quiet guitar. She says: ‘We should go.’”

FAQ

Do I need my own prompt?

You can use a prepared preset or write a custom motion prompt — both work for image to video.

Why does sound cost more?

With-sound generations use 30 tokens instead of 25 because audio (SFX, music, or speech) is rendered too.

Can the person in the photo talk?

Yes. Describe dialogue in the sound prompt. Speech works across popular languages; longer scripts do not raise the token price.

Where do I generate the video?

Preview on the site if you like, then continue to the Playbox app to run the final generation.