AI Talking Photo Generator

Talking photo animates a single portrait so it speaks your audio: either a recording you upload or text read by an AI voice. It works from a still image — for re-syncing an existing video to new audio, use lip sync instead.

Photo

Uploads are private to your account, so they need an account. Sign in

Voice
0/300
Complete the inputs to see the price

Credits are spent only when a video is delivered. If a job fails, the held credits are returned automatically.

New accounts get 130 credits once after sign-in. What they cover

Sign in to generate

Examples

Real outputs generated with this tool — not edited.

Oil-painting portrait speaks

Input photo

Input: I painted this portrait last summer. It took me three weeks and a lot of coffee.

Model: Kling AI Avatar v2 Standard + Kokoro TTS

Try this example

Portrait says hello

Input photo

Input: Hi there! Thanks for stopping by. Today I want to tell you about my favourite place in the city.

Model: Kling AI Avatar v2 Standard + Kokoro TTS

Try this example

Birthday wishes from a photo

Input photo

Input: Happy birthday! I hope this year brings you lots of laughs, good food and great friends.

Model: Kling AI Avatar v2 Standard + Kokoro TTS

Try this example

A portrait reads a fun fact

Input photo

Input: Did you know that honey never spoils? Archaeologists found edible honey in ancient tombs.

Model: Kling AI Avatar v2 Standard + Kokoro TTS

Try this example

Cartoon explorer starts talking

Input photo

Input: The weather tomorrow will be sunny in the morning, with light rain after four.

Model: Kling AI Avatar v2 Standard + Kokoro TTS

Try this example

How to use it

  1. Step 1

    Upload a portrait

    A clear, front-facing face. Photos, paintings and illustrated characters can work.

  2. Step 2

    Add the voice

    Type up to 300 characters and pick a voice, or upload MP3/WAV/M4A audio up to 30 seconds.

  3. Step 3

    Generate & download

    The video length follows the audio. Preview it and download the MP4 (the voice track is downloadable too).

What people use it for

Greeting messages

A portrait that says happy birthday in your words.

Character intros

Give an illustrated character a voice for a short.

Explainer hosts

A talking presenter for a quick product explanation.

Input requirements

  • Portrait: JPG/PNG/WebP up to 10 MB; one face, front-facing, mouth visible.
  • Audio: MP3, WAV or M4A, up to 30 seconds — or text up to 300 characters.
  • Use only images and voices you own or have permission to use. No impersonation of real people without consent.

Limits & cost

  • Maximum 30 seconds of speech per video.
  • Side profiles, covered mouths and very small faces reduce lip-sync quality.
  • Burned-in subtitles are not added yet.
  • The exact credit price is shown before you generate; credits for failed jobs are returned automatically. See credit pricing.

Typical credit cost

  • Kling AI Avatar v2 Standard, 10s of speech85 credits
  • Kling AI Avatar v2 Standard, 30s of speech253 credits
  • Kling AI Avatar v2 Standard + Kokoro TTS (American English), 300 characters (≈25s of speech)254 credits

Common problem: The mouth barely moves.

Fix: Use a larger, front-facing face with the mouth clearly visible and audio without background music.

Related tools

FAQ

How much does it cost?

Every generation shows its exact credit price before you submit. Credits are held when the job starts and only spent if a video is delivered — failed, timed-out or rejected jobs return the credits automatically.

How is this different from lip sync?

Talking photo starts from one still image. Lip sync changes the mouth in an existing video to match new audio — that is a separate tool still in testing.

Which voices are available?

A small set of English AI voices. You can always upload your own recording instead.

Who can see my uploads and results?

Only you. Uploads and results are stored privately and opened through short-lived signed links after a sign-in check. Nothing is published to a gallery.

Make a photo talk

Upload a portrait and add a voice above.

Generate