AI Talking Photo Generator
Talking photo animates a single portrait so it speaks your audio: either a recording you upload or text read by an AI voice. It works from a still image — for re-syncing an existing video to new audio, use lip sync instead.
Examples
Real outputs generated with this tool — not edited.
Oil-painting portrait speaks
Input: I painted this portrait last summer. It took me three weeks and a lot of coffee.
Model: Kling AI Avatar v2 Standard + Kokoro TTS
Try this examplePortrait says hello
Input: Hi there! Thanks for stopping by. Today I want to tell you about my favourite place in the city.
Model: Kling AI Avatar v2 Standard + Kokoro TTS
Try this exampleBirthday wishes from a photo
Input: Happy birthday! I hope this year brings you lots of laughs, good food and great friends.
Model: Kling AI Avatar v2 Standard + Kokoro TTS
Try this exampleA portrait reads a fun fact
Input: Did you know that honey never spoils? Archaeologists found edible honey in ancient tombs.
Model: Kling AI Avatar v2 Standard + Kokoro TTS
Try this exampleCartoon explorer starts talking
Input: The weather tomorrow will be sunny in the morning, with light rain after four.
Model: Kling AI Avatar v2 Standard + Kokoro TTS
Try this exampleHow to use it
- Step 1
Upload a portrait
A clear, front-facing face. Photos, paintings and illustrated characters can work.
- Step 2
Add the voice
Type up to 300 characters and pick a voice, or upload MP3/WAV/M4A audio up to 30 seconds.
- Step 3
Generate & download
The video length follows the audio. Preview it and download the MP4 (the voice track is downloadable too).
What people use it for
Greeting messages
A portrait that says happy birthday in your words.
Character intros
Give an illustrated character a voice for a short.
Explainer hosts
A talking presenter for a quick product explanation.
Input requirements
- Portrait: JPG/PNG/WebP up to 10 MB; one face, front-facing, mouth visible.
- Audio: MP3, WAV or M4A, up to 30 seconds — or text up to 300 characters.
- Use only images and voices you own or have permission to use. No impersonation of real people without consent.
Limits & cost
- Maximum 30 seconds of speech per video.
- Side profiles, covered mouths and very small faces reduce lip-sync quality.
- Burned-in subtitles are not added yet.
- The exact credit price is shown before you generate; credits for failed jobs are returned automatically. See credit pricing.
Typical credit cost
- Kling AI Avatar v2 Standard, 10s of speech85 credits
- Kling AI Avatar v2 Standard, 30s of speech253 credits
- Kling AI Avatar v2 Standard + Kokoro TTS (American English), 300 characters (≈25s of speech)254 credits
Common problem: The mouth barely moves.
Fix: Use a larger, front-facing face with the mouth clearly visible and audio without background music.
Related tools
FAQ
How much does it cost?
Every generation shows its exact credit price before you submit. Credits are held when the job starts and only spent if a video is delivered — failed, timed-out or rejected jobs return the credits automatically.
How is this different from lip sync?
Talking photo starts from one still image. Lip sync changes the mouth in an existing video to match new audio — that is a separate tool still in testing.
Which voices are available?
A small set of English AI voices. You can always upload your own recording instead.
Who can see my uploads and results?
Only you. Uploads and results are stored privately and opened through short-lived signed links after a sign-in check. Nothing is published to a gallery.
Make a photo talk
Upload a portrait and add a voice above.
Generate