Creating a Single Talking Actor
The Talking Actor tool generates a lip-synced video by animating an actor from a reference image or video using a voice script via Text-to-Speech (TTS) with voice cloning or an uploaded voice.
Required Elements
- Image / Video: Reference image or video of the actor. Supported formats: PNG, JPG, JPEG, WebP, MP4, WMV, MOV, and AVI.
- Voice: Load an audio file for the actor’s voice, or use a preset or reference audio for TTS voice generation. Supported formats: MP3 and WAV.
AI Settings
Select the Talking Actor
Tool from the left AI Toolbar and complete the tasks in each section under the Single Talking Actor tab.
You can collapse or expand a section by clicking its caption.
- In the Reference Image / Video section, click the Media slot to browse and upload a reference image (.png, .jpg, .jpeg, or webp) or video (.mp4, .wmv, .mov, .avi) containing the target actor for speech generation.
You can also drag an image or video directly from the History panel on the right or from Windows File Explorer.

Note that long videos will be automatically trimmed to 30 seconds. - The media will be inserted into the Media slot.

Hover over the media thumbnail to view an enlarged version, or click the Trash button to remove the current input.
The actor's voice can be uploaded via Audio to Voice or generated via Text to Voice (TTS).
Text to VoiceText to Voice uses Text-to-Speech (TTS) technology to generate audio from a text script, using either preset voices or voice cloning based on reference audio.
- Under the Text to Voice tab in the Voice Generation section, select a voice source method for speech generation: Presets or Reference Audio.

- Presets: A selection of male and female preset voices is available in the drop-down list across different AI Actors, allowing you to assign a voice to the actor in the video.

Click the Play button to preview the selected AI voice. - Reference Audio: Click the Folder icon (A) to upload a pre-recorded voice file (.mp3 or .wav) for AI voice cloning of the speaker’s tone, pitch, and accent.
You can also drag and drop audio directly from the History panel on the right (B) or from Windows File Explorer into the field.

- Presets: A selection of male and female preset voices is available in the drop-down list across different AI Actors, allowing you to assign a voice to the actor in the video.
- Enter the actor’s dialogue in the Speech Text field.
Use expressive interjections or emotional prefixes to enhance vocal emotion.
For example, "Hey, you’re late! I thought you guys weren’t coming!".

The AI model will automatically detect the input language. - Click Generate Voice to create audio from the current text using the selected voice source.
Note the task price displayed above the button before proceeding.

- Track progress on the AI Render View or via the new entry in History on the right.
When the submission completes successfully, the AI-generated audio will appear.
Click Play to play the audio, Loop to repeat playback, or drag the playhead to jump to a specific point.
- Review the Preview Voice information to see the generated audio length, as it will be automatically trimmed to match the video duration (up to 30 seconds).
Click Play next to it to preview the voice audio.

- Under the Text to Voice tab in the Voice Generation section, select a voice source method for speech generation: Presets or Reference Audio.
Audio to VoiceUpload an audio file to use as the actor’s voice, preserving the original speaker’s voice and spoken content.
- In the Voice Generation section, go to the Audio to Voice tab and click the Media slot (A) to upload a voice script (.mp3 or .wav).
You can also drag and drop audio from the History panel (B) or Windows File Explorer into the Media slot.

- The audio will appear in the Media slot as a waveform indicating its total duration.
It will be automatically trimmed to match the video length (up to 30 seconds).
Click Play to preview the audio, or X to remove it.
- In the Voice Generation section, go to the Audio to Voice tab and click the Media slot (A) to upload a voice script (.mp3 or .wav).
You can also drag and drop audio from the History panel (B) or Windows File Explorer into the Media slot.
Enter prompt descriptions in the Prompt section to define the scene and control the actor’s state and emotion in the generated video.

For example, "The boy sat on a fallen log, his legs dangling and swinging in the air. Fireflies danced around him, and the boy looked up at them with delight, his movements and expressions animated. He looked left and right, gesturing wildly as he talked to the fireflies.
The camera work is a slow, steady dolly-in towards his upper body".
Drag the horizontal divider downward to expand the text field and display more content.
Click the Reset Prompt button to clear the prompt field and start over if needed.
The output video is limited to 720p resolution and 30 seconds in duration.
Hover over the exclamation mark
icon in the Voice Generation section to view additional operational notes.

Each submitted task is charged using your available AI, Bonus, or DA Points.
Click the account icon in the bottom-left corner of AI Studio to view your current point balance.
Points are deducted in the following order: AI Points → Bonus Points → DA Points.

To obtain additional credits, subscribe to an AI Service Plan for more AI Points, or purchase DA Points to top up your account.
Ensure both an image / video and voice audio are provided to enable the GENERATE button.
Note the task price displayed above the button before clicking.

Track progress on the AI Render View or via the new entry in History on the right side of the AI Workspace.
When the submission completes successfully, the AI-generated video will appear.
Click Play to play the video, Loop to repeat it, or drag the playhead to a specific point.
Next, use the Video options in the Quick Access Menu at the top of the AI Render View to edit the video or upscale the generated output to a higher resolution.
