# Create AI Avatar and Talking Head Videos

Producing presenter videos requires coordinated image, speech, and lip sync tools. This skill builds inference.sh workflows for avatars, localization, and batch content.

## Install

```bash
npx skillstore add infsh-skills/ai-avatar-video
```

## Metadata

- Status: approved
- Slug: infsh-skills-ai-avatar-video
- Skillstore revision: r2
- Version status: missing
- Tree hash: 4649b11d874126d484157e56354978a97fdf02eb0535cd752851f66c09082db7
- Author: infsh-skills
- GitHub username: infsh-skills
- License: MIT
- Repository: https://github.com/infsh-skills/skills/tree/main/tools/video/ai-avatar-video/
- Ref: 4121de961d1b6f2ffca856260e239505c302452c
- Supported tools: Claude, Codex, Claude Code
- Audit status: complete
- Agent install advisory: allowed
- Manual install advisory: allowed
- Artifact signature: available
- Audit attestation: unavailable
- Human verification: not\_verified
- Risk factors: external\_commands, network
- Quality score: 69
- Public page: https://skillstore.pages.dev/skills/infsh-skills-ai-avatar-video
- Manifest: https://skillstore.pages.dev/api/skills/infsh-skills-ai-avatar-video/manifest

## Capabilities

- Creates talking head videos from a portrait and written script.
- Animates portraits with user-provided audio for lip sync.
- Combines portrait generation, speech generation, and avatar rendering.
- Builds transcription, translation, speech, and lip sync workflows for dubbing.
- Generates multiple presenter variants with fixed voice and style settings.

## Use Cases

- Produce Product Presenters: Create scripted presenter videos for product demos, announcements, and short advertisements.
- Localize Existing Videos: Transcribe, translate, synthesize speech, and synchronize a video for another language.
- Animate Course Material: Turn lesson scripts and approved portraits into consistent educational presenter clips.

## Prompt Templates

### Create a Basic Avatar

```
Create a 720p talking head video from my portrait and script. Use a natural English voice and show the planned model before running it.
```

### Match a Presenter Style

```
Create a vertical avatar video with a calm professional voice, restrained gestures, and a neutral office background. Use my approved portrait and script.
```

### Localize a Video

```
Transcribe my video, translate the speech into Spanish, generate matching audio, and lip sync the original. Pause for approval before each paid stage.
```

### Plan a Presenter Batch

```
Design three presenter variants for this campaign. Compare suitable models, estimate job count and cost, confirm media rights, then run only my approved variants.
```

## Limitations

- Requires the belt CLI, an inference.sh account, network access, and available credits.
- Uploads media references and prompts to third-party model services.
- Output quality depends on source portrait lighting, framing, and audio clarity.
- Model availability, prices, voices, languages, and performance can change.

## Best Practices

- Obtain consent and usage rights for every face, voice, script, and source video.
- Review the selected model, media destination, job count, and estimated cost before execution.
- Use front-facing portraits with clear lighting and clean audio for stable lip sync.

## Anti Patterns

- Do not upload confidential, biometric, or licensed media without authorization.
- Do not install unpinned skills or follow mutable remote instructions without review.
- Do not launch paid batch jobs before the user approves scope and expected charges.

## Security Audit

- Audited at: 2026-08-06T11:43:14.144\+00:00
- Summary: Most static findings are false positives caused by Markdown backticks, code fences, placeholder media URLs, or ordinary documentation links. Confirmed findings concern unpinned skill installation and mutable remote installation guidance. The workflows also send potentially biometric media to third-party processing and can incur paid service usage.

## Stats

- Views: 62
- Downloads: 12
- Favorites: 0
- Popularity score: 0
