HeyGen

Assemble a labeled avatar-style video from user-owned media without cloning real people

NOT REALLY · consider alternatives
price $29/mosubscription / year $348estimated build time closest consolation build: one sittingreplaced by 0 people

A consolation build is possible, but the paid product's decisive value sits outside a solo rebuild. For HeyGen, assemble a labeled avatar-style video from user-owned media without cloning real people. The hard boundary is proprietary avatars, lip sync, translation, rendering, templates, and rights operations, plus models, compute, rights, and safety operations.

Build verification: not recorded. How we judge buildability

What you give up

  • proprietary avatars, lip sync, translation, rendering, templates, and rights operations
  • frontier voice or avatar model
  • licensed voice catalog
  • real-time rendering fleet
  • moderation, consent verification, and enterprise rights

Why people still pay

People still pay for HeyGen because customers pay for output quality, production speed, licensed voices, consent workflows, and a provider that carries the operational risk. The recurring cost buys model licensing, consent records, impersonation risk, watermarking, GPU queues, media storage, abuse response, and rapid model changes, not just the visible interface.

Your build guide

The stack, security requirements, and agent rules for a focused replacement.

Before you start

  • local TTS model
  • ffmpeg
  • GPU recommended
  • voices the user has rights and consent to use
01
Use Python 3.12, FastAPI, SQLite, ffmpeg, and a user-owned local TTS model.
02
Model source audio, processing jobs, transcript or edit segments, and exported media; store source IDs and timestamps for each.
03
Interface for this synthetic voice, dubbing and AI avatars workflow: a local web page with input, progress, review, and export views.
engineering roadmap

Implementation plan

1

Phase 1, architecture and data

Use Python 3.12, FastAPI, SQLite, ffmpeg, and a user-owned local TTS model. Model source audio, processing jobs, transcript or edit segments, and exported media; store source IDs and timestamps for each.

2

Phase 2, implement

Generate speech from text with voice, speed, pause, pronunciation, and segment controls.

3

Phase 3, implement

Create a timeline for audio, captions, uploaded visuals, and simple transitions.

4

Phase 4, review and output

Store prompts, model identifiers, consent notes, and output hashes in a local provenance log. Embed project metadata and a visible synthetic-media disclosure in exported assets.

5

Phase 5, recovery and acceptance

On invalid input or interrupted processing, keep the original record, show the failed step, and permit a safe retry. Verify this invariant with a saved fixture: An invalid input or interrupted operation must retain the source and show a recoverable state; exported records must reload with the same IDs. State the practical limit: proprietary avatars, lip sync, translation, rendering, templates, and rights operations.

the pro prompt
Build me a focused synthetic voice, dubbing and AI avatars workflow for the personal core of HeyGen. Requirements:

- Use Python 3.12, FastAPI, SQLite, ffmpeg, and a user-owned local TTS model. Model source audio, processing jobs, transcript or edit segments, and exported media; store source IDs and timestamps for each.
- Paid product context: Assemble a labeled avatar-style video from user-owned media without cloning real people. Build only this DIY scope: Turn user-authored text into a clearly labeled avatar-style video assembled from user-owned media with a local model, never cloning real people, and retain provenance for every output.
- Generate speech from text with voice, speed, pause, pronunciation, and segment controls.
- Create a timeline for audio, captions, uploaded visuals, and simple transitions.
- Store prompts, model identifiers, consent notes, and output hashes in a local provenance log.
- Use a local web page with input, progress, review, and export views. Required input or access: local TTS model; ffmpeg.
- Recovery: On invalid input or interrupted processing, keep the original record, show the failed step, and permit a safe retry.
- Acceptance: with one labelled sample, show the input, saved intermediate state, and exported result; verify this invariant: An invalid input or interrupted operation must retain the source and show a recoverable state; exported records must reload with the same IDs.
- Out of scope: proprietary avatars, lip sync, translation, rendering, templates, and rights operations; frontier voice or avatar model. Keep this a personal, inspectable workflow.
- Include a README with setup, a sample input, required keys or permissions, data location, and the supported scope.

$ open in your agent (prompt prefilled, you press enter), copy the prompt or copy AGENTS.md · generated from this app's build plan

prior art · use these instead of building, if you'd ratherPiperActive community continuation of the fast local Piper text-to-speech engine.↗
share on X ↗

HeyGen pricing

planmonthlyannual (per mo)what you get
free$0$03 videos/month; 1 minute/video; 1 custom Digital Twin; 30+ languages; up to 3 photo avatars; 1 voice clone.
creator$29$24600 credits/month; 30-minute video max; 1080p; 1 seat; unlimited photo avatars.
pro$49—Base tier: 1,000 credits/month; 30-minute video max; 4K; 1 seat.Same features scale through credit tiers up to 100,000 credits/month at $4,300/month; exact annual base price is not exposed.
business$149/workspace—1,500 credits/month; 60-minute video max; 4K; 5 custom Digital Twins.Additional seats are $20/seat/month; exact annual effective price is not exposed.
enterprise——Flexible generation, no video-duration maximum, 4K, multi-workspace controls, custom seats, and invoice billing.Custom price.

free tier3 videos/month; maximum 1 minute each; 1 custom Digital Twin; 30+ languages; up to 3 photo avatars

billingmonthly + annual; only Creator's annual effective amount is publicly stated; Enterprise is custom

hidden costsCredits are feature-metered: Avatar III 3 credits/minute, Avatar IV/V 20, audio dubbing 2, lip-sync translation 5, and Video Agent 20. Monthly credits roll one extra month; annual credits accumulate until renewal. Business top-ups/auto-reload and $20 seats cost extra.

pricing sources checked 2026-08-12 · pricing source ↗

Questions about HeyGen

Can you build your own HeyGen with AI?

A full replacement is not the recommended project. A consolation build is possible, but the paid product's decisive value sits outside a solo rebuild. For HeyGen, assemble a labeled avatar-style video from user-owned media without cloning real people. The hard boundary is proprietary avatars, lip sync, translation, rendering, templates, and rights operations, plus models, compute, rights, and safety operations.

What does the HeyGen build prompt cover?

The prompt starts with this scope: Turn user-authored text into a clearly labeled avatar-style video assembled from user-owned media with a local model, never cloning real people, and retain provenance for every output. Full-product capabilities excluded from the comparison include: proprietary avatars, lip sync, translation, rendering, templates, and rights operations; frontier voice or avatar model; licensed voice catalog. Follow the implementation plan and its prerequisites before expanding the build.

How do I use the prompt, AGENTS.md and agent skills?

Start with the HeyGen prerequisites and stack, then copy the prompt into your coding agent. Save the project rules as AGENTS.md in the project root. Linked skills are optional packages or source instructions for specific tasks; review their current contents and install only those matching the chosen stack. A skill does not supply API credentials or verify the finished app.

How long will this HeyGen project take?

The catalogue estimate is closest consolation build: one sitting for the limited scope. Setup, integration approvals, debugging, deployment and ongoing maintenance can add time. This is an estimate, not a delivery guarantee.

What would I give up by replacing HeyGen?

proprietary avatars, lip sync, translation, rendering, templates, and rights operations; frontier voice or avatar model; licensed voice catalog; real-time rendering fleet; moderation, consent verification, and enterprise rights. People still pay for HeyGen because customers pay for output quality, production speed, licensed voices, consent workflows, and a provider that carries the operational risk. The recurring cost buys model licensing, consent records, impersonation risk, watermarking, GPU queues, media storage, abuse response, and rapid model changes, not just the visible interface.

What price is this guide comparing against?

The recorded Creator plan is $29/mo (monthly), checked 2026-07-31. Check the linked pricing source before buying. Building your own also has hosting, API and maintenance costs; the recorded amount is not a guaranteed saving.

What can I use instead of building HeyGen?

The prior-art section lists Piper as starting points. Review their current scope, license and maintenance before adopting one.

Every week, more subscriptions die.

New verdicts, new prompts, the week's most-doomed apps.
One email. Unsubscribe in one click.

last week:100 Questions · KINDA1of10 · KINDA1Password · KINDA+1090 more

free forever · no scanner spam · the prompt stays on the site, the deaths come to you

$weekly: what got a verdict, what died.