Vibeleaderboard
Heygen

ai lab

Heygen

HeyGen matters because it turned personalized synthetic video into a repeatable business workflow for people who do not operate video models themselves. Avatar V also gives the company a more legible technical identity: identity-preserving reference conditioning, long-form avatar generation, and production inference are treated as one specialist model system rather than a collection of editing features.2,4,3

Profile

Overview

From camera products to generated presenters

HeyGen is a Los Angeles AI video company founded in 2020 by Joshua Xu and Wayne Liang. The founders studied at Tongji University and Carnegie Mellon University before Xu worked on camera products at Snap and Liang worked in product roles at Smule and ByteDance. Xu started the company after leaving Snap, and the pair developed its generated-video product for business users.1,3

A practical avatar workflow

The initial product turned scripts into presenter videos using stock or personalized avatars. By 2023, a user could record a short consent clip and smartphone video, then generate a reusable likeness that spoke new text. HeyGen paired that workflow with generated voices, lip synchronization, and video translation. The combination targeted marketing, training, localization, and instructional content rather than the cinematic text-to-video work pursued by companies such as Runway.2,1

Avatar V exposes the model stack

HeyGen later made its underlying model work more visible. Avatar V conditions a diffusion transformer on the token sequence of a reference video, using sparse reference attention to preserve both appearance and patterns of movement without making computation grow quadratically with reference length. Its technical report also describes staged training, distilled inference, chunked long-form generation, and a consent and moderation system for production avatars.4,5

Distribution is the strongest evidence

Commercial distribution remains the company's clearest advantage. HeyGen reported reaching $200 million in annual recurring revenue in June 2026, twice its level eight months earlier, with more than 30 million users and customers across most of the Fortune 100. Those figures were reported through an independent interview but remain company-supplied. The same report describes a business focused on talking presenters and localized communication, which gives HeyGen a more specific market position than general video generation.3

Notable contributions

  1. 01Productizing smartphone-captured presenter avatarsHeyGen compressed a workflow that once required professional capture and days of processing into a smartphone recording, explicit consent step, and reusable generated presenter. The claim is about productization and accessibility, not the invention of digital humans.1,2
  2. 02Avatar V video-reference conditioningAvatar V conditions directly on long reference-video token sequences so the generated speaker can preserve facial detail, cadence, gestures, and expressions instead of relying only on a compressed identity embedding.4,5
  3. 03Documenting a long-form avatar inference stackThe company documented a serving system that combines chunked generation, bounded-memory decoding, distributed inference, caching, and super-resolution to produce long 1080p avatar videos at production scale.6,4