ai lab
Canopy Labs
Canopy Labs matters because Orpheus made expressive, streaming speech generation available as weights, training code, and adaptable checkpoints rather than only as a hosted voice API. Independent evaluation found it competitive among open speech systems, while the project's local-inference and fine-tuning paths lowered the barrier to building customized voices.2,6,7
Profile
Overview
A company represented by one public model program
Canopy Labs is an AI company building interactive models for digital humans and avatars. Its public research record centers on Orpheus, an open-weight text-to-speech model family released in March 2025. The released artifacts are narrower than the company's current ambition: speech-model weights, adaptation code, inference software, and deployment examples. Public information about financing, ownership, and internal organization remains limited, so this dossier separates the documented Orpheus program from broader avatar work that has not yet appeared as a public model release.1,2
Speech generation through a language-model backbone
Orpheus adapts a Llama 3.2 3B language-model backbone to generate speech tokens. Canopy published both a pretrained checkpoint, trained on more than 100,000 hours of English speech, and a production fine-tune. The repository code is Apache 2.0, while the published weights inherit the Llama 3.2 Community License. The model streams audio as it is generated and accepts inline cues such as laughter, sighing, coughing, and yawning. It can also condition on prior text-and-speech pairs for voice matching, although the project cautions that this was not an explicit zero-shot voice-cloning training objective.2,3,8
Weights, training tools, and multilingual adaptation
The release was designed for adaptation rather than demonstration alone. Canopy provided data-preparation notebooks, sample datasets, fine-tuning code, a Python streaming package, a llama.cpp path for CPU inference, and a research-preview collection covering seven additional languages. The multilingual release included both pretrained and fine-tuned checkpoints and a guide explaining how the team extended the English system. Those materials made Orpheus a practical base for local speech experiments and downstream ports.2,4,5
Independent evidence is stronger than the company record
Independent evaluation gives the project more substance than its company biography alone would suggest. EmergentTTS-Eval, published in the NeurIPS 2025 datasets and benchmarks track, recorded the highest overall win rate for Orpheus among the open systems it evaluated. Later streaming-speech research continued to use Orpheus as an open comparison system, while Vapi's Humanness Index ranks it through blind listener votes against hosted voice models. The evidence supports Orpheus as a meaningful open speech release, but not a claim that Canopy has already released a broad digital-human foundation model.6,7,8
Notable contributions
- 01Inline control of vocal expressionOrpheus exposes laughter, sighs, coughs, yawns, and related vocal events as tags inside ordinary text prompts. The contribution is a simple control surface in an open speech model, not the invention of expressive synthesis.2,3
- 02An adaptable open speech training stackCanopy released a pretrained checkpoint, a production fine-tune, preprocessing notebooks, sample data, and training scripts so developers could adapt Orpheus to a voice or language instead of relying only on a fixed demo.2,4
- 03Streaming and local deployment pathsThe project paired model weights with streaming inference, CPU-oriented llama.cpp examples, and an optimized Baseten deployment, covering experimentation and production serving without making the hosted route mandatory.2,5
Sources · 9+−
- 1Canopy LabsCanopy Labs · primary ↗
- 2Orpheus TTSCanopy Labs on GitHub · primary · Mar 19, 2025 ↗
- 3Orpheus 3B 0.1 fine-tuneCanopy Labs on Hugging Face · primary · Mar 19, 2025 ↗
- 4Orpheus multilingual research releaseCanopy Labs on Hugging Face · primary ↗
- 5Canopy Labs selects Baseten for Orpheus TTS inferenceBaseten · primary · May 19, 2025 ↗
- 6EmergentTTS-Eval: Evaluating TTS Models on Complex Prosodic, Expressive, and Linguistic ChallengesNeurIPS 2025 Datasets and Benchmarks Track · paper ↗
- 7Streaming Sequence-to-Sequence Learning with Delayed Streams ModelingarXiv · paper · Sep 10, 2025 ↗
- 8Canopy Orpheus on the Humanness IndexVapi · independent ↗
- 9Canopy Labs AI Limited officersUK Companies House · independent ↗