ai lab
Also indexed as fishaudio
Fish Audio
Fish Audio matters because it links an adopted open-weight speech project to a fast-growing hosted voice platform and publishes unusually concrete architecture and inference details for a young commercial vendor. Its progress also exposes the governance problem at the center of voice cloning: expressive capability is not durable unless identity, consent, attribution, and removal work at platform speed.5,1,2,6
Profile
Overview
From open project to voice company
Fish Audio is the voice AI product and company brand operated by Hanabi AI Inc. in Palo Alto. It grew out of Fish Speech, a source-available text-to-speech project started by former Nvidia researcher Shijia Liao. Independent reporting says Liao first trained the system on a single GPU after finding existing synthetic voices insufficiently expressive over longer passages. Rissa Cao joined him as co-founder and chief executive, and the project became the basis of a hosted product for creators, developers, and enterprise customers.1,3
A documented speech architecture
The original Fish Speech research describes a multilingual text-to-speech system built around a slow and fast dual autoregressive architecture. The design separates higher-level linguistic structure from acoustic detail and avoids a mandatory phoneme front end. The later S2 report extends the system with natural-language instructions for emotion and delivery, multi-speaker and multi-turn generation, streaming inference, and a staged data and reward-modeling pipeline.4,5
Open research and paid deployment
Fish Audio has kept a mixed access model rather than treating every release as either fully closed or fully open. The company released code and weights for Fish Speech and S2 under research-oriented terms, while S2.1 Pro is available through a paid API. TechCrunch reported in July 2026 that the repository had exceeded 31,000 GitHub stars, the company had more than eight million users across open and hosted versions, and it had raised a $52 million seed round after launching five speech models in roughly a year.7,8,1
Consent is a product requirement
Voice cloning also makes consent part of the technical and commercial record. TechCrunch reported complaints that some creator voices had been uploaded without permission, and the performers' union Equity separately demanded removal and disclosure after receiving complaints from members. Fish Audio said it automated takedowns and could remove a disputed voice quickly after receiving a sample or contract, but the process remains reactive because an unauthorized upload can be used until its owner discovers it. That unresolved ownership problem is material to evaluating a platform whose catalog is partly built from user-submitted voices.1,2
Notable contributions
- 01Fish Speech's dual autoregressive designFish Speech used separate slow and fast autoregressive components to model linguistic structure and acoustic tokens while supporting multilingual synthesis without a required phoneme front end.4
- 02S2 instruction-controlled expressive speechS2 made natural-language descriptions of emotion, pace, and delivery part of the model interface and documented multi-speaker, multi-turn, and streaming generation in the same system.5
- 03MIKU-PAL emotional-speech data labelingFish Audio researchers co-authored MIKU-PAL, an Interspeech 2025 pipeline that uses multimodal analysis to extract and label consistent emotional speech from unlabeled video data.6
Sources · 8+−
- 1Fish Audio raises $52M seed to build AI voice models for creators and enterprisesTechCrunch · independent · Jul 28, 2026 ↗
- 2Equity demands Fish Audio removes unauthorised AI voicesEquity · independent · Jun 3, 2026 ↗
- 35 Models, 22 People, 1 YearFish Audio · primary · Jul 27, 2026 ↗
- 4Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech SynthesisarXiv · paper · Nov 2, 2024 ↗
- 5Fish Audio S2 Technical ReportarXiv · paper · Mar 9, 2026 ↗
- 6MIKU-PAL: An Automated and Standardized Multimodal Method for Speech Paralinguistic and Affect LabelingInterspeech 2025 · paper · Aug 17, 2025 ↗
- 7Fish Speech repositoryGitHub · primary ↗
- 8What We Mean by Open Source, and Why It Matters for S2Fish Audio · primary · Mar 12, 2026 ↗