Funding

Fish Audio Raises $50M Seed for AI Voice Models

Fish Audio, a Palo Alto-based startup building AI voice models, has raised $50 million in a seed round led by Coreline Ventures and Capital Today. The company offers over 15,000 natural language controls and serves more than 8 million users, generating $21 million in annual recurring revenue. It plans to use the funding to develop more advanced models and expand enterprise offerings.

Neura News

Neura News

Neura Market Editorial

July 28, 20265 min read
Fish Audio Raises $50M Seed for AI Voice Models

The market for AI-generated voice models is large. Creative users need voices that sound more expressive. Enterprises automating customer support and sales operations require models that are easier to steer.

Fish Audio, based in Palo Alto, aims to serve all of these needs with its library of more than 15,000 natural language controls. Since launching last year, the startup now has over 8 million people using either the open-source or hosted versions of its models. The company generates annual recurring revenue of $21 million.

On Tuesday, the startup announced it has raised $50 million in a seed round. Coreline Ventures and Capital Today led the round. Other investors included 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0.

From a Single GPU Project to a Growing Startup

Fish Audio began as a small project by Shijia Liao, a former researcher at NVIDIA. Liao was frustrated by the lack of expressive synthetic voices available on the market. He trained a voice generation model on a single GPU and open-sourced it. The Fish Speech repository on GitHub now has more than 31,000 stars. It is used by indie developers, video game designers, and creators.

The company has launched five models in the past year. Four are speech generation models, and one is a speech-to-text model. Fish Audio has open-sourced three of its speech generation models. However, its latest S2.1 Pro model is only available through its paid API.

Fish Audio offers paid monthly plans designed for creators and teams. These plans unlock a set number of minutes of generation and include voice cloning features. The company also offers an enterprise version of its APIs and platform. Organizations like HeyGen, Sanas, and Plaud are already using it.

Addressing Diverse Enterprise Needs

"Every enterprise has different use cases and different preferences. For example, companies like HeyGen, which use our voices to power AI avatars, want realism in voices; a gaming studio would want expressive voice for their characters; and voice agent companies like LiveKit want more natural-sounding and low-latency voices that are expressive enough for calls," said Rissa Cao, CEO and co-founder of Fish Audio.

One way the startup built its library of voices was by asking users to submit their own voices for training its models. The company compensates users if their voices are used. This approach caused some trouble a few months ago. Some creators alleged that their voices were uploaded to Fish Audio without their consent. The startup had a DMCA content takedown process in place, but the takedowns took a long time.

Automating the Takedown Process

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

Cao told TechCrunch that the company has now automated the takedown process. Creators can easily submit a short voice sample or a contract to prove that an uploaded voice belongs to them. Their voice will be removed from the platform in less than three minutes, she said.

Still, this does not prevent anyone from uploading an artist's voice without their knowledge. Until the artist finds out, their voice will continue to be used on the platform until they file for it to be taken down.

Oskue Honda, a partner at Coreline Ventures, said a community-driven model only works when creators trust the platform.

"A community-centric approach can only become a durable advantage if creators trust the platform. That means consent, transparency, and attribution must be built into the product rather than treated as afterthoughts. I believe the industry needs to move toward verified voice ownership, clear licensing terms, easy reporting and takedown processes, and eventually revenue-sharing models where creators benefit financially when their voices are licensed or used commercially," he said.

Why Fish Audio Sought Funding

Cao said that when the startup was only offering its product as an open-source project with plans for creators, it was running efficiently and did not need money. However, the company wanted to develop more advanced models and also wanted to accommodate enterprises as investor interest grew. This led it to seek capital.

Looking ahead, Fish Audio plans to release an audio understanding model this year. It is also building a speech-to-speech model.

The speech generation market is crowded. Companies like ElevenLabs, WellSaid, Cartesia, Speechify, Async (previously Podcastle), and Krisp are all competing for creators and enterprise budgets.

According to Rico Mallozzi, a partner at 359 Capital, fine-grained controls for developers and cost-efficient model training will help Fish Audio compete better with big AI labs.

"I think what they've been able to build, state-of-the-art models, with the team they have, compared to some of these other well-funded AI labs or companies, is incredible. It shows their technical acumen in closing the gap between artificial-sounding and human-like voices," Mallozzi told TechCrunch over a call.

Related on Neura Market

More from Neura News