Fish Audio has raised $52 million in seed funding as the voice AI company moves beyond text-to-speech and voice cloning toward a broader audio-native AI stack. Coreline Ventures and Capital Today led the round, with participation from 359 Capital, Play Time, HF0, 645 Ventures, Parable, Carya Venture Partners, Alphalist Partners and angel investors. The financing arrives only a year after Fish Audio emerged from an open-source voice project and gives the company substantial capital to expand models, enterprise operations and developer infrastructure. The company says it has already reached $21 million in annual recurring revenue and more than 8 million users across creators, developers and enterprises.
Fish Audio Turns Open-Source Voice Into AI Infrastructure
Fish Audio’s trajectory began with a relatively small technical problem. Co-founder and Chief Scientist Shijia Liao, a former NVIDIA video researcher and longtime VTuber and anime fan, wanted to overcome the flat and mechanical quality that defined many synthetic voices. He trained early models on a single gaming GPU, then released the resulting Fish Speech project as open source. The project gained more than 31,000 GitHub stars and attracted developers, game designers and creators who wanted more expressive control over generated speech.
That early community now forms part of a much larger commercial platform. Fish Audio says its technology can clone a voice from a five-second recording in roughly 15 seconds and supports more than 83 languages. Its systems also offer word-level control over emotion and more than 15,000 natural-language controls, giving developers a way to direct delivery rather than simply generate spoken words. The shift matters because voice AI increasingly sits inside products rather than operating only as a standalone media-generation tool.
The company’s latest S2.1 Pro model illustrates that change. Fish Audio says the model is preferred by nearly 67% of listeners over leading competitors in blind listening tests, while enterprise customers can use on-premises deployment, zero-data-retention policies and HIPAA-compliant configurations. Those capabilities move the product closer to infrastructure for production AI systems, where latency, control, security and deployment flexibility can matter as much as raw voice quality. For enterprise buyers, the distinction is important because voice generation increasingly connects directly to customer interactions, digital agents and other business workflows.
$52 Million Backs A Broader Audio Model Strategy
The new capital signals that Fish Audio does not intend to remain a conventional text-to-speech company. The startup plans to expand beyond speech generation into voice-native large language models, speech-to-speech capabilities and other parts of an audio-native stack. It also plans to strengthen its enterprise sales organization and expand developer tooling and integrations with companies including LiveKit and Retell. The strategy places model development and inference infrastructure at the center of the company’s next phase.
This is a significant change in positioning. Traditional text-to-speech systems primarily convert written language into spoken output, while newer voice systems need to understand timing, emotion, conversational context and interaction. Fish Audio is therefore competing for a larger role in the AI application stack, where voice can become an interface between users and models rather than simply the final output layer. The opportunity extends across AI avatars, games, customer-service agents, dubbing, entertainment and other applications that require real-time speech.
Fish Audio’s existing customer base reflects that broader market. The company identifies HeyGen, Retell, LiveKit, OpenArt, Telnyx and Sanas among organizations using its platform, while its models remain available for self-hosting by teams that want greater control over data and infrastructure. That combination gives the company exposure to both developer-led adoption and enterprise deployments. It also creates a path for Fish Audio to compete on infrastructure economics as voice workloads become more demanding.
Enterprise Voice AI Is Becoming A Different Market
For enterprise customers, realistic speech alone is no longer enough. AI avatars may prioritize natural and visually convincing voices, while gaming companies need character-specific expression and voice-agent platforms need low latency alongside conversational realism. Rissa Cao, Fish Audio’s co-founder and CEO, described those differences in the company’s customer requirements: “Every enterprise has different use cases and different preferences. For example, companies like HeyGen, which use our voices to power AI avatars, want realism in voices; a gaming studio would want expressive voices for their characters; and voice agent companies like LiveKit want more natural-sounding and low-latency voices that are expressive enough for calls.”
That diversity could shape the economics of the voice model market. A single general-purpose model may not satisfy every workload, especially when customers trade off latency, quality, controllability, privacy and cost. Fish Audio’s emphasis on natural-language controls and multiple deployment models gives it room to position voice generation as an adaptable infrastructure layer rather than a fixed API. The company’s next challenge will be turning that technical flexibility into durable enterprise contracts and predictable model economics.
Cao framed the company’s original objective in similarly direct terms: “We built Fish Audio because we wanted voice AI that sounded human, not chunky or robotic, and we wanted that quality to be accessible at any scale.” She added: “We make high-quality, human-sounding voices available to every user, from beginner creatives to million-dollar enterprises, so communication is not only more efficient, but more trustworthy. We’ve always believed that if we kept making the models better, people would notice. Eight million of them did. There’s a lot of work left to do, and now we have the resources to do it.”
S2.1 Pro Pushes Fish Audio Toward Developers
Fish Audio is also using its anniversary to lower the barrier for developers testing its newest model. The company plans to make S2.1 Pro available free through its API through the end of August, giving developers an opportunity to build applications around the model without an initial usage cost. That move could help Fish Audio expand the developer ecosystem around its technology while collecting more real-world feedback from production workloads. It also fits the company’s broader strategy of using open-source development and hosted infrastructure together.
The approach reflects a larger shift in AI infrastructure. Model companies increasingly need developers to build the application layer around their technology, while developers want models that provide enough control to differentiate their products. Fish Audio’s open models, hosted services and enterprise deployment options give it several routes into that ecosystem. The company can therefore compete not only for end users but also for the infrastructure position underneath voice-enabled applications.
Fish Audio’s technical history gives it an unusual foundation for that expansion. Fish Speech began as an open-source project rather than a conventional enterprise product, allowing developers to experiment with the technology before the company built a larger commercial platform around it. The company now says its platform serves more than 8 million users, while its open models support real-time text-to-speech, voice cloning and voice agents across more than 83 languages. That progression from community project to commercial infrastructure is central to the company’s current funding story.
Investors See Voice As An AI Interface
Coreline Ventures Managing Partner Osuke Honda views the opportunity as larger than synthetic speech alone. “Voice is becoming the default interface for AI, and Fish Audio is unlocking this opportunity to a new generation of creators, developers, and enterprises,” Honda said. He added: “In its short history, Fish Audio has built an unbeatable track record of pushing the envelope on performance, multilingual support, emotional expression, and cost. All factors that have quickly made Fish Audio the default choice for creators, developers, and now enterprises globally, and we expect them to continue to lead the way.”
The investment arrives as competition across AI voice infrastructure intensifies. Companies are racing to make generated speech faster, more expressive and cheaper while also supporting increasingly complex agentic applications. Fish Audio’s funding gives it the capital to pursue that competition across several layers at once, from foundational models and inference to APIs, enterprise sales and integrations. That breadth could prove more important than any individual voice-generation benchmark as customers begin building voice directly into AI products.
Fish Audio’s Next Test Is Scale
The $52 million round gives Fish Audio considerable room to accelerate after an unusually fast first year. Reaching $21 million in annual recurring revenue and more than 8 million users within that period provides evidence of early market demand, although the company’s next phase will require it to translate rapid adoption into sustainable infrastructure economics. The company must also maintain model quality as workloads expand across creators, developers and enterprise customers. Its ability to serve those segments without sacrificing latency, reliability or control will determine how far the platform can move up the AI infrastructure stack.
For now, Fish Audio is betting that voice will become more than an output format for AI. Its funding plan points toward systems that can generate, transform and understand audio while supporting real-time interaction across a growing range of applications. If that strategy works, the company’s competitive position will depend less on making an AI voice sound convincing in isolation and more on becoming the infrastructure layer that allows other AI products to speak, listen and respond naturally. The $52 million seed round marks the beginning of that larger ambition rather than the culmination of the company’s first year.
