Fish Audio
Fish Audio is a voice AI platform from Hanabi AI that turns text into expressive, emotion-tagged speech and clones a voice from a 10-second sample. Built on the open-source Fish-Speech research line, it markets itself as a faster, dramatically cheaper alternative to ElevenLabs for developers and creators.
The platform crystallized around two model launches: OpenAudio S1 debuted in June 2025, undercutting ElevenLabs pricing 6x while topping the TTS-Arena leaderboard and pushing Hanabi AI's ARR past $5 million, and Fish Audio S2 shipped March 10, 2026 as a fully open-sourced, sub-150ms dual-autoregressive model with inline emotion tags.
Think of Fish Audio as an open-source engine bolted onto ElevenLabs' business model — same voice quality, a fraction of the price.
See nascent terms 7 days before everyone, unlock every stage filter, and get weekly early alerts.
Search Interest
-
Nascent0–7 days
-
Emergent8–30 days
-
Validating31–90 days
-
Rising91–180 days
-
Established ← now180 days +
Why is it emerging now?
Fish Audio open-sourced its S2 model on March 10, 2026 — a sub-150ms, dual-autoregressive TTS system Hanabi AI positions as a free alternative to ElevenLabs, building on S1's 2025 debut that already drove roughly $5M ARR and 20,000 developers at 6x lower prices.
Outlook
6-month signal projection and commercial timeline.
Open pricing and OSS weights keep it the default cheap-TTS benchmark, but Gemini/OpenAI native voice and rivals like KugelAudio erode differentiation.
Risk · Big-lab native TTS (Gemini, OpenAI) and funded rivals like KugelAudio could commoditize its price advantage fast.
Analogs · ElevenLabs · open-source LLMs vs proprietary APIs · Stable Diffusion (open-weights commoditization)
-
nowCheap OSS TTS default
S2 weights are free; hosted API runs ~11x cheaper than ElevenLabs at roughly $5M ARR.
-
3-6moWrapper apps proliferate
ComfyUI nodes, dubbing tools, and voice-cloning apps build on the open S2 weights.
-
6-12moNative TTS squeezes margins
Bundled Gemini/OpenAI voice and funded rivals pressure Fish Audio's price edge.
Competition & Opportunity for term “Fish Audio”
Signals derived from the tracked queries, the term's monetization cards, and its cluster neighbors. Heuristic except where marked measured (Google KD).
Ideas for term “Fish Audio”
Buildable pitches — turn this term into an article, site, product, post, newsletter, video, or course. Steal any card and run with it.
Byte-billing quirk (3x cost for CJK text) and A/B win-rate data are underexplored; high commercial search intent.
Open weights mean a real self-hosting tutorial is possible; most existing coverage only reviews the hosted API.
Sub-150ms latency enables live dubbing; no dedicated build tutorial exists yet.
Byte-billing vs character-billing makes cost comparison non-obvious for non-English text; a calculator fills a real gap.
Wraps the open model into a hosted dubbing product for streamers and course creators who don't want to self-host.
Audio quality claims are easy to demo and highly shareable; blind A/B format drives watch time.
First-person or analysis angle on how a 4-person team's open model reshaped premium voice-AI pricing.
A four-person Gen Z team open-sourced a model that beats a $3B company in blind A/B tests — for free.
Open-sourcing your flagship model sounds like giving up your moat. Fish Audio's revenue says otherwise.
Ten seconds of audio, zero dollars, and a voice clone that passed a blind listening test against a paid competitor.
What People Search
Long-tail queries from Google Suggest + Trends. Volume and competition are heuristics — directional, not audited. Content Type comes from query shape.
SERP of term “Fish Audio”
What searchers see today — organic results on top, paid ads if anyone's bidding. Ad density is a real-time commercial signal.
FAQ
What is Fish Audio?
Fish Audio is a voice AI platform from Hanabi AI that turns text into expressive, emotion-tagged speech and clones a voice from a 10-second sample.
Why is Fish Audio emerging now?
Fish Audio open-sourced its S2 model on March 10, 2026 — a sub-150ms, dual-autoregressive TTS system Hanabi AI positions as a free alternative to ElevenLabs, building on S1's 2025 debut that already drove roughly $5M ARR and 20,000 developers at 6x lower prices.
When did Fish Audio emerge?
Publicly emerged around 2024-05-13 (about 801 days ago as of 2026-07-23). EarlyTerms first recorded a pipeline signal on 2026-07-02.
Related Terms
Other terms in the same space — aliases, subtypes, competitors, and neighbors to explore next.
- Competitor ElevenLabs ElevenLabs is a voice AI company that builds text-to-speech, speech-to-text, voice cloning, and voice-agent infrastructure used by… →
- Competitor Gemini 3.1 Flash TTS Gemini 3.1 Flash TTS is Google DeepMind's text-to-speech model that generates expressive speech in 70+ languages, steered by 200+ audio… →
- Also known as
- Part of
- Includes ·
- Competitor ··
- Related ··
Sources
Primary URLs this report cites — open any to verify the claim yourself.
- 01 Fish Audio — official site fish.audio ↗
- 02 Fish Audio Blog: Open-Sourcing S2 fish.audio ↗
- 03 Fish Audio S2 product page fish.audio ↗
- 04 MarkTechPost: Fish Audio S2 launch coverage marktechpost.com ↗
- 05 StartupHub.ai: Fish Audio S1 vs ElevenLabs startuphub.ai ↗
- 06 fishaudio/fish-speech (GitHub, 31.3k stars) github.com ↗
- 07 TextToLab: Fish Audio pricing breakdown texttolab.com ↗
- 08 Hacker News: Ask HN — What's KugelAudio's (YC P26) Moat? news.ycombinator.com ↗