EarlyTerms

Fish Audio

Established · Emerged · 801 days old · Last reviewed

Fish Audio is a voice AI platform from Hanabi AI that turns text into expressive, emotion-tagged speech and clones a voice from a 10-second sample. Built on the open-source Fish-Speech research line, it markets itself as a faster, dramatically cheaper alternative to ElevenLabs for developers and creators.

The platform crystallized around two model launches: OpenAudio S1 debuted in June 2025, undercutting ElevenLabs pricing 6x while topping the TTS-Arena leaderboard and pushing Hanabi AI's ARR past $5 million, and Fish Audio S2 shipped March 10, 2026 as a fully open-sourced, sub-150ms dual-autoregressive model with inline emotion tags.

Think of Fish Audio as an open-source engine bolted onto ElevenLabs' business model — same voice quality, a fraction of the price.

EarlyTerms Pro

See nascent terms 7 days before everyone, unlock every stage filter, and get weekly early alerts.

Search Interest

peak ~17K/mo
updated 2026-07-21
~17K/mo ~8.7K/mo 0
2026-06-22 2026-07-07 2026-07-21
Term Lifecycle
  1. Nascent
    0–7 days
  2. Emergent
    8–30 days
  3. Validating
    31–90 days
  4. Rising
    91–180 days
  5. Established ← now
    180 days +

Why is it emerging now?

TL;DR

Fish Audio open-sourced its S2 model on March 10, 2026 — a sub-150ms, dual-autoregressive TTS system Hanabi AI positions as a free alternative to ElevenLabs, building on S1's 2025 debut that already drove roughly $5M ARR and 20,000 developers at 6x lower prices.

5 forces driving coverage — scroll →

Outlook

6-month signal projection and commercial timeline.

Signal medium
Revenue strong

Open pricing and OSS weights keep it the default cheap-TTS benchmark, but Gemini/OpenAI native voice and rivals like KugelAudio erode differentiation.

Risk · Big-lab native TTS (Gemini, OpenAI) and funded rivals like KugelAudio could commoditize its price advantage fast.

Analogs · ElevenLabs · open-source LLMs vs proprietary APIs · Stable Diffusion (open-weights commoditization)

Monetization timeline
  1. now
    Cheap OSS TTS default

    S2 weights are free; hosted API runs ~11x cheaper than ElevenLabs at roughly $5M ARR.

  2. 3-6mo
    Wrapper apps proliferate

    ComfyUI nodes, dubbing tools, and voice-cloning apps build on the open S2 weights.

  3. 6-12mo
    Native TTS squeezes margins

    Bundled Gemini/OpenAI voice and funded rivals pressure Fish Audio's price edge.

Competition & Opportunity for term “Fish Audio”

Signals derived from the tracked queries, the term's monetization cards, and its cluster neighbors. Heuristic except where marked measured (Google KD).

Content Gap
24 queries tracked
Led by General (19), Reference (4)
8 Suggest-only tails — long-tail opening
Revenue Potential
0% commercial-intent queries
2 monetization angles mapped
Mostly informational — pre-commercial
Build Difficulty
Very High (heuristic)
Stage: established — late entry — expect incumbents
6 / 12 default TLDs taken · oldest incumbent fishaudio.com (2017-01-01)
2 related terms already published
Heuristic · signals: tracked queries, term monetization cards, cluster neighbors

Ideas for term “Fish Audio”

Buildable pitches — turn this term into an article, site, product, post, newsletter, video, or course. Steal any card and run with it.

Article
Fish Audio vs ElevenLabs: Full 2026 Pricing and Quality Comparison

Byte-billing quirk (3x cost for CJK text) and A/B win-rate data are underexplored; high commercial search intent.

Article
How to Self-Host Fish Audio S2 for Free Voice Cloning

Open weights mean a real self-hosting tutorial is possible; most existing coverage only reviews the hosted API.

Article
Fish Audio S2 API Tutorial: Building a Real-Time Dubbing Pipeline

Sub-150ms latency enables live dubbing; no dedicated build tutorial exists yet.

Product
TTS pricing calculator: Fish Audio vs ElevenLabs vs KugelAudio

Byte-billing vs character-billing makes cost comparison non-obvious for non-English text; a calculator fills a real gap.

Product
Multilingual live-dubbing SaaS built on S2's low latency

Wraps the open model into a hosted dubbing product for streamers and course creators who don't want to self-host.

Video
'I tested Fish Audio S2 against ElevenLabs and Cartesia — blind listening test' — 15-min YouTube comparison

Audio quality claims are easy to demo and highly shareable; blind A/B format drives watch time.

Post
The Open-Source TTS Model That Made ElevenLabs Cut Its Prices

First-person or analysis angle on how a 4-person team's open model reshaped premium voice-AI pricing.

Post HN / r/LocalLLaMA
The Open-Source TTS Model That's Quietly Eating ElevenLabs' Lunch

A four-person Gen Z team open-sourced a model that beats a $3B company in blind A/B tests — for free.

Post Newsletter / Voice-AI Twitter
Fish Audio Went From $400K to $5M ARR in Four Months — By Giving the Model Away

Open-sourcing your flagship model sounds like giving up your moat. Fish Audio's revenue says otherwise.

Post YouTube / Tech media
I Cloned My Voice for Free With Fish Audio. ElevenLabs Should Be Worried.

Ten seconds of audio, zero dollars, and a voice clone that passed a blind listening test against a paid competitor.

What People Search

Long-tail queries from Google Suggest + Trends. Volume and competition are heuristics — directional, not audited. Content Type comes from query shape.

Keyword
Competition
Content Type
fish audio
Very Low
General
fish audio s2
Very Low
General
fish audio s2 pro
Very Low
General
fish audio jjk narrator
Very Low
General
fish audio super smash bros
Very Low
General
fish audio github
Very Low
Showcase
fish audio tts
Very Low
General
fish audio singapore
Very Low
General
1–8 of 24
1 / 3
Updated 2026-07-21 · sources: Google Trends, Google Suggest · Competition is heuristic

SERP of term “Fish Audio”

What searchers see today — organic results on top, paid ads if anyone's bidding. Ad density is a real-time commercial signal.

FAQ

What is Fish Audio?

Fish Audio is a voice AI platform from Hanabi AI that turns text into expressive, emotion-tagged speech and clones a voice from a 10-second sample.

Why is Fish Audio emerging now?

Fish Audio open-sourced its S2 model on March 10, 2026 — a sub-150ms, dual-autoregressive TTS system Hanabi AI positions as a free alternative to ElevenLabs, building on S1's 2025 debut that already drove roughly $5M ARR and 20,000 developers at 6x lower prices.

When did Fish Audio emerge?

Publicly emerged around 2024-05-13 (about 801 days ago as of 2026-07-23). EarlyTerms first recorded a pipeline signal on 2026-07-02.

Related Terms

Other terms in the same space — aliases, subtypes, competitors, and neighbors to explore next.

Explore next
Also mentioned
  • Also known as Fish-Speech
  • Part of text-to-speech
  • Includes OpenAudio S1·Fish Audio S2
  • Competitor KugelAudio·CosyVoice·Cartesia
  • Related Hanabi AI·voice cloning·zero-shot voice cloning

Sources

Primary URLs this report cites — open any to verify the claim yourself.

  1. 01 Fish Audio — official site fish.audio
  2. 02 Fish Audio Blog: Open-Sourcing S2 fish.audio
  3. 03 Fish Audio S2 product page fish.audio
  4. 04 MarkTechPost: Fish Audio S2 launch coverage marktechpost.com
  5. 05 StartupHub.ai: Fish Audio S1 vs ElevenLabs startuphub.ai
  6. 06 fishaudio/fish-speech (GitHub, 31.3k stars) github.com
  7. 07 TextToLab: Fish Audio pricing breakdown texttolab.com
  8. 08 Hacker News: Ask HN — What's KugelAudio's (YC P26) Moat? news.ycombinator.com