Case file — C969F6D5

SHIP IT
?/10

Launch roast · Launch HN

Launch HN: Speko (YC S26) – OpenRouter for Voice AI by abdik · HN thread

Unsolicited, from public launch info only. Founder? Reply or ask us to take it down.

The idea

“Launch HN: Speko (YC S26) – OpenRouter for Voice AI https://speko.ai/ Founder's Show HN post: Hi HN! I'm Bek, founder of Speko, a platform that finds an optimal combination of speech-to-text, LLM, and text-to-speech models, given your constraints, among all our public benchmarked options, and tells you why. Demo: https://www.youtube.com/watch?v=no2LY2gRh-c Typical production voice agent is an ensemble of three models: STT, an LLM, and TTS. Each of those layers offers a dozen credible vendors, and each month there are new models on the market. Almost everyone evaluates once, picks a stack of their choice, and never rechecks because switching from a vendor to another involves yet another integration and arguments about the numbers. The result is that you use voice agents running last quarter's models while better and cheaper options are available. Before founding Speko, I spent four years as cofounder and CTO building voice agents for enterprises across Asia in 10+ languages. Each time a new speech model would arrive, we repeated the same ritual: hire native-speaking raters, benchmark it against our existing stack, and update production if it improved. Speko turns this process into an API. A team running thousands of calls a day told us: "we can literally go to this dashboard, switch the model, and it will do it for us." How it works: you send a request with your optimization criteria (accuracy, latency, cost or balanced), language and region. The router filters to models which we measured for the given combination of constraints, benchmarks them, selects the winner, and returns a response with headers containing provider, model names, and the scores. The gateway prefetches signed session plans, so a new session dials the provider straight from memory; no control-plane round trip while a caller waits. Failover happens only during connection setup stage: if the provider refuses the connection attempt, we start connecting to the runners-up. Some of the customer stories: one founder came to us not knowing what to pick at all: he gave us his use case and n Landing page content (The Router for Voice AI): > ## Speko page index > > The complete index of every page on this site is at: https://speko.ai/llms.txt > Read it before exploring further. It lists the exact Markdown URL for every canonical HTML page. # The Router for Voice AI Backed by Y Combinator Every speech model, benchmarked language by language, wired into one API. Get API key ## Router A hosted, provider-neutral STT, LLM and TTS data plane at router.speko.dev, with typed contracts and managed routing. ### Benchmark coverage by language Model EN (English) ES (Spanish) DE (German) FR (French) AR (Arabic) FIL (Filipino) NB (Norwegian) HI (Hindi) TA (Tamil) TE (Telugu) KO (Korean) ZH (Chinese (Mandarin)) JA (Japanese) TH (Thai) deepgram:nova-3 openai:gpt-4o-transcribe soniox:stt-rt-v5 cartesia:ink-whisper Alibaba:Qwen3-ASR bench google:chirp_3 OpenAI:GPT-4o-mini Transcribe bench assemblyai:universal-3-5-pro openai:gpt-transcribe Smallest AI:Pulse Pro bench Deepgram:Nova-2 bench xai:grok-stt Alibaba:Qwen-ASR bench ElevenLabs:Scribe v2 bench Gladia:Solaria-1 bench gradium:default modulate:velma-2-stt-streaming ElevenLabs:Scribe v1 bench OpenAI:GPT-4o mini Transcribe bench alibaba:qwen3-asr-flash-realtime elevenlabs:scribe_v2_realtime AssemblyAI:Universal (batch tier) bench azure:MAI-Transcribe-2 cartesia:ink-2 fish:transcribe-1 fish:transcribe-1-pro Hamsa:S3 bench modulate:velma-2-stt-streaming-english-v2 nari:qwen3-asr nari:qwen3-asr-fast smallest:pulse speechmatics:enhanced speechmatics:standard Hover a cell for its number. worse better 20 of 47 are only measured in English Their rank in any other language is unknown — including the model that sits at the top of the English table. 5 different models win across 14 languages No single model is best everywhere, so the right pick changes with the language your users speak. ### Score against cost, per stage Accuracy (WER · conversational) 25ms 4.57s 12.0% 35.2% Qwen3-ASR Fast 12.0% · 33ms Qwen3-ASR 12.0% · 25ms Finalize (ms, p50) Model Qwen3-ASR Fast nari:qwen3-asr-fast 12.0% 33ms Qwen3-ASR nari:qwen3-asr 12.0% 25ms Qwen3-ASR Realtime alibaba:qwen3-asr-flash-realtime 12.2% 889ms Scribe v2 Realtime elevenlabs:scribe_v2_realtime 12.5% 1.75s Gemini 3.5 Transcribe Live gemini:gemini-3.5-transcribe-live 13.1% 1.28s GPT Live Transcribe openai:gpt-live-transcribe 13.5% 620ms Grok Voice Transcribe 1.0 xai:grok-stt 13.7% 1.13s Universal-3.6 Pro assemblyai:universal-3-6-pro 13.7% 505ms Muse”