227 points gmays 2 hours ago 59 comments

andy_ppp 2 hours ago | parent

Wow certainly in English this is incredibly accurate I tried to break it and it understood me perfectly!

I know it's slightly off topic but surely it must be easy by now to train a spell checker that doesn't annoy the crap out of everyone using it (looking at you here Apple)!

tecleandor 2 hours ago | parent

Spanish is not good (seems to write non existing words and/or with terrible typos...) but English seem to work good even with my (Spanish) accent...

kaoD 2 hours ago | parent

Spanish from where? Here (Castilian Spanish) it seemed to work fine.

chilicuil 1 hour ago | parent

Mexican and venezuelan aren't detected correctly

tecleandor 1 hour ago | parent

Madrid. But it will only work properly if I'm clearly dictating with a very regular rhythm (ViaVoice dictation, if anyone remembers...). If I use a more natural/conversational rhythm (no slang, no abbreviations...) it easily confuses words.

saturn8601 2 hours ago | parent

Initial tests make this feel just like iPhone's terrible text to speech. It is the one thing I utterly hate about iPhone. Ive tried apps that try to embed themselves into the iPhone keyboard and they always don't work out well. Hopefully this gets better and we can somehow get it into the iPhone more seamlessly.

MayeulC 1 hour ago | parent

Speech to text I assume? Maybe it has to do with your a accent or pronunciation? You could contribute a bit to Mozilla's Common voice, if that's the case. I assume it is part of every STT training corpus.

saturn8601 1 hour ago | parent

Yes sorry Speech to Text. I have a standard US East coast accent but sometimes I speak a little mumbly. When I made an effort to speak more slowly and with a cleared throat there was some improvement but still not writing all words.

INTPenis 2 hours ago | parent

I don't think the challenge with speech to text was size of the binary. In my experience the challenge is understanding my 84 year old Croatian father with a sagging mouth after a stroke, when he's trying to write his autobiography.

I just setup Windows speech to text for him last week and it's great to see how he can write an entire page in 10 minutes, it would take him days using the keyboard.

But every single sound he makes with his mouth ends up on the page too.

ComputerGuru 2 hours ago | parent

Sorry about your father. He needs a dictation model, not a general purpose speech-to-text model. They ignore umms and ahhs, change things like “an elephant, no a monkey, went up the tree” to “a monkey went up the tree,” support saying punctuation aloud sometimes, etc.

Gemini team just released Gemini 3.5 Transcribe that’s supposed to be good at this; it’s available via api: https://blog.google/innovation-and-ai/models-and-research/ge...

cgbur 1 hour ago | parent

For essentially infinite and fast dictation I use https://github.com/cjpais/Handy on Parakeet streaming (cohere is far better, but slower and has a token output limit so you cant ramble for many minutes). And then just do a cleanup pass with a cheap LLM, it will in my experience, do far better than trying to voice control to go edit a sentence or change words. I just weave instructions into my writing. I understand this requires technical know-how, but for those with it, this is the best solution I have found to long form writing without my hands.

Imustaskforhelp 43 minutes ago | parent

+1 for handy and then using LLM's for the cleanup pass, though what are your observations on feeling as if sharing that output though?

Because I have seemingly mixed opinions on it, on one hand, I did put the effort but on the other, the output is AI generated so I am unsure about sharing it with others (because they might think its AI generated)

Do you use it for very small edits (removing just the uhhm's?) or for slightly more edits.

The way that I use it sometimes is that while thinking, I will write something which can sometimes make me feel as if a better re-write can better explain my thoughts or rephrasing it as such. For example. I will think about X topic, connect it to Y, then try to add some more points about X again.

I found LLM's to do a really decent job at generating the final outputs as such, but as I said, I am left sometimes feeling a little confused as to sharing it or not because of it being AI generated and the end user not knowing if I put an actual effort into creation of it or not.

Should I try to share the actual transcript of it as well, I really wish if some good ethics and internet ettiquette could be established about it.

asa123 7 minutes ago | parent

what do you mean:

"what are your observations on feeling as if sharing that output though?"

and "Should I try to share the actual transcript of it as well, I really wish if some good ethics and internet ettiquette could be established about it."

could you rephrase the question

for STT, it's literally you saying it, with a model transcribing, and then another model correcting a little, and you also have control over editing it. I don't think any of the arguments on "etiquette re: sharing AI output" apply here.

dv35z 21 minutes ago | parent

Another plug for Handy, and wanted to share something cool about it.

You can set it "Push to talk" mode (like a walkie-talkie radio), and when you're done talking and release the button, it can paste the text into any text field.

You can even replicate ChatGPT voice conversation mode, by having Handy as your speech input, and then (I forgot the extension) enabling a speech-to-text model for OpenCode. Surprisingly relaxing flow for certain tasks, like tweaking a website's styles.

nvtop 1 hour ago | parent

I've been using Gemini Desktop App purely for dictation. It's a miracle! For the first time in my life I'm blown away by the quality of my (heavy accent) speech recognition. Just be sure to disable the "speak to window -> reasoning" option to make it purely dictation and stop from writing whole emails for you.

boplicity 1 hour ago | parent

There are many different challenges, each requiring their own solution. I, for one, really miss the old Google Assistant on my Android phone. It would very reliably play most songs that I wanted to hear on Spotify. Gemini fails at this almost every time, and is significantly slower. It's actually a difficult problem, as the songs people want to hear are regularly being released, are often associated with uncommon names, or have words in unusual orders, so normal LLM style tools just don't cut it.

xp84 49 minutes ago | parent

This reminds me of how the thing iPhones had pre-Siri (so we're talking pre-2010), which was entirely offline, did a better job than even the most modern thing at "Play [one of the finite set of songs in my library]." I sometimes get absurd matches from bands I've never heard of, when the right answer is something right there in my library.

testycool 1 hour ago | parent

Unrelated: I love your username.

yymir 1 hour ago | parent

i mean for something this small, it can be fit into a l3 cache on a cpu and be essentially always on various purposes

yu3zhou4 28 minutes ago | parent

I hope you will find a solid solution for your father.

I was researching STT for people with speech disorders two years ago and essentially everything was boiling down to three problems at the end of the day - data scarcity, irregularity of way of speaking and thus constant ambiguity in translation, and individual differences in speech patterns among patients.

joewhale 2 hours ago | parent

I initially read this as whistle to text, which would be way cooler.

mejutoco 1 hour ago | parent

A good project for training an llm

https://en.wikipedia.org/wiki/Silbo_Gomero

jasonwatkinspdx 1 hour ago | parent

I've met folks that descend from the Zapotec in southern Mexico, and they still use whistling language to talk to each other across mountain valleys.

stymaar 57 minutes ago | parent

There are still people speaking a whisking language in the south-west of France[1], not many but it's being taught in school there to preserve it[2]

[1]: https://fr.wikipedia.org/wiki/Langage_siffl%C3%A9_d%27Aas

[2]: https://www.dailymotion.com/video/x4lxhsi

charv 1 hour ago | parent

Just like Marvel's Yondu!

armcat 2 hours ago | parent

Those are insane benchmarks at this size. Well done!

mrkn1 2 hours ago | parent

love seeing more sub-20MB, CPU-first models. if anyone wants a CLI built on the same ethos (no GPU, no cloud), been using yapsnap streaming Zipformer ASR, plus diarization and timestamps all on CPU! It supports 10 languages. Unlimited transcription for free.

kamranjon 1 hour ago | parent

Sooo I haven't really been super impressed with the needle models before, but this is very impressive. It transcribed multiple sentences I gave it with complex timing and words and in such a small footprint, I'm super impressed. Excited to see what types of things can be built with something like this, the performance seems very good.

aidotguru 1 hour ago | parent

eager to see if working in android phones

rpdillon 1 hour ago | parent

FUTO keyboard (open-source, free) runs entirely on-device and has extremely good STT accurary, especially with the 70M parameter model. I've used it for years now and love it.

https://futo.tech/

Edit: As others have pointed out, this is not actually open source. It's source-available, which is quite a bit different because folks can't fork and distribute it as easily. The license also appears to be revocable and non-transferable, which makes it different from open source licenses.

robertlane0 1 hour ago | parent

I'd classify it as "source-available" given the noncommercial clause in the license.

https://github.com/futo-org/android-keyboard/blob/master/LIC...

rpdillon 51 minutes ago | parent

Good call out. I tend to be very sensitive about these things, and this is a case where I messed up. Thanks for the correction.

lrvick 1 hour ago | parent

Futo only produces source-available proprietary software. They most certainly are not Open Source, though they unfortunately lied about this a lot before they got called out enough times.

https://github.com/futo-org/voice-input/blob/master/LICENSE....

rpdillon 51 minutes ago | parent

Good call... I got that one wrong. Thanks for the pointer!

helterskelter2 1 hour ago | parent

I've used FUDO keyboard for a long time, but I never tried using the voice input, so I'm testing it now. Let's see how well it transcribes everything.

...Okay that was pretty good.

mo2art 1 hour ago | parent

RuntimeError: audio limit is 30 s

jjice 1 hour ago | parent

Did you record over 30 seconds?

agilek 1 hour ago | parent

Can we have more languages?

mrkn1 1 hour ago | parent

yapsnap has 10 languages

jayshah5696 1 hour ago | parent

This is actually a really great release. Congratulations team. I just tried few words. My Indian accent also was able to pick up.I'm gonna run it on my Linux Box.

albert_e 1 hour ago | parent

What the demo does not do is show streaming output of transcribed text as we are speaking and recording (before we hit stop). That is an essential feature IMO for most general purpose live STT apps.

iforgotmypasswo 58 minutes ago | parent

This is the main reason I lean on Deepgram over local services.

solarkraft 40 minutes ago | parent

My mind is boggled by how many implementations miss this.

Handy has Nemotron Streaming and it works fabulously, FWIW. I’ve vibed a kind-of-working Deepgram API server into it but haven’t gotten around to finishing it. It’s something that should exist IMO!

wkcheng 1 hour ago | parent

How does this compare with Parakeet? I've been using that locally in my projects on an M-series macbook and it's been working great. It's fast and accurate enough for my use cases (meeting transcription, audio transcription for demo videos, etc.)

This definitely seems lighter and faster. How does accuracy compare?

theturtletalks 1 hour ago | parent

Parakeet is the gold standard. With models like moonshine and koroko (TTS model), it’s more about embedding the model in the application itself. If you’re using Parakeet, embedding it in the application is not feasible.

I use parakeet with superwhisper, and I’m making another app that has SST and TTS built in, and I want to use my downloaded parakeet model, but it seems there’s so many different implementations from ONNX to whisper, it’s not easy to use your downloaded models. So models like moonshine and this one allow you to just embed it into your application simply. It might not be as good as parakeet, but it gets you 80% of the way there.

jwr 20 minutes ago | parent

People keep praising Parakeet, but I've found it to be worse than Whisper Large. Yes, it is much, much faster and smaller, but accuracy matters a lot if you are to use dictation regularly and seriously.

I ended up having AI optimize Whisper Large and create a plugin for TypeWhisper, and that's what I use (feeding the results through local Qwen 3.8 running under MTPLX).

e12e 1 hour ago | parent

Hm. I saw language=detect and tried some Japanese - which (given the actual list of supported languages) unsurprisingly turned into some mangled Spanish.

Since it doesn't support Norwegian - I tried English - and it mis-transcribed "cleaning" for "training" - probably a failure due to context/training (Hello everyone, today we are going to do some cleaning).

So, reasonable, but limited?

rafaelm 1 hour ago | parent

Huh, this was really confusing. I already had an STT app called Whistle on my Android phone.

Centigonal 1 hour ago | parent

This is quite good, especially given the size and the fact that it runs on the CPU.

zimpenfish 52 minutes ago | parent

Tried it on a random TV episode and it seems to get stuck sometimes where it just outputs "Thank you." as a default - at one point emitting that for 60s of dialogue (and no, the episode does not have someone repeating "Thank you." for 60s.) Happens several times during the transcription.

jwr 38 minutes ago | parent

It's funny how so many models tend to generate "Thank you. Don't forget to subscribe" or "Thanks for watching" if there is silence. Shows you what they've been trained on :-)

MisterMunchkin 45 minutes ago | parent

For something you could plausibly ship inside a webapp, it’s very good. I can definitely think of some cool uses for this.

thomkaar 44 minutes ago | parent

i can't even tell when i say my booty thick

thomkaar 44 minutes ago | parent

it can't even tell when i say "my booty thick"

sfpk 38 minutes ago | parent

If you need speech2text try this, https://ccoreilly.github.io/vosk-browser/

pzo 25 minutes ago | parent

tested in polish and unless you speak very loud, clear and slow is not that good, parakeet definitely better.

try-working 14 minutes ago | parent

I built an STT plugin for DeepSeek Harness that uses this Whistle model as well as a larger one from Desert Ant Labs: https://github.com/try-works/dsh-stt

rshemet 10 minutes ago | parent

hey, Roman here from Cactus, thank you for the feature!

opening this thread for questions/feedback if you have any