227 points gmays 2 hours ago 59 comments
andy_ppp 2 hours ago | parent
I know it's slightly off topic but surely it must be easy by now to train a spell checker that doesn't annoy the crap out of everyone using it (looking at you here Apple)!
tecleandor 2 hours ago | parent
kaoD 2 hours ago | parent
chilicuil 1 hour ago | parent
tecleandor 1 hour ago | parent
saturn8601 2 hours ago | parent
MayeulC 1 hour ago | parent
saturn8601 1 hour ago | parent
INTPenis 2 hours ago | parent
I just setup Windows speech to text for him last week and it's great to see how he can write an entire page in 10 minutes, it would take him days using the keyboard.
But every single sound he makes with his mouth ends up on the page too.
ComputerGuru 2 hours ago | parent
Gemini team just released Gemini 3.5 Transcribe that’s supposed to be good at this; it’s available via api: https://blog.google/innovation-and-ai/models-and-research/ge...
cgbur 1 hour ago | parent
Imustaskforhelp 43 minutes ago | parent
Because I have seemingly mixed opinions on it, on one hand, I did put the effort but on the other, the output is AI generated so I am unsure about sharing it with others (because they might think its AI generated)
Do you use it for very small edits (removing just the uhhm's?) or for slightly more edits.
The way that I use it sometimes is that while thinking, I will write something which can sometimes make me feel as if a better re-write can better explain my thoughts or rephrasing it as such. For example. I will think about X topic, connect it to Y, then try to add some more points about X again.
I found LLM's to do a really decent job at generating the final outputs as such, but as I said, I am left sometimes feeling a little confused as to sharing it or not because of it being AI generated and the end user not knowing if I put an actual effort into creation of it or not.
Should I try to share the actual transcript of it as well, I really wish if some good ethics and internet ettiquette could be established about it.
asa123 7 minutes ago | parent
"what are your observations on feeling as if sharing that output though?"
and "Should I try to share the actual transcript of it as well, I really wish if some good ethics and internet ettiquette could be established about it."
could you rephrase the question
for STT, it's literally you saying it, with a model transcribing, and then another model correcting a little, and you also have control over editing it. I don't think any of the arguments on "etiquette re: sharing AI output" apply here.
dv35z 21 minutes ago | parent
You can set it "Push to talk" mode (like a walkie-talkie radio), and when you're done talking and release the button, it can paste the text into any text field.
You can even replicate ChatGPT voice conversation mode, by having Handy as your speech input, and then (I forgot the extension) enabling a speech-to-text model for OpenCode. Surprisingly relaxing flow for certain tasks, like tweaking a website's styles.
nvtop 1 hour ago | parent
boplicity 1 hour ago | parent
xp84 49 minutes ago | parent
testycool 1 hour ago | parent
yymir 1 hour ago | parent
yu3zhou4 28 minutes ago | parent
I was researching STT for people with speech disorders two years ago and essentially everything was boiling down to three problems at the end of the day - data scarcity, irregularity of way of speaking and thus constant ambiguity in translation, and individual differences in speech patterns among patients.
joewhale 2 hours ago | parent
mejutoco 1 hour ago | parent
jasonwatkinspdx 1 hour ago | parent
stymaar 57 minutes ago | parent
[1]: https://fr.wikipedia.org/wiki/Langage_siffl%C3%A9_d%27Aas
charv 1 hour ago | parent
armcat 2 hours ago | parent
mrkn1 2 hours ago | parent
mrkn1 1 hour ago | parent
kamranjon 1 hour ago | parent
aidotguru 1 hour ago | parent
rpdillon 1 hour ago | parent
Edit: As others have pointed out, this is not actually open source. It's source-available, which is quite a bit different because folks can't fork and distribute it as easily. The license also appears to be revocable and non-transferable, which makes it different from open source licenses.
robertlane0 1 hour ago | parent
https://github.com/futo-org/android-keyboard/blob/master/LIC...
rpdillon 51 minutes ago | parent
lrvick 1 hour ago | parent
https://github.com/futo-org/voice-input/blob/master/LICENSE....
rpdillon 51 minutes ago | parent
helterskelter2 1 hour ago | parent
...Okay that was pretty good.
jayshah5696 1 hour ago | parent
albert_e 1 hour ago | parent
iforgotmypasswo 58 minutes ago | parent
solarkraft 40 minutes ago | parent
Handy has Nemotron Streaming and it works fabulously, FWIW. I’ve vibed a kind-of-working Deepgram API server into it but haven’t gotten around to finishing it. It’s something that should exist IMO!
wkcheng 1 hour ago | parent
This definitely seems lighter and faster. How does accuracy compare?
theturtletalks 1 hour ago | parent
I use parakeet with superwhisper, and I’m making another app that has SST and TTS built in, and I want to use my downloaded parakeet model, but it seems there’s so many different implementations from ONNX to whisper, it’s not easy to use your downloaded models. So models like moonshine and this one allow you to just embed it into your application simply. It might not be as good as parakeet, but it gets you 80% of the way there.
jwr 20 minutes ago | parent
I ended up having AI optimize Whisper Large and create a plugin for TypeWhisper, and that's what I use (feeding the results through local Qwen 3.8 running under MTPLX).
e12e 1 hour ago | parent
Since it doesn't support Norwegian - I tried English - and it mis-transcribed "cleaning" for "training" - probably a failure due to context/training (Hello everyone, today we are going to do some cleaning).
So, reasonable, but limited?
rafaelm 1 hour ago | parent
Centigonal 1 hour ago | parent
zimpenfish 52 minutes ago | parent
jwr 38 minutes ago | parent
MisterMunchkin 45 minutes ago | parent
thomkaar 44 minutes ago | parent
thomkaar 44 minutes ago | parent
sfpk 38 minutes ago | parent
pzo 25 minutes ago | parent
try-working 14 minutes ago | parent
rshemet 10 minutes ago | parent
opening this thread for questions/feedback if you have any