79 points swolpers 1 hour ago 41 comments
thangalin 52 minutes ago | parent
https://www.youtube.com/watch?v=WAeHgE94rVo
No cloud, no tokens to pay. Reads a book using a full cast of characters. Quotation attribution detection (for my novel) is at 97.2% accuracy (485/499 quotes identified and assigned correctly). The autofill of character voice descriptions uses the prose to determine how the character sounds.
Employs Gemma 4[1] for the prose analysis (voice fills, quotation detection) and Qwen3 TTS Voice Design[2] for creating voice samples. Runs on an 8GB NVIDIA T1000 GPU card, 96 GB RAM, and a AMD Ryzen 5 7600.
[1]: https://deepmind.google/models/gemma/gemma-4/
[2]: https://huggingface.co/spaces/Qwen/Qwen3-TTS-Voice-Design
Multicomp 50 minutes ago | parent
The title of the video is 'KeenLore - Emotive Audiobook Creator Demo' and it appears to be a web UI and some local stack that reads text files.
loremm 46 minutes ago | parent
I understand audiobook narrators often do it, and that's fun. But it's not so critical in my opinion
talon8635 44 minutes ago | parent
simonw 51 minutes ago | parent
I guess voice cloning is widely enough available now from other providers that Google are no longer hesitant to ship it.
Multicomp 49 minutes ago | parent
and/or local voice cloning is good enough as is so Google doesn't grant a uniquely liable ability?
perrohunter 48 minutes ago | parent
112233 47 minutes ago | parent
burkaman 16 minutes ago | parent
Multicomp 46 minutes ago | parent
Getting GPT-Live to have unique enough voices and to be expressive with how I imagine the voices going in my head is hard to direct, there's not enough control there.
So this Gemini 3.8 specific large voice library and ability to tightly control (if you are willing to write a script) is nice to find, and while I'm not sure which of the 5,286 Gemini products this is, nor how to onboard and get started feeding this my own text files, nor what training will happen to my data if I did somehow use it, I love that the state of the industry is such that Google can do this and release it publicly, because that means eventually an equivalent product can come from someone else and be used locally / confidently that the generated audio or inputs won't be retained and misused.
ghostbrainalpha 32 minutes ago | parent
exhilaration 23 minutes ago | parent
Also the Qwen3-TTS demo is cool, you can describe the voice you want: https://huggingface.co/spaces/Qwen/Qwen3-TTS
I came across both on this subreddit, it's very active: https://www.reddit.com/r/TextToSpeech/
I'm personally using this locally: https://github.com/mateogon/pdf-narrator (it's a Python frontend for Kokoro) on my M1 Macbook Air (from 2020, with 8GB RAM) and it's incredible. I make my own audiobooks now - for free!
My favorite voice is am_michael and here's a sample: https://voicerankings.com/voice/kokoro-82M/male/am_michael/s...
talon8635 45 minutes ago | parent
sgc 42 minutes ago | parent
thevinter 39 minutes ago | parent
Price per hour:
- 3.8 Flash TTS, standard: $0.81
- 3.8 Flash TTS, batch: $0.41
- 3.8 Flash‑Lite TTS, standard: $0.54
- 3.8 Flash‑Lite TTS, batch: $0.27
xnx 38 minutes ago | parent
mamudo 32 minutes ago | parent
laweijfmvo 31 minutes ago | parent
andrewstuart 36 minutes ago | parent
They all sound like Americans putting in their best fake British accent.
maelito 31 minutes ago | parent
Having a voice under 1Mo is crazy, even if it sounds robotic.
drewbitt 26 minutes ago | parent
burkaman 22 minutes ago | parent
avazhi 13 minutes ago | parent
Um, what?
nater5000 17 minutes ago | parent
Also weird that there are no "neutral gender" voices in the English language. There's also limited "use cases," like the "Gaming" use case is empty?
And there's no pricing listed anywhere.
I don't know, I guess their roll out is a bit sloppy. It's a bit of a shame, though, since the voices which are available all sound like generic Gemini voices to me. Nothing stands out is being particularly interesting or impressive about this.
accountrequired 17 minutes ago | parent
How long is this stored? What could go wrong? :P
m3kw9 14 minutes ago | parent
seemaze 13 minutes ago | parent
Is there a good browser extension that does this with a flexible TTS backend? I know Qwen, Kokoro, and VibeVoice all have decent quality..
kyrra 11 minutes ago | parent
xmorse 12 minutes ago | parent
https://storage.googleapis.com/gweb-uniblog-publish-prod/ori...
fullstackwife 12 minutes ago | parent
nitroedge 7 minutes ago | parent
$0.50 per hour pricing could last a long time with back and forth conversation use.