Back to News
Advertisement
Advertisement

⚡ Community Insights

Discussion Sentiment

92% Positive

Analyzed from 369 words in the discussion.

Trending Topics

#tts#quality#speech#voice#model#inflect#micro#amazing#small#text

Discussion (27 Comments)Read Original on HackerNews

modinfo1 day ago
This is amazing, the quality blow my mind for such small model! I just replaced my old onnx model with yours!

here my implementation with speech dispatcher and server: https://github.com/skorotkiewicz/inflect-speechd

thanks for shearing!

yjftsjthsd-h1 day ago
Couple highlights:

> Complete local text-to-waveform speech synthesis under 10M parameters.

In case, like me, you hoped "complete" voice might mean both stt and tts. Not to speak poorly of it, just clarifying.

> English only, with one fixed male voice. This is not zero-shot voice cloning.

(And then a bunch of statements on limitations that I read as 'quality can be spotty but if you play with it it should be fine') But like. In <10M params I'm not judging:)

semiquaverabout 14 hours ago
When would “text-to-waveform speech synthesis” ever imply speech to text?
yjftsjthsd-habout 12 hours ago
The HN title is "Inflect-Micro-v2: complete voice in 9.36M parameters".
billdueberabout 16 hours ago
I keep seeing tts stories here. Is it just an interesting subset of the llm world, or is there a huge use case I’m somehow missing?
eightysixfourabout 15 hours ago
I use STT/TTS to interface with a local LLM for Home Assistant in my house.
NetOpWibbyabout 23 hours ago
The inflections are weird but this doesn't sound like a robot. Not bad!
K0baltabout 12 hours ago
How heavy in inference on this? The model would easily fit on many microcontroller modules, I wonder if they could run it?
tmaly1 day ago
This is impressive. I wish there were a voice clone option.
fastball1 day ago
With so few parameters, I imagine a voice fine-tune might be readily tractable.
sudbabout 14 hours ago
this is extremely encouraging for individuals/small companies being able to train pareto-frontier TTS models (specifically compute required to run vs quality of model output)
da-xabout 13 hours ago
I think we need more neurons in the human brain than parameters in this model for speech. I wonder what it says about the human brain vs LLM efficiency.
StilesCrisisabout 17 hours ago
I'd love to hear it but it seems your quota is exhausted.
g58892881about 16 hours ago
StilesCrisisabout 14 hours ago
Nice! Strangely, "Nano" sounds a lot better than "Micro" to me.

On my iPhone 14 Pro the page crashes after 2-3 plays. I wonder if it uses too much memory?

g58892881about 9 hours ago
also, i double checked, nano is nano and micro is micro. didnt fuck that up
g58892881about 12 hours ago
right. happens on my 13 too. memory leak confirmed, not sure yet what's causing it
jsomedon1 day ago
amazing quality for such small size!
itakeabout 24 hours ago
Amazing quality for small size, but definitely not that enjoyable to listen to.

IMHO, its at about the same quality level of historic TTS tools.

stavrosabout 23 hours ago
I'm not sure which historic tools you mean, but to me this sounds much better than anything older than ten years ago.
itakeabout 22 hours ago
I compared the macos Samantha just now and I guess the inflect-micro is marginally better...
leobgabout 22 hours ago
Ivona „Joey“, „Amy“
Advertisement
phoenixrangerabout 14 hours ago
amazing! was looking for something similar
mcbetzabout 21 hours ago
Alternative title: Text to speech in 9.36M, English only.