Opensource modern TTS engines for Android are now available! - SLRPNK
slrpnk.net
external-link
I just found this. Main page [https://k2-fsa.github.io/sherpa/onnx/index.html] This is huge! As a german, I use thorsten medium [https://huggingface.co/csukuangfj/sherpa-onnx-apk/resolve/main/tts-engine-new/1.10.26/sherpa-onnx-1.10.26-arm64-v8a-de-tts-engine-vits-piper-de_DE-thorsten-medium.apk] as he simply made the best dataset. Mixing english with german, speaking numbers, single letters, pausing without a “.” but just a linebreak, all those can be essential. And… it is nearly perfect! And all local! This is crazy! eSpeak can finally go to rest!

Is there a way to install multiple languages? I am would be using German and English, sometimes Italian.

Lemongrab
link
fedilink
84M

Piper TTS on Android is great, the project is called Sherpa Onnx. It doesn’t require permissions either.

RHvoice

boredsquirrel
creator
link
fedilink
24M

Only english and a few others

That isn’t a problem for English speaking countries. Maybe try espeak?

I’m not sure I completely understand the need to run an AI voice on your phone. It uses a lot of resources for little gain.

boredsquirrel
creator
link
fedilink
24M

I mentioned espeak at least somewhere. It is 32bit and likely unmaintained. And the Android version is already better than what I had on Fedora KDE.

Yes, something like improved espeak would be fine and extremely efficient. But this works fine too.

Modern phones have NPUs for low power neural network tasks like those. Older ones hopefully use the GPU for that.

The apps are pretty flawed in that they have only one model.

And it is very impressive how espeak can do so much, so efficiently.

Its simply that nobody created a better voice model (and removed all the silly joke okes) for it, and ported the APK to modern Android with armv8a (64bit)

I’ve been using this for a few months now with Organic Maps and it’s incredibly good. A few versions ago I noticed some issues with certain letter sounds but with frequent updates it now sounds so natural.

I found “vits-piper-en_GB-southern_english_female-medium” to be the most natural sounding voice for English.

boredsquirrel
creator
link
fedilink
24M

Nice! It all depends on the model.

Thorsten has spoken 23+h of audios into his main dataset, crazy

Um ok but is this really a big deal? Good sounding TTS is nice but crappy sounding TTS is good enough for most purposes and is fairly easy. Speech to text is way harder.

boredsquirrel
creator
link
fedilink
34M

Speech to text exists with FUTO keyboard and whisper.

Espeak is the only TTS for german, and it is 32bit and sounds awful. It is quite embarassing to use that for navigating with other people used to… modern voices.

CubitOom
link
fedilink
124M

For anyone that is dyslexic and privacy conscious it’s a pretty big deal.

Its also nice for anyone that is only privacy conscious.

Create a post

A place to discuss privacy and freedom in the digital world.

Privacy has become a very important issue in modern society, with companies and governments constantly abusing their power, more and more people are waking up to the importance of digital privacy.

In this community everyone is welcome to post links and discuss topics related to privacy.

Some Rules

  • Posting a link to a website containing tracking isn’t great, if contents of the website are behind a paywall maybe copy them into the post
  • Don’t promote proprietary software
  • Try to keep things on topic
  • If you have a question, please try searching for previous discussions, maybe it has already been answered
  • Reposts are fine, but should have at least a couple of weeks in between so that the post can reach a new audience
  • Be nice :)

Related communities

much thanks to @gary_host_laptop for the logo design :)

  • 0 users online
  • 124 users / day
  • 1.05K users / week
  • 1.3K users / month
  • 4.58K users / 6 months
  • 1 subscriber
  • 3.35K Posts
  • 85.2K Comments
  • Modlog