Voice -> Text transcription.

One can save alot of time and meticulous finger gymnastics by using a good voice transcription software. And if the AI that does the transcribing is run locally on the device there is no big privacy issue involved.

FUTO Keyboard on Google Play: https://play.google.com/store/apps/details?id=org.futo.inputmethod.latin.playstore&hl=en

FUTO seems excellent. I also tried Localwhisper for iOS. Quite good too.

More research here:-

1 Like

Very interesting :grin:

I I tried FUTO Keyboard, it has saved me a lot of time, and I was surprised at how much easier and faster typing became.

Why spend so much time typing slowly when there’s a tool that can make typing quicker and easier?

There are people here in Africa who can actually speak English but can't write, this would really be very helpful for them or those having problems with typing especially the old people with their shaky hands

FUTO Keyboard is interesting mainly because it treats the keyboard as a local input device, not as a cloud service.

The strongest privacy point is simple: it is designed to work without Internet access at all. Autocorrect, swipe typing and even voice input can run locally on the phone. Google Play currently says FUTO collects no data and shares no data with third parties. Google Play

That matters more for a keyboard than for most apps, because the keyboard potentially sees almost everything you type: searches, messages, names, addresses, fragments of passwords, private conversations, etc. Android does isolate password fields somewhat, but fundamentally an IME is a highly privileged piece of software.

Compared with something like Gboard, the distinction is architectural rather than simply “Google bad / FUTO good.” Gboard has substantial on-device processing too, and Google provides privacy controls. But Gboard is still network-capable and has mechanisms such as personalization, federated learning and optional improvement features. Google documents that some typing/voice-derived information can participate in those systems depending on settings. Google Support

FUTO instead makes the stronger design choice:

keyboard → local models → text

rather than potentially:

keyboard → local processing + cloud/account/service infrastructure → text

Another advantage is inspectability. FUTO Keyboard is based on Android's LatinIME and publishes its source code, so people can inspect how it works and build it themselves. One caveat: its current FUTO Source First License is source-available, not conventional OSI open-source, because it has restrictions on commercial use. GitHub

In practice

If the built-in keyboard is Gboard or Samsung Keyboard, FUTO gives you a substantially cleaner trust model:

  • no Internet permission by design;
  • no account required for normal operation;
  • offline speech recognition;
  • offline swipe/autocorrect;
  • no telemetry/data collection according to its Play declaration;
  • auditable source.

The trade-off is that Gboard is generally more mature: broader language support, better integrations, more refined prediction and some convenience features. FUTO itself still labels the keyboard as under active development. Google Play

So the real attraction is not merely “a private keyboard.” It is removing the network from the keyboard's threat model altogether. For software that receives nearly every sentence you type, that's a fairly compelling architectural choice.

1 Like

There are basically a few different ways to get thoughts into a message.

Normal typing is precise, but relatively slow. Swipe typing is already much faster once you get used to it, because you are composing continuously instead of pecking at individual keys.

Then there is voice. Voice is extremely fast for the sender, but a voice message transfers a lot of the cost to the receiver. A 30-second voice message takes roughly 30 seconds to listen to. Maybe only one sentence in it is actually new or relevant. With text, a fast reader can scan the whole thing in a few seconds, immediately see the important part, skip what they already know, search it later, quote one line, translate it, or feed it into another system.

So voice-to-text is interesting because it combines much of the bandwidth of speaking with the properties of text.

In Envoy/Botchi terms, this is partly a question of semantic transport. The goal isn't merely to transmit more data; it is to get the intended meaning across the boundary with as little unnecessary decoding work as possible.

A raw voice message says, in effect: here is a linear stream; consume it at approximately my speed and extract the useful information yourself.

Text gives the receiver much more agency. They can change the reading speed, jump around, inspect structure and decide what deserves attention.

That is also where the Babel problem appears: communication becomes worse when the medium adds unnecessary friction, ambiguity or reconstruction work. Botchi is closer to preserving the useful semantic object across the boundary.

So for many everyday messages:

speech → local voice-to-text → editable text → recipient

is probably a substantially better communication primitive than simply sending the recording.

You retain the speed and naturalness of speaking without forcing everyone else to listen at your pace.

1 Like

When I'm tired and I don't want to type, I just activate it​:blush:

[quote="cybe, post:6, topic:7316"]

Swipe typing is already much faster once you get used to it, because you are composing continuously instead of pecking at individual keys.

[/quote]

You do the talking and voice keyboard does the typing.

We have old friends who can't even see probably but can speak good English, this is good for them

Or for some, sometimes it’s the other way around.

When you are full of energy, social energiske wirkint mode you feel like talking and then later not.