There is a small, slightly humiliating thing a lot of people do with their phones, and almost nobody talks about it. They put on an accent.
Not a funny one. A flattened, careful, vaguely American one, adopted for the thirty seconds it takes to dictate a message, and dropped the moment the keyboard comes back. Scottish, Nigerian, Indian, broad Australian, West Country, Singaporean — if your vowels are not the ones the model was mostly trained on, you have probably learned to move them. You may not have noticed you were doing it.
The first thing people try is the settings screen. It does not work, and the reason is worth saying plainly: locale and accent are not the same thing. Setting your phone to English (Australia) tells it which spellings to prefer, which place names to expect, and roughly which dictionary to reach for. It does not tell it that you say things the way you say them. Two people in the same city, on the same locale setting, can get wildly different accuracy from the same model. The setting describes where you are. The problem is how you sound.
The number everybody repeats
So the obvious question is why the phone cannot just learn your voice. You have spoken into it thousands of times. It knows what you meant, because you corrected it.
Ask around, and you get a number: about half an hour of clean recorded audio, carefully read, to adapt a speech model to one speaker. It is repeated confidently, it sounds like a fact about speech recognition, and it is the reason the idea usually stops there. Half an hour of studio-quiet reading is not something a normal person is going to do for the privilege of texting in their own voice.
Here is the thing. That number is real, and it is measuring the wrong technique.
Roughly half an hour is what you need for full fine-tuning — taking the entire model and nudging every one of its parameters toward one speaker. Full fine-tuning is the sledgehammer. It is also, specifically, the technique that behaves worst when you have very little data, because a model with that many free parameters and a few minutes of audio will happily memorise the few minutes rather than learn the voice. The half-hour figure is not a law of speech recognition. It is the price of one method, quoted as though it were the price of the goal.
The cheaper rungs
Between "do nothing" and "fine-tune the whole model" there are rungs, and they have been getting cheaper for a while. Three worth knowing about:
- Zero-shot adaptation from a prompt. You give the model a short sample of the speaker's audio together with its correct transcript, at the moment of recognition, and it conditions on that — the way you might show someone one example before asking them to do a task. No training run. Nothing is saved. Seconds of audio, not half an hour.
- Speaker-embedding conditioning. The model is built to accept a compact numerical description of a voice alongside the audio. Produce that description once from a small sample, hand it over with every request, and the model leans toward that speaker. The description is small enough to sit on the device.
- LoRA. Instead of moving every parameter, you train a small set of extra ones bolted onto the side, and freeze the rest. Far fewer things to learn means far less data needed to learn them, and the adapter is small enough to keep per person.
None of these is exotic. All three are ordinary technique in adjacent parts of machine learning. What has not happened is anyone packaging them into the thing on your phone.
Four things this does not say
This is the part I would want if I were reading it, so it is going in the middle rather than buried at the end.
The published gains are small. Not "accent problem solved". Measurable improvements, of the kind that turn a sentence you would have retyped into one you can send. Worth having, not transformative.
They were measured on unrelated populations. The studies behind these techniques were not run on the accents I am talking about. So the honest statement is that the technique is cheap and shows gains somewhere, not that it has been shown to work for Glaswegian or Tamil-inflected English specifically. Nobody has run that test.
Nobody has packaged any of it for Android. I went looking for an implementation you could actually install, and did not find one for any of the three rungs. This is an argument that something is possible and cheap, not a pointer to a download.
And the hardest one: the models on your phone are the wrong models. Dictation has to produce words as you speak, which means a streaming model — one that commits to text before the sentence is finished. Streaming models are less accurate than the non-streaming ones, and the adaptation research largely improves the non-streaming ones. So the cheap technique and the model that needs it most are, at the moment, not the same thing.
Why it is still worth writing down
I have deliberately not claimed that nobody has thought of this. I do not know that, nobody ran the search that would establish it, and "I could not find it" is a much weaker statement than "it does not exist" — the kind of gap that looks like an opportunity right up until you find the six people already in it.
What I will claim is narrower and, I think, more useful. There is a widely repeated number that makes a fix look impossible. The number is measuring the most expensive method available. Cheaper methods exist, they are well understood, and the reason they are not on your phone is not that someone tried them and they failed.
That is the whole point. A wrong number, repeated often enough, stops people attempting things — and a discouraging fact that is not actually true about the thing you want to do is worth correcting, even when you cannot hand anybody the fix. Somebody reading this knows more about on-device speech than I do. If the answer is "here is why that will not work either", I would genuinely like to know, because that answer is not written down anywhere I could find.
In the meantime: if you catch yourself putting on a voice for your phone, you are not being precious, and it is not your accent that is wrong.