Read Their Lips
Upload a silent clip and get a transcript read from the speaker's mouth, not the audio track.
4 min read · updated Aug 2026
💡 In plain words
You give it a video with no usable sound, and it works out what the person said by watching their mouth move. No audio is involved at any point — it is reading lips, the way a person would, only faster.
🎯 A real example
You have footage of a press conference where the microphone cut out for thirty seconds. You upload the clip, mark that stretch, point the tool at the speaker's face, and get back a text guess at what was said while the sound was gone.
🤔 Is it for you?
- Recovering dialogue from clips where the audio failed or was never recorded
- Video editors working with muted or corrupted source footage
- Curiosity — it is genuinely interesting to watch a model attempt this
- Anything you intend to quote, publish, or rely on as fact
- Video where the speaker's mouth is small, angled, blurred, or partly hidden
- Evidence, legal work, or accusations — the failure mode is a confident wrong sentence
Lip reading has been a human skill for as long as there have been deaf people and noisy rooms, and it has always been an unreliable one — the best human lip readers work at roughly 30–45% word accuracy on unfamiliar speakers. Read Their Lips, from Symphonic Labs, is an attempt to do the same job with a neural network: upload a silent clip, and it returns what it believes was said.
It is a genuinely impressive demonstration and a genuinely dangerous tool to trust. Both of those need saying on the same page, because the output does not look uncertain — it looks like a transcript.
What it actually does
You upload a video, optionally narrow it to a time range, and select which face to track when more than one person is on screen. The tool then analyses the movement of that person’s mouth and returns text.
No audio is used at any stage. That is the whole point — it works on footage where the sound failed, was never captured, or was stripped. If your clip has usable audio, ordinary speech-to-text will be far more accurate and you should use that instead.
How it works, and why that matters
Under the hood a convolutional neural network maps mouth shapes to phonemes — the individual sound units of speech — and then assembles those phonemes into plausible words. It was trained on a large body of video where the phonemes were labelled.
The important consequence sits in that description: it is matching pictures of mouths, and many distinct sounds produce identical pictures.
| These sounds | Look like this on the mouth |
|---|---|
| p · b · m | lips pressed together, then released |
| f · v | top teeth on bottom lip |
| t · d · n | tongue behind the teeth, invisible from outside |
“Pat”, “bat” and “mat” are the same image. So are “fan” and “van”. A human lip reader closes that gap with context, expectation, and knowing the speaker. A model closes it with a guess — and the guess arrives phrased as a fact.
That is the mechanism behind the tool’s most-quoted failure: testing by Newsweek found it rendering “island view” as “I love you”. Those two phrases genuinely do look near-identical on a mouth. The model was not malfunctioning; it hit the ceiling the task has.
What a good clip looks like
Accuracy swings enormously with input quality. It needs:
- A large, sharp face — mouth detail is the entire signal
- Good, even lighting on the lower face
- A roughly front-on angle — profile shots lose most of the information
- An unobstructed mouth — no hand, microphone, or heavy moustache
- Unhurried, clearly articulated speech
Miss two or three of those and the output degrades from “roughly right” to “fluent invention”.
Where it earns its place
Recovering your own footage. You shot an interview, the lav mic died for forty seconds, and you need to know roughly what was covered so you can re-record or caption around it. Here a rough guess is worth a lot, and you can check it against your memory of the shoot.
Editing muted source material. Working out which take is which, or where a sentence begins, without hunting for a separate audio file.
Accessibility experiments. Real-time lip reading for deaf users is the ambition this technology serves, and tools like this are how it gets tested in public.
Where it does real harm
The tool’s output is fluent English. That fluency is the danger — a wrong transcript does not look wrong. It looks like a quote.
Do not use it on footage of people to work out what they “really said.” This is the most common use it gets put to and the worst. A confidently generated sentence attributed to a real person is defamation whether a human or a model wrote it, and you would have no way to prove the words were ever spoken.
Do not treat it as evidence. It is not admissible, it is not reproducible, and its failure mode is producing something plausible rather than producing nothing.
Honest limitations
- A hard accuracy ceiling — set by the task, not the model. Identical mouth shapes cannot be separated visually.
- Steep quality dependence — angle, lighting, and resolution move accuracy more than anything else.
- No uncertainty signal in the output — you get sentences, not confidence scores, so a guess reads exactly like a correct reading.
- Language and accent limits — training data determines what it can handle, and coverage outside clear standard English is thin.
The verdict
Read Their Lips is worth trying, and worth understanding before you rely on a word of it. As a way to recover the gist of dialogue from footage you shot yourself and can sanity-check, it does something no other tool does. As a source of truth about what a person said, it is not one, and the confident phrasing of its output makes that easy to forget.
A rule that will keep you out of trouble: if you could not independently confirm the sentence some other way, do not repeat it as something the person said.
Official resource: Read Their Lips
Get the best new tools — before everyone else
One short, friendly email whenever we add a tool worth your time. No spam, unsubscribe anytime.