Models

Lipflow Translates Silent Mouthing Into Mac Text

Lipflow, a new open-source macOS application, uses local AI to translate silent lip movements into text, offering a private and noise-free dictation alternative for shared workspaces.

AlphaSignal1 day agoModels
Image: AlphaSignal

Lipflow has launched as an open-source macOS proof of concept that allows users to dictate text silently into any active application. By capturing a user's lip movements through a standard webcam, the application runs visual speech recognition locally on Apple Silicon to insert text directly at the cursor. This system provides a silent dictation option for practitioners working in shared offices or public spaces where speaking aloud is not feasible.

The application is built on the Auto-AVSR model trained on the LRS3 dataset, utilizing MediaPipe FaceLandmarker for real-time facial tracking and MLX for optimized performance on Apple hardware. During an initial eight-minute setup, Lipflow fine-tunes its lip-reading capabilities to the user's specific face and adapts its language model to their phrasing. For video processing, MediaPipe tracks facial landmarks to align the user's face with a mean training face using the eyes, nose base, and mouth as anchors. It then extracts a 96 x 96 grayscale mouth crop and resamples the video feed to 25 frames per second.

To resolve look-alike phonemes, Lipflow offers multiple text cleanup options. Users can process transcriptions offline using local rules, run a local Qwen3-0.6B model through MLX or Ollama, or send data to Anthropic's Claude. While Auto-AVSR achieves a 19.1% word error rate on the LRS3 research corpus, silent mouthing typically produces smaller movements than spoken speech. To address this, Lipflow features a Whisper mode that fuses visual lip tracking with soft audio, which successfully drops the word error rate from 31.9% to 6.9% on test clips.

Users control the application using the Right Option key, holding it down to mouth words and releasing it to paste the text, or double-tapping for a hands-free session lasting up to 60 seconds. For developers and practitioners looking to build upon this technology, the Lipflow code is available under an MIT license. However, the downloaded LRS3-trained model weights are restricted to non-commercial research use only.

This is our own summary of reporting by AlphaSignal

More in Models