For Windows users, the most useful detail is less glamorous than an AI benchmark: Hayamimi can capture the PC’s audio output directly. That makes videos, streamed presentations and call audio potential transcription sources—not just whatever reaches a microphone.
What the “under 2GB” claim actually means
Hayamimi’s repository describes models loaded on demand and evicted using a least-recently-used cache. Its reported sub-2GB memory result uses three resident non-tier-0 models; that is not a guarantee that Windows, a browser and the entire workflow will fit on a PC with only 2GB of installed RAM.
GIGAZINE reports that the recognizer identifies each utterance’s language and routes it to a specialist model. The system uses quantized INT8 ONNX models through sherpa-onnx rather than requiring PyTorch or CUDA.
Its language routes cover Japanese, Chinese, Korean, Cantonese, English and 24 European languages, with approximately 1,600 other languages covered through a Meta Omnilingual ASR fallback. That broad catalog should not be confused with equally strong results in every language.
The practical takeaway: CPU-only describes the processing architecture, not universal performance on every CPU.
Live captions and cleaned-up transcripts are different outputs
According to GIGAZINE, draft subtitles refresh roughly every half-second during speech. Japanese confirmed lines are reported to arrive approximately 100 milliseconds after speech ends, while a second pass revisits recent utterances after two seconds of silence.
The repository’s headline Japanese result—3.8% character error rate—comes from a 15-clip test. Its performance table explicitly excludes two-pass refinement, so that figure should not be combined with the separate 15.5%-to-12.0% refinement result as though they were one experiment.
These are developer-reported measurements, not independent WindowsForum benchmarks. For a live presentation, the useful test is whether captions remain readable and timely on the actual machine while its other applications are running.
A practical Windows starting point
GIGAZINE lists Python 3.10 or later and ffmpeg available on PATH as prerequisites. Git is also needed if using the clone command below.
The maintainers identify Windows 11 as their developed-and-tested platform. They expect macOS and Linux to work, but say those platforms have not been CI-tested end to end; speaker-loopback capture is Windows-only.
Run these commands from a terminal:
git clone [GitHub - oboroge0/hayamimi: 早耳 - Real-time multilingual speech-to-text on CPU only. Live subtitles, browser dashboard, speaker labels, translation. No GPU, no cloud. · GitHub](https://github.com/oboroge0/hayamimi.git)
cd hayamimi
python -m venv.venv
.venv\Scripts\pip install -r requirements.txt
.venv\Scripts\python scripts/download_models.py
For the smaller Japanese-and-English configuration, substitute this model-download command:
.venv\Scripts\python scripts/download_models.py --minimal
GIGAZINE describes that configuration as approximately 1.1GB. This is a model-download footprint, not a statement about total runtime memory.
Start with microphone transcription and the browser interface:
.venv\Scripts\python scripts/realtime_transcribe.py --serve
The local views are:
[url]http://localhost:8833/dashboard[/url]— in-progress captions, confirmed text, translations and refined output.[url]http://localhost:8833/[/url]— the caption overlay for an OBS Browser Source.[url]http://localhost:8833/transcript[/url]— transcript history.
Successful operation means speech produces draft text, followed by confirmed lines and later refined output—not merely that the server starts.
For PC playback or combined meeting audio, use:
.venv\Scripts\python scripts/realtime_transcribe.py --input speaker --serve
.venv\Scripts\python scripts/realtime_transcribe.py --input mix --serve
The repository confirms that these modes use WASAPI loopback without requiring Stereo Mix. Mixed capture has no acoustic echo cancellation, so the maintainers recommend headphones to avoid capturing call audio twice.
Labels, translation and developer controls
GIGAZINE describes --speakers as enabling live speaker tagging through CAM++, followed by pyannote-based re-segmentation during refinement. These labels indicate speaker turns; they do not identify people by name.
Its --translate option translates Japanese captions into selected target languages, including English, Chinese, Korean and Spanish. The report cautions that translation can mishandle numbers and amounts. Treat translated captions as an aid to understanding, not an authoritative financial transcript.
Other reported controls include:
--replacefor post-recognition substitutions.- Number normalization for Japanese, Chinese and Cantonese.
- Runtime replacement and normalization dictionaries.
- Structured pipeline events through an application-level
EventHub. - WebSocket audio input through
--input ws. - An optional English Parakeet TDT v2 tier.
- An optional four-class Japanese punctuation model.
One important correction: the repository says --hotwords currently has no effect on its Japanese recognizer. It recommends --replace for Japanese proper nouns instead. It also warns that WebSocket ingestion has no authentication when users deliberately expose it beyond localhost.
The boundaries matter as much as the speed
GIGAZINE reports three significant recognition limitations: mixed-language sentences are unsupported, short speech following music or sound effects can trigger incorrect language detection, and overlapping speakers cannot be separated.
Those restrictions define sensible expectations. A clearly spoken presentation is a better starting test than a noisy meeting full of interruptions and Japanese-English code-switching.
Hayamimi’s attraction is straightforward: it offers a locally controlled captioning pipeline with useful Windows audio capture and browser integration. The sensible deployment approach is equally straightforward—test representative audio, verify names and numbers, and distinguish fast live captions from the transcript that emerges after refinement. A quick ear is useful; a careful human review still earns its place.
References
- 'Hayami,' a real-time multilingual speech recognition system that operates solely on the CPU, can perform live subtitle display, browser display, speaker labeling, and translated subtitles with less than 2GB of memory, without using a GPU or cloud API. - GIGAZINE GIGAZINE · 2026-10-11T03:00:00+00:00
- GitHub - oboroge0/hayamimi: 早耳 - Real-time multilingual speech-to-text on CPU only. Live subtitles, browser dashboard, speaker labels, translation. No GPU, no cloud. · GitHub github.com