← All projects
Open Source 2026

LocalEar

Realtime microphone-to-text transcription with speaker diarization, running entirely in the browser and on CPU.

Python FastAPI ONNX Runtime WebSocket
Repository
Context

Live meeting transcription usually means sending audio to a cloud API or requiring a GPU for local ASR. The goal was a self-hosted option that runs on CPU alone, keeps audio processing in the browser as much as possible, and still separates who said what.

Result

Opening the browser and speaking produces a running, speaker-labelled transcript with no GPU and no terminal step — segment boundaries are decided client-side, transcription and diarization happen on the server, and everything from ASR model to login is optional and swappable via env vars.

Architecture

The browser buffers mic audio in an AudioWorklet, runs Silero VAD and a NeXt-TDNN speaker-change detector client-side (both vendored ONNX models via onnxruntime-web) to decide segment boundaries, then streams raw blocks over a WebSocket. A FastAPI server transcribes each segment through a resident GGUF ASR model (any model, via transcribe-cpp/ggml, CPU or Vulkan/CUDA), applies fuzzy custom-vocabulary correction, and diarizes speakers server-side with the same embedding model via cosine-centroid clustering. Session history (PostgreSQL), audio storage (any S3-compatible store), summaries (any OpenAI-compatible endpoint), and login (OIDC) are all optional — the app runs stateless with none of them configured.

LocalEar architecture diagram

Want something similar?

Talk about your project