Skip to content
Rafi Akmal Widikta

backend

Elysia Voice Assistant

Offline-first speech-to-speech assistant for Linux: local wake word, VAD and transcription, with the LLM only choosing the action.

Role
Deep Learning
Date
August 2026

Technologies

  • python
  • asyncio
  • faster-whisper
  • silero vad
  • openwakeword
  • gemini
  • edge-tts
  • piper
  • pytorch
  • pydantic
  • structlog
  • pytest

Problem

Cloud voice assistants send everything you say to a server, and general-purpose ones will run whatever they are asked to. For an assistant that keeps a microphone open on your own machine all day, both defaults are wrong: the audio should not leave the laptop, and the assistant should only be able to do the things you decided it may do. Neither is negotiable when it sits next to your shell.

Solution

It is a local pipeline with a hard latency budget (two to three seconds end to end) and that budget is what shapes the architecture. An offline wake word model arms the microphone; Silero VAD decides when you have actually stopped speaking instead of trusting a fixed timer; faster-whisper transcribes Indonesian on the CPU with int8 quantisation. Only then does anything leave the machine, and only the transcribed text: the model has exactly one job, choosing which tool to call.

Every stage carries a timeout and a fallback, which is how the budget survives a slow component. Speech tries an online voice first, falls back to Piper offline with raw PCM, and if both fail it prints the answer to the console rather than going silent. Transcription is guarded by its own lock and raises a busy error instead of running twice, because the underlying runtime is not thread-safe; a silently corrupted transcript is worse than a rejected request.

The audio callback does almost nothing: it reads the current state and branches. The state machine moves through idle, listening, processing and speaking, and the heavy work runs on a separate thread so capture never stalls. A second rule is that the assistant must not hear itself, so the microphone is muted, the wake word paused and a cooldown applied while it speaks.

Nothing reaches the shell as free-form text. Commands go through a strict allowlist with no shell interpolation, destructive actions such as shutdown or reboot need a spoken two-step confirmation, and the system prompt carries an injection guard. When an app is not on the list the assistant works through exact match, session cache and fuzzy matching before asking the model to guess, then confirms out loud: did you mean X? Transient model failures are deliberately distinguished from an unrecognised request and are not cached, because caching a failure would teach it the wrong thing. 184 test cases cover the pipeline.

Contact

Interested in working together?

Tell me what you are building and what you need. A short note is enough.

Back to top