Files
2026-09-30 00:50:44 +07:00
..
2026-09-30 00:50:44 +07:00
2026-09-30 00:50:44 +07:00
2026-09-30 00:50:44 +07:00
2026-09-30 00:50:44 +07:00

MeetVault Local Speech Engine

This process is local, not the future MeetVault backend. It provides the demo's real-time Whisper + diart functionality.

Use Python 3.10 or 3.11. diart/PyTorch/pyannote compatibility is much better there than on very new Python releases.

python -m venv .venv
# Windows
.venv\Scripts\activate
# Linux/macOS
source .venv/bin/activate

pip install -r requirements.txt

For diart, install the native audio dependencies required by your OS and authenticate with Hugging Face if the pyannote models are gated in your environment. Follow the current diart/pyannote model instructions and accept their model terms before starting real mode.

Run real mode

python server.py

The desktop application connects to:

ws://127.0.0.1:8765/ws/live

Run mock mode

Mock mode requires only the FastAPI stack and is intended to validate UI behavior:

python server.py --mock

You can also choose Mock demo under MeetVault Settings → Local speech engine.

Audio protocol

  • mono
  • signed int16 little-endian PCM
  • 16 kHz
  • binary WebSocket frames

The first WebSocket message is JSON configuration. The engine replies with JSON events such as:

{
  "type": "final",
  "segment": {
    "id": "uuid",
    "speakerId": "speaker_1",
    "start": 12.1,
    "end": 15.4,
    "text": "We should deploy tomorrow.",
    "final": true
  }
}

Speaker names such as Alice are not generated by the engine. The engine produces stable session labels (speaker_1, speaker_2, ...), and the MeetVault UI lets the user rename those labels live.