1.6 KiB
MeetVault Local Speech Engine
This process is local, not the future MeetVault backend. It provides the demo's real-time Whisper + diart functionality.
Recommended environment
Use Python 3.10 or 3.11. diart/PyTorch/pyannote compatibility is much better there than on very new Python releases.
python -m venv .venv
# Windows
.venv\Scripts\activate
# Linux/macOS
source .venv/bin/activate
pip install -r requirements.txt
For diart, install the native audio dependencies required by your OS and authenticate with Hugging Face if the pyannote models are gated in your environment. Follow the current diart/pyannote model instructions and accept their model terms before starting real mode.
Run real mode
python server.py
The desktop application connects to:
ws://127.0.0.1:8765/ws/live
Run mock mode
Mock mode requires only the FastAPI stack and is intended to validate UI behavior:
python server.py --mock
You can also choose Mock demo under MeetVault Settings → Local speech engine.
Audio protocol
- mono
- signed int16 little-endian PCM
- 16 kHz
- binary WebSocket frames
The first WebSocket message is JSON configuration. The engine replies with JSON events such as:
{
"type": "final",
"segment": {
"id": "uuid",
"speakerId": "speaker_1",
"start": 12.1,
"end": 15.4,
"text": "We should deploy tomorrow.",
"final": true
}
}
Speaker names such as Alice are not generated by the engine. The engine produces stable session labels (speaker_1, speaker_2, ...), and the MeetVault UI lets the user rename those labels live.