Files
2026-09-30 00:50:44 +07:00

67 lines
1.6 KiB
Markdown

# MeetVault Local Speech Engine
This process is **local**, not the future MeetVault backend. It provides the demo's real-time Whisper + diart functionality.
## Recommended environment
Use Python **3.10 or 3.11**. `diart`/PyTorch/pyannote compatibility is much better there than on very new Python releases.
```bash
python -m venv .venv
# Windows
.venv\Scripts\activate
# Linux/macOS
source .venv/bin/activate
pip install -r requirements.txt
```
For diart, install the native audio dependencies required by your OS and authenticate with Hugging Face if the pyannote models are gated in your environment. Follow the current diart/pyannote model instructions and accept their model terms before starting real mode.
## Run real mode
```bash
python server.py
```
The desktop application connects to:
```text
ws://127.0.0.1:8765/ws/live
```
## Run mock mode
Mock mode requires only the FastAPI stack and is intended to validate UI behavior:
```bash
python server.py --mock
```
You can also choose **Mock demo** under MeetVault Settings → Local speech engine.
## Audio protocol
- mono
- signed int16 little-endian PCM
- 16 kHz
- binary WebSocket frames
The first WebSocket message is JSON configuration. The engine replies with JSON events such as:
```json
{
"type": "final",
"segment": {
"id": "uuid",
"speakerId": "speaker_1",
"start": 12.1,
"end": 15.4,
"text": "We should deploy tomorrow.",
"final": true
}
}
```
Speaker names such as `Alice` are **not** generated by the engine. The engine produces stable session labels (`speaker_1`, `speaker_2`, ...), and the MeetVault UI lets the user rename those labels live.