67 lines
1.6 KiB
Markdown
67 lines
1.6 KiB
Markdown
# MeetVault Local Speech Engine
|
|
|
|
This process is **local**, not the future MeetVault backend. It provides the demo's real-time Whisper + diart functionality.
|
|
|
|
## Recommended environment
|
|
|
|
Use Python **3.10 or 3.11**. `diart`/PyTorch/pyannote compatibility is much better there than on very new Python releases.
|
|
|
|
```bash
|
|
python -m venv .venv
|
|
# Windows
|
|
.venv\Scripts\activate
|
|
# Linux/macOS
|
|
source .venv/bin/activate
|
|
|
|
pip install -r requirements.txt
|
|
```
|
|
|
|
For diart, install the native audio dependencies required by your OS and authenticate with Hugging Face if the pyannote models are gated in your environment. Follow the current diart/pyannote model instructions and accept their model terms before starting real mode.
|
|
|
|
## Run real mode
|
|
|
|
```bash
|
|
python server.py
|
|
```
|
|
|
|
The desktop application connects to:
|
|
|
|
```text
|
|
ws://127.0.0.1:8765/ws/live
|
|
```
|
|
|
|
## Run mock mode
|
|
|
|
Mock mode requires only the FastAPI stack and is intended to validate UI behavior:
|
|
|
|
```bash
|
|
python server.py --mock
|
|
```
|
|
|
|
You can also choose **Mock demo** under MeetVault Settings → Local speech engine.
|
|
|
|
## Audio protocol
|
|
|
|
- mono
|
|
- signed int16 little-endian PCM
|
|
- 16 kHz
|
|
- binary WebSocket frames
|
|
|
|
The first WebSocket message is JSON configuration. The engine replies with JSON events such as:
|
|
|
|
```json
|
|
{
|
|
"type": "final",
|
|
"segment": {
|
|
"id": "uuid",
|
|
"speakerId": "speaker_1",
|
|
"start": 12.1,
|
|
"end": 15.4,
|
|
"text": "We should deploy tomorrow.",
|
|
"final": true
|
|
}
|
|
}
|
|
```
|
|
|
|
Speaker names such as `Alice` are **not** generated by the engine. The engine produces stable session labels (`speaker_1`, `speaker_2`, ...), and the MeetVault UI lets the user rename those labels live.
|