422 lines
9.8 KiB
Markdown
422 lines
9.8 KiB
Markdown
# MeetVault Desktop Demo
|
|
|
|
This repository is a runnable desktop-demo scaffold based on the supplied MeetVault wireframes and desktop architecture notes.
|
|
|
|
## Demo scope
|
|
|
|
Included now:
|
|
|
|
- Tauri + React + TypeScript desktop UI
|
|
- Dashboard
|
|
- Meetings list
|
|
- Meeting details
|
|
- Local video/audio playback
|
|
- Live meeting screen
|
|
- Real-time Whisper transcription through a local Python speech engine
|
|
- Real-time speaker diarization through diart
|
|
- `Speaker 1`, `Speaker 2`, ... labels
|
|
- Live speaker rename
|
|
- Transparent always-on-top overlay window
|
|
- Microphone capture
|
|
- Optional system-audio/screen capture through the WebView capture API
|
|
- Local recording files
|
|
- Local transcript data
|
|
- Functional Settings screen
|
|
- Mock speech-engine mode for UI testing without ML models
|
|
|
|
Not included in this demo:
|
|
|
|
- User login
|
|
- User identification/account verification
|
|
- Subscription/payment
|
|
- PostgreSQL
|
|
- SeaweedFS
|
|
- Remote cloud storage
|
|
- Real summary/task/reminder backend
|
|
|
|
The backend-facing demo state is intentionally:
|
|
|
|
```text
|
|
Processing (The job). Please wait a moment
|
|
```
|
|
|
|
Summary/task/reminder controls show this processing state but do not call a production backend yet.
|
|
|
|
---
|
|
|
|
## Architecture
|
|
|
|
```text
|
|
Desktop UI (Tauri + React)
|
|
│
|
|
├── microphone / optional system audio + screen
|
|
│
|
|
├── MediaRecorder ───────────► local recording.webm
|
|
│
|
|
└── 16 kHz PCM
|
|
│
|
|
▼
|
|
Local Speech Engine
|
|
Python WebSocket server
|
|
│
|
|
┌──────┴───────┐
|
|
▼ ▼
|
|
faster-whisper diart
|
|
transcription diarization
|
|
│ │
|
|
└──────┬───────┘
|
|
▼
|
|
speaker-aware segments
|
|
│
|
|
┌──────┴──────────┐
|
|
▼ ▼
|
|
Live transcript Overlay window
|
|
```
|
|
|
|
The local speech engine is **not** the future MeetVault backend. It is a local ML side process used by the desktop demo.
|
|
|
|
---
|
|
|
|
## Project structure
|
|
|
|
```text
|
|
MeetVault-Desktop-Demo/
|
|
├── src/ React desktop UI
|
|
│ ├── components/
|
|
│ ├── lib/
|
|
│ │ ├── capture.ts microphone/system capture + MediaRecorder
|
|
│ │ ├── speechClient.ts local WebSocket speech client
|
|
│ │ ├── liveBus.ts main-window ↔ overlay events
|
|
│ │ └── localFiles.ts local recording/transcript persistence
|
|
│ ├── pages/
|
|
│ │ ├── Dashboard.tsx
|
|
│ │ ├── LiveMeeting.tsx
|
|
│ │ ├── Meetings.tsx
|
|
│ │ ├── MeetingDetails.tsx
|
|
│ │ ├── Settings.tsx
|
|
│ │ └── Overlay.tsx
|
|
│ └── store/
|
|
├── src-tauri/ Tauri 2 desktop shell
|
|
├── speech-engine/ local Whisper + diart process
|
|
├── docs/ supplied source wireframes/README
|
|
└── scripts/
|
|
```
|
|
|
|
---
|
|
|
|
## Prerequisites
|
|
|
|
### Desktop UI
|
|
|
|
- Node.js 20+
|
|
- Rust toolchain compatible with your installed Tauri 2 release
|
|
- Platform-specific Tauri prerequisites
|
|
|
|
### Real speech engine
|
|
|
|
Recommended:
|
|
|
|
- Python 3.10 or 3.11
|
|
- ffmpeg
|
|
- PortAudio
|
|
- libsndfile
|
|
- Python packages in `speech-engine/requirements.txt`
|
|
- Hugging Face/pyannote model access required by your diart setup
|
|
|
|
For UI-only testing, use mock speech mode instead.
|
|
|
|
---
|
|
|
|
## 1. Install frontend dependencies
|
|
|
|
```bash
|
|
npm install
|
|
```
|
|
|
|
---
|
|
|
|
## 2. Test the UI in a browser
|
|
|
|
```bash
|
|
npm run dev
|
|
```
|
|
|
|
Open:
|
|
|
|
```text
|
|
http://127.0.0.1:1420
|
|
```
|
|
|
|
Browser mode is useful for UI development. Tauri-only functions such as the native overlay window and persistent AppLocalData storage use reduced fallbacks in browser mode.
|
|
|
|
---
|
|
|
|
## 3. Install the local speech engine
|
|
|
|
Create a Python 3.10/3.11 virtual environment inside `speech-engine`:
|
|
|
|
```bash
|
|
cd speech-engine
|
|
python -m venv .venv
|
|
```
|
|
|
|
Windows:
|
|
|
|
```powershell
|
|
.venv\Scripts\activate
|
|
pip install -r requirements.txt
|
|
```
|
|
|
|
Linux/macOS:
|
|
|
|
```bash
|
|
source .venv/bin/activate
|
|
pip install -r requirements.txt
|
|
```
|
|
|
|
### Real mode
|
|
|
|
```bash
|
|
python server.py
|
|
```
|
|
|
|
### Mock mode
|
|
|
|
```bash
|
|
python server.py --mock
|
|
```
|
|
|
|
The application expects the local engine at:
|
|
|
|
```text
|
|
ws://127.0.0.1:8765/ws/live
|
|
```
|
|
|
|
This can be changed from Settings.
|
|
|
|
---
|
|
|
|
## 4. Run as a desktop application
|
|
|
|
From the repository root:
|
|
|
|
```bash
|
|
npm run desktop
|
|
```
|
|
|
|
On Windows, convenience scripts are also provided:
|
|
|
|
```powershell
|
|
scripts\start-demo.ps1
|
|
```
|
|
|
|
or for UI testing with mock speech:
|
|
|
|
```powershell
|
|
scripts\start-demo-mock.ps1
|
|
```
|
|
|
|
---
|
|
|
|
# Live meeting behavior
|
|
|
|
When a meeting starts:
|
|
|
|
1. MeetVault requests microphone access.
|
|
2. If **Capture system audio / screen** is enabled, it also opens the system screen/audio picker.
|
|
3. Audio is mixed locally.
|
|
4. The mixed audio is recorded through `MediaRecorder`.
|
|
5. The same audio is converted to mono 16 kHz signed PCM.
|
|
6. PCM is streamed to `speech-engine/server.py`.
|
|
7. Whisper produces timestamped text.
|
|
8. diart produces speaker time ranges.
|
|
9. The local engine merges the two timelines.
|
|
10. The UI receives segments such as:
|
|
|
|
```json
|
|
{
|
|
"speakerId": "speaker_1",
|
|
"start": 18.2,
|
|
"end": 21.7,
|
|
"text": "We should finish the API before Friday.",
|
|
"final": true
|
|
}
|
|
```
|
|
|
|
11. The user can rename `speaker_1` to `Alice` without changing the underlying diarization ID.
|
|
12. The main transcript and overlay immediately display the renamed speaker.
|
|
|
|
---
|
|
|
|
# Live overlay
|
|
|
|
The Tauri build creates a second transparent window.
|
|
|
|
Properties:
|
|
|
|
- Always on top
|
|
- Transparent
|
|
- Resizable
|
|
- No decorations
|
|
- Hidden from the taskbar
|
|
- Shows the latest transcript lines
|
|
- Speaker name can be clicked to rename
|
|
- Receives the same live transcript event stream as the main window
|
|
|
|
The overlay is opened from **Show overlay** on the live meeting screen.
|
|
|
|
---
|
|
|
|
# Local files
|
|
|
|
There is no cloud storage in this demo.
|
|
|
|
In Tauri mode, media is stored below the application's local-data directory:
|
|
|
|
```text
|
|
meetings/{meeting_uuid}/recording.webm
|
|
meetings/{meeting_uuid}/meeting.json
|
|
meetings/{meeting_uuid}/transcript.json
|
|
```
|
|
|
|
The Meeting Details page can play the stored video/audio file.
|
|
|
|
If screen capture is enabled, the recording normally contains video + mixed audio.
|
|
|
|
If only microphone capture is enabled, the recording is audio-only.
|
|
|
|
---
|
|
|
|
# Video playback
|
|
|
|
The Meeting Details screen includes a real HTML5 video/audio player.
|
|
|
|
The Tauri asset protocol is scoped only to MeetVault's local `meetings` directory.
|
|
|
|
This allows locally recorded media to be played without uploading it anywhere.
|
|
|
|
---
|
|
|
|
# Settings
|
|
|
|
The Settings screen is functional and persisted locally.
|
|
|
|
Current settings include:
|
|
|
|
### General
|
|
|
|
- Launch at startup preference
|
|
- Keep in tray preference
|
|
- Global hotkey value
|
|
- Theme preference
|
|
|
|
The startup/tray/hotkey preferences are stored now; OS registration is intentionally left for the packaging phase.
|
|
|
|
### Audio
|
|
|
|
- Microphone device ID
|
|
- System audio/screen capture toggle
|
|
- Microphone level monitoring
|
|
|
|
### Transcription
|
|
|
|
- Whisper model: `base`, `small`, `medium`
|
|
- Compute device: auto/CPU/CUDA
|
|
- Language
|
|
- Speaker diarization toggle
|
|
|
|
### Overlay
|
|
|
|
- Enabled
|
|
- Opacity
|
|
- Font size
|
|
- Visible line count
|
|
|
|
### Speech engine
|
|
|
|
- Real Whisper + diart mode
|
|
- Mock mode
|
|
- Local WebSocket URL
|
|
|
|
---
|
|
|
|
# Backend placeholder
|
|
|
|
This demo deliberately has no remote backend implementation.
|
|
|
|
Any backend-only action returns/displays:
|
|
|
|
```text
|
|
Processing (The job). Please wait a moment
|
|
```
|
|
|
|
This currently applies to example actions such as:
|
|
|
|
- Generate/regenerate summary
|
|
- Create reminder
|
|
- Action-item processing
|
|
|
|
The UI is structured so a real API can replace `src/lib/backend.ts` later.
|
|
|
|
---
|
|
|
|
# Current limitations
|
|
|
|
## System audio
|
|
|
|
System-audio capture depends on the operating system, WebView implementation, selected source, and whether that source permits audio capture.
|
|
|
|
Microphone capture is the reliable baseline for the demo.
|
|
|
|
## diart model setup
|
|
|
|
Real diart execution can require pyannote model terms/authentication. The exact dependency combinations can also be sensitive to Python, PyTorch, torchaudio, and platform versions.
|
|
|
|
Use the recommended Python version and validate the speech-engine environment on the target development machine before packaging.
|
|
|
|
## Realtime accuracy
|
|
|
|
Speaker diarization can temporarily swap speaker labels, especially with:
|
|
|
|
- overlapping speakers
|
|
- short utterances
|
|
- background noise
|
|
- similar voices
|
|
|
|
The rename function changes the display name but does not perform biometric identity recognition.
|
|
|
|
## Long recordings
|
|
|
|
The media recorder writes the final Blob when the meeting ends in this initial demo. Before production use, change this to incremental/chunked file writing so multi-hour meetings do not accumulate the entire media Blob in memory.
|
|
|
|
---
|
|
|
|
# Recommended next development steps
|
|
|
|
1. Validate real Whisper + diart on the target Windows machine.
|
|
2. Replace final-Blob recording with incremental local file writes.
|
|
3. Add proper microphone device enumeration.
|
|
4. Add native Windows loopback audio capture if system-audio reliability is required.
|
|
5. Add session crash recovery.
|
|
6. Add transcript segment editing/reassignment.
|
|
7. Package the Python speech engine as a sidecar or replace it with native whisper.cpp + a native/ONNX diarization runtime.
|
|
8. Implement the real backend jobs/API later.
|
|
9. Replace local-only media with SeaweedFS when backend/storage work begins.
|
|
|
|
---
|
|
|
|
# Demo design decisions from the supplied specification
|
|
|
|
```text
|
|
Login / subscription REMOVED
|
|
Remote backend PLACEHOLDER ONLY
|
|
Backend message "Processing (The job). Please wait a moment"
|
|
Media storage LOCAL
|
|
Video playback INCLUDED
|
|
Live transcription INCLUDED
|
|
Speaker diarization INCLUDED
|
|
Live speaker rename INCLUDED
|
|
Realtime overlay INCLUDED
|
|
Settings INCLUDED
|
|
```
|