Files
MeetVault/Wireframes/Desktop/MeetVault-Desktop-Demo/README.md
2026-09-30 00:50:44 +07:00

422 lines
9.8 KiB
Markdown

# MeetVault Desktop Demo
This repository is a runnable desktop-demo scaffold based on the supplied MeetVault wireframes and desktop architecture notes.
## Demo scope
Included now:
- Tauri + React + TypeScript desktop UI
- Dashboard
- Meetings list
- Meeting details
- Local video/audio playback
- Live meeting screen
- Real-time Whisper transcription through a local Python speech engine
- Real-time speaker diarization through diart
- `Speaker 1`, `Speaker 2`, ... labels
- Live speaker rename
- Transparent always-on-top overlay window
- Microphone capture
- Optional system-audio/screen capture through the WebView capture API
- Local recording files
- Local transcript data
- Functional Settings screen
- Mock speech-engine mode for UI testing without ML models
Not included in this demo:
- User login
- User identification/account verification
- Subscription/payment
- PostgreSQL
- SeaweedFS
- Remote cloud storage
- Real summary/task/reminder backend
The backend-facing demo state is intentionally:
```text
Processing (The job). Please wait a moment
```
Summary/task/reminder controls show this processing state but do not call a production backend yet.
---
## Architecture
```text
Desktop UI (Tauri + React)
│
├── microphone / optional system audio + screen
│
├── MediaRecorder ───────────► local recording.webm
│
└── 16 kHz PCM
│
▼
Local Speech Engine
Python WebSocket server
│
┌──────┴───────┐
▼ ▼
faster-whisper diart
transcription diarization
│ │
└──────┬───────┘
▼
speaker-aware segments
│
┌──────┴──────────┐
▼ ▼
Live transcript Overlay window
```
The local speech engine is **not** the future MeetVault backend. It is a local ML side process used by the desktop demo.
---
## Project structure
```text
MeetVault-Desktop-Demo/
├── src/ React desktop UI
│ ├── components/
│ ├── lib/
│ │ ├── capture.ts microphone/system capture + MediaRecorder
│ │ ├── speechClient.ts local WebSocket speech client
│ │ ├── liveBus.ts main-window ↔ overlay events
│ │ └── localFiles.ts local recording/transcript persistence
│ ├── pages/
│ │ ├── Dashboard.tsx
│ │ ├── LiveMeeting.tsx
│ │ ├── Meetings.tsx
│ │ ├── MeetingDetails.tsx
│ │ ├── Settings.tsx
│ │ └── Overlay.tsx
│ └── store/
├── src-tauri/ Tauri 2 desktop shell
├── speech-engine/ local Whisper + diart process
├── docs/ supplied source wireframes/README
└── scripts/
```
---
## Prerequisites
### Desktop UI
- Node.js 20+
- Rust toolchain compatible with your installed Tauri 2 release
- Platform-specific Tauri prerequisites
### Real speech engine
Recommended:
- Python 3.10 or 3.11
- ffmpeg
- PortAudio
- libsndfile
- Python packages in `speech-engine/requirements.txt`
- Hugging Face/pyannote model access required by your diart setup
For UI-only testing, use mock speech mode instead.
---
## 1. Install frontend dependencies
```bash
npm install
```
---
## 2. Test the UI in a browser
```bash
npm run dev
```
Open:
```text
http://127.0.0.1:1420
```
Browser mode is useful for UI development. Tauri-only functions such as the native overlay window and persistent AppLocalData storage use reduced fallbacks in browser mode.
---
## 3. Install the local speech engine
Create a Python 3.10/3.11 virtual environment inside `speech-engine`:
```bash
cd speech-engine
python -m venv .venv
```
Windows:
```powershell
.venv\Scripts\activate
pip install -r requirements.txt
```
Linux/macOS:
```bash
source .venv/bin/activate
pip install -r requirements.txt
```
### Real mode
```bash
python server.py
```
### Mock mode
```bash
python server.py --mock
```
The application expects the local engine at:
```text
ws://127.0.0.1:8765/ws/live
```
This can be changed from Settings.
---
## 4. Run as a desktop application
From the repository root:
```bash
npm run desktop
```
On Windows, convenience scripts are also provided:
```powershell
scripts\start-demo.ps1
```
or for UI testing with mock speech:
```powershell
scripts\start-demo-mock.ps1
```
---
# Live meeting behavior
When a meeting starts:
1. MeetVault requests microphone access.
2. If **Capture system audio / screen** is enabled, it also opens the system screen/audio picker.
3. Audio is mixed locally.
4. The mixed audio is recorded through `MediaRecorder`.
5. The same audio is converted to mono 16 kHz signed PCM.
6. PCM is streamed to `speech-engine/server.py`.
7. Whisper produces timestamped text.
8. diart produces speaker time ranges.
9. The local engine merges the two timelines.
10. The UI receives segments such as:
```json
{
"speakerId": "speaker_1",
"start": 18.2,
"end": 21.7,
"text": "We should finish the API before Friday.",
"final": true
}
```
11. The user can rename `speaker_1` to `Alice` without changing the underlying diarization ID.
12. The main transcript and overlay immediately display the renamed speaker.
---
# Live overlay
The Tauri build creates a second transparent window.
Properties:
- Always on top
- Transparent
- Resizable
- No decorations
- Hidden from the taskbar
- Shows the latest transcript lines
- Speaker name can be clicked to rename
- Receives the same live transcript event stream as the main window
The overlay is opened from **Show overlay** on the live meeting screen.
---
# Local files
There is no cloud storage in this demo.
In Tauri mode, media is stored below the application's local-data directory:
```text
meetings/{meeting_uuid}/recording.webm
meetings/{meeting_uuid}/meeting.json
meetings/{meeting_uuid}/transcript.json
```
The Meeting Details page can play the stored video/audio file.
If screen capture is enabled, the recording normally contains video + mixed audio.
If only microphone capture is enabled, the recording is audio-only.
---
# Video playback
The Meeting Details screen includes a real HTML5 video/audio player.
The Tauri asset protocol is scoped only to MeetVault's local `meetings` directory.
This allows locally recorded media to be played without uploading it anywhere.
---
# Settings
The Settings screen is functional and persisted locally.
Current settings include:
### General
- Launch at startup preference
- Keep in tray preference
- Global hotkey value
- Theme preference
The startup/tray/hotkey preferences are stored now; OS registration is intentionally left for the packaging phase.
### Audio
- Microphone device ID
- System audio/screen capture toggle
- Microphone level monitoring
### Transcription
- Whisper model: `base`, `small`, `medium`
- Compute device: auto/CPU/CUDA
- Language
- Speaker diarization toggle
### Overlay
- Enabled
- Opacity
- Font size
- Visible line count
### Speech engine
- Real Whisper + diart mode
- Mock mode
- Local WebSocket URL
---
# Backend placeholder
This demo deliberately has no remote backend implementation.
Any backend-only action returns/displays:
```text
Processing (The job). Please wait a moment
```
This currently applies to example actions such as:
- Generate/regenerate summary
- Create reminder
- Action-item processing
The UI is structured so a real API can replace `src/lib/backend.ts` later.
---
# Current limitations
## System audio
System-audio capture depends on the operating system, WebView implementation, selected source, and whether that source permits audio capture.
Microphone capture is the reliable baseline for the demo.
## diart model setup
Real diart execution can require pyannote model terms/authentication. The exact dependency combinations can also be sensitive to Python, PyTorch, torchaudio, and platform versions.
Use the recommended Python version and validate the speech-engine environment on the target development machine before packaging.
## Realtime accuracy
Speaker diarization can temporarily swap speaker labels, especially with:
- overlapping speakers
- short utterances
- background noise
- similar voices
The rename function changes the display name but does not perform biometric identity recognition.
## Long recordings
The media recorder writes the final Blob when the meeting ends in this initial demo. Before production use, change this to incremental/chunked file writing so multi-hour meetings do not accumulate the entire media Blob in memory.
---
# Recommended next development steps
1. Validate real Whisper + diart on the target Windows machine.
2. Replace final-Blob recording with incremental local file writes.
3. Add proper microphone device enumeration.
4. Add native Windows loopback audio capture if system-audio reliability is required.
5. Add session crash recovery.
6. Add transcript segment editing/reassignment.
7. Package the Python speech engine as a sidecar or replace it with native whisper.cpp + a native/ONNX diarization runtime.
8. Implement the real backend jobs/API later.
9. Replace local-only media with SeaweedFS when backend/storage work begins.
---
# Demo design decisions from the supplied specification
```text
Login / subscription REMOVED
Remote backend PLACEHOLDER ONLY
Backend message "Processing (The job). Please wait a moment"
Media storage LOCAL
Video playback INCLUDED
Live transcription INCLUDED
Speaker diarization INCLUDED
Live speaker rename INCLUDED
Realtime overlay INCLUDED
Settings INCLUDED
```