# MeetVault Desktop Demo This repository is a runnable desktop-demo scaffold based on the supplied MeetVault wireframes and desktop architecture notes. ## Demo scope Included now: - Tauri + React + TypeScript desktop UI - Dashboard - Meetings list - Meeting details - Local video/audio playback - Live meeting screen - Real-time Whisper transcription through a local Python speech engine - Real-time speaker diarization through diart - `Speaker 1`, `Speaker 2`, ... labels - Live speaker rename - Transparent always-on-top overlay window - Microphone capture - Optional system-audio/screen capture through the WebView capture API - Local recording files - Local transcript data - Functional Settings screen - Mock speech-engine mode for UI testing without ML models Not included in this demo: - User login - User identification/account verification - Subscription/payment - PostgreSQL - SeaweedFS - Remote cloud storage - Real summary/task/reminder backend The backend-facing demo state is intentionally: ```text Processing (The job). Please wait a moment ``` Summary/task/reminder controls show this processing state but do not call a production backend yet. --- ## Architecture ```text Desktop UI (Tauri + React) │ ├── microphone / optional system audio + screen │ ├── MediaRecorder ───────────► local recording.webm │ └── 16 kHz PCM │ ▼ Local Speech Engine Python WebSocket server │ ┌──────┴───────┐ ▼ ▼ faster-whisper diart transcription diarization │ │ └──────┬───────┘ ▼ speaker-aware segments │ ┌──────┴──────────┐ ▼ ▼ Live transcript Overlay window ``` The local speech engine is **not** the future MeetVault backend. It is a local ML side process used by the desktop demo. --- ## Project structure ```text MeetVault-Desktop-Demo/ ├── src/ React desktop UI │ ├── components/ │ ├── lib/ │ │ ├── capture.ts microphone/system capture + MediaRecorder │ │ ├── speechClient.ts local WebSocket speech client │ │ ├── liveBus.ts main-window ↔ overlay events │ │ └── localFiles.ts local recording/transcript persistence │ ├── pages/ │ │ ├── Dashboard.tsx │ │ ├── LiveMeeting.tsx │ │ ├── Meetings.tsx │ │ ├── MeetingDetails.tsx │ │ ├── Settings.tsx │ │ └── Overlay.tsx │ └── store/ ├── src-tauri/ Tauri 2 desktop shell ├── speech-engine/ local Whisper + diart process ├── docs/ supplied source wireframes/README └── scripts/ ``` --- ## Prerequisites ### Desktop UI - Node.js 20+ - Rust toolchain compatible with your installed Tauri 2 release - Platform-specific Tauri prerequisites ### Real speech engine Recommended: - Python 3.10 or 3.11 - ffmpeg - PortAudio - libsndfile - Python packages in `speech-engine/requirements.txt` - Hugging Face/pyannote model access required by your diart setup For UI-only testing, use mock speech mode instead. --- ## 1. Install frontend dependencies ```bash npm install ``` --- ## 2. Test the UI in a browser ```bash npm run dev ``` Open: ```text http://127.0.0.1:1420 ``` Browser mode is useful for UI development. Tauri-only functions such as the native overlay window and persistent AppLocalData storage use reduced fallbacks in browser mode. --- ## 3. Install the local speech engine Create a Python 3.10/3.11 virtual environment inside `speech-engine`: ```bash cd speech-engine python -m venv .venv ``` Windows: ```powershell .venv\Scripts\activate pip install -r requirements.txt ``` Linux/macOS: ```bash source .venv/bin/activate pip install -r requirements.txt ``` ### Real mode ```bash python server.py ``` ### Mock mode ```bash python server.py --mock ``` The application expects the local engine at: ```text ws://127.0.0.1:8765/ws/live ``` This can be changed from Settings. --- ## 4. Run as a desktop application From the repository root: ```bash npm run desktop ``` On Windows, convenience scripts are also provided: ```powershell scripts\start-demo.ps1 ``` or for UI testing with mock speech: ```powershell scripts\start-demo-mock.ps1 ``` --- # Live meeting behavior When a meeting starts: 1. MeetVault requests microphone access. 2. If **Capture system audio / screen** is enabled, it also opens the system screen/audio picker. 3. Audio is mixed locally. 4. The mixed audio is recorded through `MediaRecorder`. 5. The same audio is converted to mono 16 kHz signed PCM. 6. PCM is streamed to `speech-engine/server.py`. 7. Whisper produces timestamped text. 8. diart produces speaker time ranges. 9. The local engine merges the two timelines. 10. The UI receives segments such as: ```json { "speakerId": "speaker_1", "start": 18.2, "end": 21.7, "text": "We should finish the API before Friday.", "final": true } ``` 11. The user can rename `speaker_1` to `Alice` without changing the underlying diarization ID. 12. The main transcript and overlay immediately display the renamed speaker. --- # Live overlay The Tauri build creates a second transparent window. Properties: - Always on top - Transparent - Resizable - No decorations - Hidden from the taskbar - Shows the latest transcript lines - Speaker name can be clicked to rename - Receives the same live transcript event stream as the main window The overlay is opened from **Show overlay** on the live meeting screen. --- # Local files There is no cloud storage in this demo. In Tauri mode, media is stored below the application's local-data directory: ```text meetings/{meeting_uuid}/recording.webm meetings/{meeting_uuid}/meeting.json meetings/{meeting_uuid}/transcript.json ``` The Meeting Details page can play the stored video/audio file. If screen capture is enabled, the recording normally contains video + mixed audio. If only microphone capture is enabled, the recording is audio-only. --- # Video playback The Meeting Details screen includes a real HTML5 video/audio player. The Tauri asset protocol is scoped only to MeetVault's local `meetings` directory. This allows locally recorded media to be played without uploading it anywhere. --- # Settings The Settings screen is functional and persisted locally. Current settings include: ### General - Launch at startup preference - Keep in tray preference - Global hotkey value - Theme preference The startup/tray/hotkey preferences are stored now; OS registration is intentionally left for the packaging phase. ### Audio - Microphone device ID - System audio/screen capture toggle - Microphone level monitoring ### Transcription - Whisper model: `base`, `small`, `medium` - Compute device: auto/CPU/CUDA - Language - Speaker diarization toggle ### Overlay - Enabled - Opacity - Font size - Visible line count ### Speech engine - Real Whisper + diart mode - Mock mode - Local WebSocket URL --- # Backend placeholder This demo deliberately has no remote backend implementation. Any backend-only action returns/displays: ```text Processing (The job). Please wait a moment ``` This currently applies to example actions such as: - Generate/regenerate summary - Create reminder - Action-item processing The UI is structured so a real API can replace `src/lib/backend.ts` later. --- # Current limitations ## System audio System-audio capture depends on the operating system, WebView implementation, selected source, and whether that source permits audio capture. Microphone capture is the reliable baseline for the demo. ## diart model setup Real diart execution can require pyannote model terms/authentication. The exact dependency combinations can also be sensitive to Python, PyTorch, torchaudio, and platform versions. Use the recommended Python version and validate the speech-engine environment on the target development machine before packaging. ## Realtime accuracy Speaker diarization can temporarily swap speaker labels, especially with: - overlapping speakers - short utterances - background noise - similar voices The rename function changes the display name but does not perform biometric identity recognition. ## Long recordings The media recorder writes the final Blob when the meeting ends in this initial demo. Before production use, change this to incremental/chunked file writing so multi-hour meetings do not accumulate the entire media Blob in memory. --- # Recommended next development steps 1. Validate real Whisper + diart on the target Windows machine. 2. Replace final-Blob recording with incremental local file writes. 3. Add proper microphone device enumeration. 4. Add native Windows loopback audio capture if system-audio reliability is required. 5. Add session crash recovery. 6. Add transcript segment editing/reassignment. 7. Package the Python speech engine as a sidecar or replace it with native whisper.cpp + a native/ONNX diarization runtime. 8. Implement the real backend jobs/API later. 9. Replace local-only media with SeaweedFS when backend/storage work begins. --- # Demo design decisions from the supplied specification ```text Login / subscription REMOVED Remote backend PLACEHOLDER ONLY Backend message "Processing (The job). Please wait a moment" Media storage LOCAL Video playback INCLUDED Live transcription INCLUDED Speaker diarization INCLUDED Live speaker rename INCLUDED Realtime overlay INCLUDED Settings INCLUDED ```