- Add vitest unit tests for Zustand app store state transitions (6 tests) - Add mock speech-engine WebSocket protocol integration test (spawns server.py --mock on a free port) - Add ESLint 10 flat config with typescript-eslint; add typecheck/lint/test scripts to package.json - Fix pre-existing tsc -b breakage via noEmit in tsconfig.node.json; add dedicated tsconfig.tests.json project for Node-side tests - Track src-tauri/Cargo.lock for reproducible Rust builds; enable protocol-asset Tauri feature - Ignore generated artifacts (*.tsbuildinfo, src-tauri/gen/, emitted vite.config.js/.d.ts); remove stale emitted vite config copies - Update AGENTS.md agent guide (demo scope lock, roadmap, verification commands) - Add HANDOFF.md documenting Phase 0 state and remaining steps for the next agent
MeetVault Desktop Demo
This repository is a runnable desktop-demo scaffold based on the supplied MeetVault wireframes and desktop architecture notes.
Demo scope
Included now:
- Tauri + React + TypeScript desktop UI
- Dashboard
- Meetings list
- Meeting details
- Local video/audio playback
- Live meeting screen
- Real-time Whisper transcription through a local Python speech engine
- Real-time speaker diarization through diart
Speaker 1,Speaker 2, ... labels- Live speaker rename
- Transparent always-on-top overlay window
- Microphone capture
- Optional system-audio/screen capture through the WebView capture API
- Local recording files
- Local transcript data
- Functional Settings screen
- Mock speech-engine mode for UI testing without ML models
Not included in this demo:
- User login
- User identification/account verification
- Subscription/payment
- PostgreSQL
- SeaweedFS
- Remote cloud storage
- Real summary/task/reminder backend
The backend-facing demo state is intentionally:
Processing (The job). Please wait a moment
Summary/task/reminder controls show this processing state but do not call a production backend yet.
Architecture
Desktop UI (Tauri + React)
│
├── microphone / optional system audio + screen
│
├── MediaRecorder ───────────► local recording.webm
│
└── 16 kHz PCM
│
▼
Local Speech Engine
Python WebSocket server
│
┌──────┴───────┐
▼ ▼
faster-whisper diart
transcription diarization
│ │
└──────┬───────┘
▼
speaker-aware segments
│
┌──────┴──────────┐
▼ ▼
Live transcript Overlay window
The local speech engine is not the future MeetVault backend. It is a local ML side process used by the desktop demo.
Project structure
MeetVault-Desktop-Demo/
├── src/ React desktop UI
│ ├── components/
│ ├── lib/
│ │ ├── capture.ts microphone/system capture + MediaRecorder
│ │ ├── speechClient.ts local WebSocket speech client
│ │ ├── liveBus.ts main-window ↔ overlay events
│ │ └── localFiles.ts local recording/transcript persistence
│ ├── pages/
│ │ ├── Dashboard.tsx
│ │ ├── LiveMeeting.tsx
│ │ ├── Meetings.tsx
│ │ ├── MeetingDetails.tsx
│ │ ├── Settings.tsx
│ │ └── Overlay.tsx
│ └── store/
├── src-tauri/ Tauri 2 desktop shell
├── speech-engine/ local Whisper + diart process
├── docs/ supplied source wireframes/README
└── scripts/
Prerequisites
Desktop UI
- Node.js 20+
- Rust toolchain compatible with your installed Tauri 2 release
- Platform-specific Tauri prerequisites
Real speech engine
Recommended:
- Python 3.10 or 3.11
- ffmpeg
- PortAudio
- libsndfile
- Python packages in
speech-engine/requirements.txt - Hugging Face/pyannote model access required by your diart setup
For UI-only testing, use mock speech mode instead.
1. Install frontend dependencies
npm install
2. Test the UI in a browser
npm run dev
Open:
http://127.0.0.1:1420
Browser mode is useful for UI development. Tauri-only functions such as the native overlay window and persistent AppLocalData storage use reduced fallbacks in browser mode.
3. Install the local speech engine
Create a Python 3.10/3.11 virtual environment inside speech-engine:
cd speech-engine
python -m venv .venv
Windows:
.venv\Scripts\activate
pip install -r requirements.txt
Linux/macOS:
source .venv/bin/activate
pip install -r requirements.txt
Real mode
python server.py
Mock mode
python server.py --mock
The application expects the local engine at:
ws://127.0.0.1:8765/ws/live
This can be changed from Settings.
4. Run as a desktop application
From the repository root:
npm run desktop
On Windows, convenience scripts are also provided:
scripts\start-demo.ps1
or for UI testing with mock speech:
scripts\start-demo-mock.ps1
Live meeting behavior
When a meeting starts:
- MeetVault requests microphone access.
- If Capture system audio / screen is enabled, it also opens the system screen/audio picker.
- Audio is mixed locally.
- The mixed audio is recorded through
MediaRecorder. - The same audio is converted to mono 16 kHz signed PCM.
- PCM is streamed to
speech-engine/server.py. - Whisper produces timestamped text.
- diart produces speaker time ranges.
- The local engine merges the two timelines.
- The UI receives segments such as:
{
"speakerId": "speaker_1",
"start": 18.2,
"end": 21.7,
"text": "We should finish the API before Friday.",
"final": true
}
- The user can rename
speaker_1toAlicewithout changing the underlying diarization ID. - The main transcript and overlay immediately display the renamed speaker.
Live overlay
The Tauri build creates a second transparent window.
Properties:
- Always on top
- Transparent
- Resizable
- No decorations
- Hidden from the taskbar
- Shows the latest transcript lines
- Speaker name can be clicked to rename
- Receives the same live transcript event stream as the main window
The overlay is opened from Show overlay on the live meeting screen.
Local files
There is no cloud storage in this demo.
In Tauri mode, media is stored below the application's local-data directory:
meetings/{meeting_uuid}/recording.webm
meetings/{meeting_uuid}/meeting.json
meetings/{meeting_uuid}/transcript.json
The Meeting Details page can play the stored video/audio file.
If screen capture is enabled, the recording normally contains video + mixed audio.
If only microphone capture is enabled, the recording is audio-only.
Video playback
The Meeting Details screen includes a real HTML5 video/audio player.
The Tauri asset protocol is scoped only to MeetVault's local meetings directory.
This allows locally recorded media to be played without uploading it anywhere.
Settings
The Settings screen is functional and persisted locally.
Current settings include:
General
- Launch at startup preference
- Keep in tray preference
- Global hotkey value
- Theme preference
The startup/tray/hotkey preferences are stored now; OS registration is intentionally left for the packaging phase.
Audio
- Microphone device ID
- System audio/screen capture toggle
- Microphone level monitoring
Transcription
- Whisper model:
base,small,medium - Compute device: auto/CPU/CUDA
- Language
- Speaker diarization toggle
Overlay
- Enabled
- Opacity
- Font size
- Visible line count
Speech engine
- Real Whisper + diart mode
- Mock mode
- Local WebSocket URL
Backend placeholder
This demo deliberately has no remote backend implementation.
Any backend-only action returns/displays:
Processing (The job). Please wait a moment
This currently applies to example actions such as:
- Generate/regenerate summary
- Create reminder
- Action-item processing
The UI is structured so a real API can replace src/lib/backend.ts later.
Current limitations
System audio
System-audio capture depends on the operating system, WebView implementation, selected source, and whether that source permits audio capture.
Microphone capture is the reliable baseline for the demo.
diart model setup
Real diart execution can require pyannote model terms/authentication. The exact dependency combinations can also be sensitive to Python, PyTorch, torchaudio, and platform versions.
Use the recommended Python version and validate the speech-engine environment on the target development machine before packaging.
Realtime accuracy
Speaker diarization can temporarily swap speaker labels, especially with:
- overlapping speakers
- short utterances
- background noise
- similar voices
The rename function changes the display name but does not perform biometric identity recognition.
Long recordings
The media recorder writes the final Blob when the meeting ends in this initial demo. Before production use, change this to incremental/chunked file writing so multi-hour meetings do not accumulate the entire media Blob in memory.
Recommended next development steps
- Validate real Whisper + diart on the target Windows machine.
- Replace final-Blob recording with incremental local file writes.
- Add proper microphone device enumeration.
- Add native Windows loopback audio capture if system-audio reliability is required.
- Add session crash recovery.
- Add transcript segment editing/reassignment.
- Package the Python speech engine as a sidecar or replace it with native whisper.cpp + a native/ONNX diarization runtime.
- Implement the real backend jobs/API later.
- Replace local-only media with SeaweedFS when backend/storage work begins.
Demo design decisions from the supplied specification
Login / subscription REMOVED
Remote backend PLACEHOLDER ONLY
Backend message "Processing (The job). Please wait a moment"
Media storage LOCAL
Video playback INCLUDED
Live transcription INCLUDED
Speaker diarization INCLUDED
Live speaker rename INCLUDED
Realtime overlay INCLUDED
Settings INCLUDED