Files
supersanta 654033fba3 Add Phase 0 test harness (vitest + ESLint) and fix build config
- Add vitest unit tests for Zustand app store state transitions (6 tests)
- Add mock speech-engine WebSocket protocol integration test (spawns server.py --mock on a free port)
- Add ESLint 10 flat config with typescript-eslint; add typecheck/lint/test scripts to package.json
- Fix pre-existing tsc -b breakage via noEmit in tsconfig.node.json; add dedicated tsconfig.tests.json project for Node-side tests
- Track src-tauri/Cargo.lock for reproducible Rust builds; enable protocol-asset Tauri feature
- Ignore generated artifacts (*.tsbuildinfo, src-tauri/gen/, emitted vite.config.js/.d.ts); remove stale emitted vite config copies
- Update AGENTS.md agent guide (demo scope lock, roadmap, verification commands)
- Add HANDOFF.md documenting Phase 0 state and remaining steps for the next agent
2026-10-03 22:33:47 +07:00
..
2026-09-30 00:50:44 +07:00
2026-09-30 00:50:44 +07:00
2026-09-30 00:50:44 +07:00
2026-09-30 00:50:44 +07:00
2026-09-30 00:50:44 +07:00
2026-09-30 00:50:44 +07:00
2026-09-30 00:50:44 +07:00
2026-09-30 00:50:44 +07:00

MeetVault Desktop Demo

This repository is a runnable desktop-demo scaffold based on the supplied MeetVault wireframes and desktop architecture notes.

Demo scope

Included now:

  • Tauri + React + TypeScript desktop UI
  • Dashboard
  • Meetings list
  • Meeting details
  • Local video/audio playback
  • Live meeting screen
  • Real-time Whisper transcription through a local Python speech engine
  • Real-time speaker diarization through diart
  • Speaker 1, Speaker 2, ... labels
  • Live speaker rename
  • Transparent always-on-top overlay window
  • Microphone capture
  • Optional system-audio/screen capture through the WebView capture API
  • Local recording files
  • Local transcript data
  • Functional Settings screen
  • Mock speech-engine mode for UI testing without ML models

Not included in this demo:

  • User login
  • User identification/account verification
  • Subscription/payment
  • PostgreSQL
  • SeaweedFS
  • Remote cloud storage
  • Real summary/task/reminder backend

The backend-facing demo state is intentionally:

Processing (The job). Please wait a moment

Summary/task/reminder controls show this processing state but do not call a production backend yet.


Architecture

Desktop UI (Tauri + React)
        │
        ├── microphone / optional system audio + screen
        │
        ├── MediaRecorder ───────────► local recording.webm
        │
        └── 16 kHz PCM
                 │
                 ▼
        Local Speech Engine
        Python WebSocket server
                 │
          ┌──────┴───────┐
          ▼              ▼
   faster-whisper       diart
   transcription       diarization
          │              │
          └──────┬───────┘
                 ▼
        speaker-aware segments
                 │
          ┌──────┴──────────┐
          ▼                 ▼
   Live transcript     Overlay window

The local speech engine is not the future MeetVault backend. It is a local ML side process used by the desktop demo.


Project structure

MeetVault-Desktop-Demo/
├── src/                       React desktop UI
│   ├── components/
│   ├── lib/
│   │   ├── capture.ts         microphone/system capture + MediaRecorder
│   │   ├── speechClient.ts    local WebSocket speech client
│   │   ├── liveBus.ts         main-window ↔ overlay events
│   │   └── localFiles.ts      local recording/transcript persistence
│   ├── pages/
│   │   ├── Dashboard.tsx
│   │   ├── LiveMeeting.tsx
│   │   ├── Meetings.tsx
│   │   ├── MeetingDetails.tsx
│   │   ├── Settings.tsx
│   │   └── Overlay.tsx
│   └── store/
├── src-tauri/                 Tauri 2 desktop shell
├── speech-engine/             local Whisper + diart process
├── docs/                      supplied source wireframes/README
└── scripts/

Prerequisites

Desktop UI

  • Node.js 20+
  • Rust toolchain compatible with your installed Tauri 2 release
  • Platform-specific Tauri prerequisites

Real speech engine

Recommended:

  • Python 3.10 or 3.11
  • ffmpeg
  • PortAudio
  • libsndfile
  • Python packages in speech-engine/requirements.txt
  • Hugging Face/pyannote model access required by your diart setup

For UI-only testing, use mock speech mode instead.


1. Install frontend dependencies

npm install

2. Test the UI in a browser

npm run dev

Open:

http://127.0.0.1:1420

Browser mode is useful for UI development. Tauri-only functions such as the native overlay window and persistent AppLocalData storage use reduced fallbacks in browser mode.


3. Install the local speech engine

Create a Python 3.10/3.11 virtual environment inside speech-engine:

cd speech-engine
python -m venv .venv

Windows:

.venv\Scripts\activate
pip install -r requirements.txt

Linux/macOS:

source .venv/bin/activate
pip install -r requirements.txt

Real mode

python server.py

Mock mode

python server.py --mock

The application expects the local engine at:

ws://127.0.0.1:8765/ws/live

This can be changed from Settings.


4. Run as a desktop application

From the repository root:

npm run desktop

On Windows, convenience scripts are also provided:

scripts\start-demo.ps1

or for UI testing with mock speech:

scripts\start-demo-mock.ps1

Live meeting behavior

When a meeting starts:

  1. MeetVault requests microphone access.
  2. If Capture system audio / screen is enabled, it also opens the system screen/audio picker.
  3. Audio is mixed locally.
  4. The mixed audio is recorded through MediaRecorder.
  5. The same audio is converted to mono 16 kHz signed PCM.
  6. PCM is streamed to speech-engine/server.py.
  7. Whisper produces timestamped text.
  8. diart produces speaker time ranges.
  9. The local engine merges the two timelines.
  10. The UI receives segments such as:
{
  "speakerId": "speaker_1",
  "start": 18.2,
  "end": 21.7,
  "text": "We should finish the API before Friday.",
  "final": true
}
  1. The user can rename speaker_1 to Alice without changing the underlying diarization ID.
  2. The main transcript and overlay immediately display the renamed speaker.

Live overlay

The Tauri build creates a second transparent window.

Properties:

  • Always on top
  • Transparent
  • Resizable
  • No decorations
  • Hidden from the taskbar
  • Shows the latest transcript lines
  • Speaker name can be clicked to rename
  • Receives the same live transcript event stream as the main window

The overlay is opened from Show overlay on the live meeting screen.


Local files

There is no cloud storage in this demo.

In Tauri mode, media is stored below the application's local-data directory:

meetings/{meeting_uuid}/recording.webm
meetings/{meeting_uuid}/meeting.json
meetings/{meeting_uuid}/transcript.json

The Meeting Details page can play the stored video/audio file.

If screen capture is enabled, the recording normally contains video + mixed audio.

If only microphone capture is enabled, the recording is audio-only.


Video playback

The Meeting Details screen includes a real HTML5 video/audio player.

The Tauri asset protocol is scoped only to MeetVault's local meetings directory.

This allows locally recorded media to be played without uploading it anywhere.


Settings

The Settings screen is functional and persisted locally.

Current settings include:

General

  • Launch at startup preference
  • Keep in tray preference
  • Global hotkey value
  • Theme preference

The startup/tray/hotkey preferences are stored now; OS registration is intentionally left for the packaging phase.

Audio

  • Microphone device ID
  • System audio/screen capture toggle
  • Microphone level monitoring

Transcription

  • Whisper model: base, small, medium
  • Compute device: auto/CPU/CUDA
  • Language
  • Speaker diarization toggle

Overlay

  • Enabled
  • Opacity
  • Font size
  • Visible line count

Speech engine

  • Real Whisper + diart mode
  • Mock mode
  • Local WebSocket URL

Backend placeholder

This demo deliberately has no remote backend implementation.

Any backend-only action returns/displays:

Processing (The job). Please wait a moment

This currently applies to example actions such as:

  • Generate/regenerate summary
  • Create reminder
  • Action-item processing

The UI is structured so a real API can replace src/lib/backend.ts later.


Current limitations

System audio

System-audio capture depends on the operating system, WebView implementation, selected source, and whether that source permits audio capture.

Microphone capture is the reliable baseline for the demo.

diart model setup

Real diart execution can require pyannote model terms/authentication. The exact dependency combinations can also be sensitive to Python, PyTorch, torchaudio, and platform versions.

Use the recommended Python version and validate the speech-engine environment on the target development machine before packaging.

Realtime accuracy

Speaker diarization can temporarily swap speaker labels, especially with:

  • overlapping speakers
  • short utterances
  • background noise
  • similar voices

The rename function changes the display name but does not perform biometric identity recognition.

Long recordings

The media recorder writes the final Blob when the meeting ends in this initial demo. Before production use, change this to incremental/chunked file writing so multi-hour meetings do not accumulate the entire media Blob in memory.


Recommended next development steps

  1. Validate real Whisper + diart on the target Windows machine.
  2. Replace final-Blob recording with incremental local file writes.
  3. Add proper microphone device enumeration.
  4. Add native Windows loopback audio capture if system-audio reliability is required.
  5. Add session crash recovery.
  6. Add transcript segment editing/reassignment.
  7. Package the Python speech engine as a sidecar or replace it with native whisper.cpp + a native/ONNX diarization runtime.
  8. Implement the real backend jobs/API later.
  9. Replace local-only media with SeaweedFS when backend/storage work begins.

Demo design decisions from the supplied specification

Login / subscription          REMOVED
Remote backend                PLACEHOLDER ONLY
Backend message               "Processing (The job). Please wait a moment"
Media storage                 LOCAL
Video playback                INCLUDED
Live transcription            INCLUDED
Speaker diarization           INCLUDED
Live speaker rename           INCLUDED
Realtime overlay              INCLUDED
Settings                      INCLUDED