Update agent
This commit is contained in:
421
Wireframes/Desktop/MeetVault-Desktop-Demo/README.md
Normal file
421
Wireframes/Desktop/MeetVault-Desktop-Demo/README.md
Normal file
@@ -0,0 +1,421 @@
|
||||
# MeetVault Desktop Demo
|
||||
|
||||
This repository is a runnable desktop-demo scaffold based on the supplied MeetVault wireframes and desktop architecture notes.
|
||||
|
||||
## Demo scope
|
||||
|
||||
Included now:
|
||||
|
||||
- Tauri + React + TypeScript desktop UI
|
||||
- Dashboard
|
||||
- Meetings list
|
||||
- Meeting details
|
||||
- Local video/audio playback
|
||||
- Live meeting screen
|
||||
- Real-time Whisper transcription through a local Python speech engine
|
||||
- Real-time speaker diarization through diart
|
||||
- `Speaker 1`, `Speaker 2`, ... labels
|
||||
- Live speaker rename
|
||||
- Transparent always-on-top overlay window
|
||||
- Microphone capture
|
||||
- Optional system-audio/screen capture through the WebView capture API
|
||||
- Local recording files
|
||||
- Local transcript data
|
||||
- Functional Settings screen
|
||||
- Mock speech-engine mode for UI testing without ML models
|
||||
|
||||
Not included in this demo:
|
||||
|
||||
- User login
|
||||
- User identification/account verification
|
||||
- Subscription/payment
|
||||
- PostgreSQL
|
||||
- SeaweedFS
|
||||
- Remote cloud storage
|
||||
- Real summary/task/reminder backend
|
||||
|
||||
The backend-facing demo state is intentionally:
|
||||
|
||||
```text
|
||||
Processing (The job). Please wait a moment
|
||||
```
|
||||
|
||||
Summary/task/reminder controls show this processing state but do not call a production backend yet.
|
||||
|
||||
---
|
||||
|
||||
## Architecture
|
||||
|
||||
```text
|
||||
Desktop UI (Tauri + React)
|
||||
│
|
||||
├── microphone / optional system audio + screen
|
||||
│
|
||||
├── MediaRecorder ───────────► local recording.webm
|
||||
│
|
||||
└── 16 kHz PCM
|
||||
│
|
||||
▼
|
||||
Local Speech Engine
|
||||
Python WebSocket server
|
||||
│
|
||||
┌──────┴───────┐
|
||||
▼ ▼
|
||||
faster-whisper diart
|
||||
transcription diarization
|
||||
│ │
|
||||
└──────┬───────┘
|
||||
▼
|
||||
speaker-aware segments
|
||||
│
|
||||
┌──────┴──────────┐
|
||||
▼ ▼
|
||||
Live transcript Overlay window
|
||||
```
|
||||
|
||||
The local speech engine is **not** the future MeetVault backend. It is a local ML side process used by the desktop demo.
|
||||
|
||||
---
|
||||
|
||||
## Project structure
|
||||
|
||||
```text
|
||||
MeetVault-Desktop-Demo/
|
||||
├── src/ React desktop UI
|
||||
│ ├── components/
|
||||
│ ├── lib/
|
||||
│ │ ├── capture.ts microphone/system capture + MediaRecorder
|
||||
│ │ ├── speechClient.ts local WebSocket speech client
|
||||
│ │ ├── liveBus.ts main-window ↔ overlay events
|
||||
│ │ └── localFiles.ts local recording/transcript persistence
|
||||
│ ├── pages/
|
||||
│ │ ├── Dashboard.tsx
|
||||
│ │ ├── LiveMeeting.tsx
|
||||
│ │ ├── Meetings.tsx
|
||||
│ │ ├── MeetingDetails.tsx
|
||||
│ │ ├── Settings.tsx
|
||||
│ │ └── Overlay.tsx
|
||||
│ └── store/
|
||||
├── src-tauri/ Tauri 2 desktop shell
|
||||
├── speech-engine/ local Whisper + diart process
|
||||
├── docs/ supplied source wireframes/README
|
||||
└── scripts/
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Prerequisites
|
||||
|
||||
### Desktop UI
|
||||
|
||||
- Node.js 20+
|
||||
- Rust toolchain compatible with your installed Tauri 2 release
|
||||
- Platform-specific Tauri prerequisites
|
||||
|
||||
### Real speech engine
|
||||
|
||||
Recommended:
|
||||
|
||||
- Python 3.10 or 3.11
|
||||
- ffmpeg
|
||||
- PortAudio
|
||||
- libsndfile
|
||||
- Python packages in `speech-engine/requirements.txt`
|
||||
- Hugging Face/pyannote model access required by your diart setup
|
||||
|
||||
For UI-only testing, use mock speech mode instead.
|
||||
|
||||
---
|
||||
|
||||
## 1. Install frontend dependencies
|
||||
|
||||
```bash
|
||||
npm install
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 2. Test the UI in a browser
|
||||
|
||||
```bash
|
||||
npm run dev
|
||||
```
|
||||
|
||||
Open:
|
||||
|
||||
```text
|
||||
http://127.0.0.1:1420
|
||||
```
|
||||
|
||||
Browser mode is useful for UI development. Tauri-only functions such as the native overlay window and persistent AppLocalData storage use reduced fallbacks in browser mode.
|
||||
|
||||
---
|
||||
|
||||
## 3. Install the local speech engine
|
||||
|
||||
Create a Python 3.10/3.11 virtual environment inside `speech-engine`:
|
||||
|
||||
```bash
|
||||
cd speech-engine
|
||||
python -m venv .venv
|
||||
```
|
||||
|
||||
Windows:
|
||||
|
||||
```powershell
|
||||
.venv\Scripts\activate
|
||||
pip install -r requirements.txt
|
||||
```
|
||||
|
||||
Linux/macOS:
|
||||
|
||||
```bash
|
||||
source .venv/bin/activate
|
||||
pip install -r requirements.txt
|
||||
```
|
||||
|
||||
### Real mode
|
||||
|
||||
```bash
|
||||
python server.py
|
||||
```
|
||||
|
||||
### Mock mode
|
||||
|
||||
```bash
|
||||
python server.py --mock
|
||||
```
|
||||
|
||||
The application expects the local engine at:
|
||||
|
||||
```text
|
||||
ws://127.0.0.1:8765/ws/live
|
||||
```
|
||||
|
||||
This can be changed from Settings.
|
||||
|
||||
---
|
||||
|
||||
## 4. Run as a desktop application
|
||||
|
||||
From the repository root:
|
||||
|
||||
```bash
|
||||
npm run desktop
|
||||
```
|
||||
|
||||
On Windows, convenience scripts are also provided:
|
||||
|
||||
```powershell
|
||||
scripts\start-demo.ps1
|
||||
```
|
||||
|
||||
or for UI testing with mock speech:
|
||||
|
||||
```powershell
|
||||
scripts\start-demo-mock.ps1
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
# Live meeting behavior
|
||||
|
||||
When a meeting starts:
|
||||
|
||||
1. MeetVault requests microphone access.
|
||||
2. If **Capture system audio / screen** is enabled, it also opens the system screen/audio picker.
|
||||
3. Audio is mixed locally.
|
||||
4. The mixed audio is recorded through `MediaRecorder`.
|
||||
5. The same audio is converted to mono 16 kHz signed PCM.
|
||||
6. PCM is streamed to `speech-engine/server.py`.
|
||||
7. Whisper produces timestamped text.
|
||||
8. diart produces speaker time ranges.
|
||||
9. The local engine merges the two timelines.
|
||||
10. The UI receives segments such as:
|
||||
|
||||
```json
|
||||
{
|
||||
"speakerId": "speaker_1",
|
||||
"start": 18.2,
|
||||
"end": 21.7,
|
||||
"text": "We should finish the API before Friday.",
|
||||
"final": true
|
||||
}
|
||||
```
|
||||
|
||||
11. The user can rename `speaker_1` to `Alice` without changing the underlying diarization ID.
|
||||
12. The main transcript and overlay immediately display the renamed speaker.
|
||||
|
||||
---
|
||||
|
||||
# Live overlay
|
||||
|
||||
The Tauri build creates a second transparent window.
|
||||
|
||||
Properties:
|
||||
|
||||
- Always on top
|
||||
- Transparent
|
||||
- Resizable
|
||||
- No decorations
|
||||
- Hidden from the taskbar
|
||||
- Shows the latest transcript lines
|
||||
- Speaker name can be clicked to rename
|
||||
- Receives the same live transcript event stream as the main window
|
||||
|
||||
The overlay is opened from **Show overlay** on the live meeting screen.
|
||||
|
||||
---
|
||||
|
||||
# Local files
|
||||
|
||||
There is no cloud storage in this demo.
|
||||
|
||||
In Tauri mode, media is stored below the application's local-data directory:
|
||||
|
||||
```text
|
||||
meetings/{meeting_uuid}/recording.webm
|
||||
meetings/{meeting_uuid}/meeting.json
|
||||
meetings/{meeting_uuid}/transcript.json
|
||||
```
|
||||
|
||||
The Meeting Details page can play the stored video/audio file.
|
||||
|
||||
If screen capture is enabled, the recording normally contains video + mixed audio.
|
||||
|
||||
If only microphone capture is enabled, the recording is audio-only.
|
||||
|
||||
---
|
||||
|
||||
# Video playback
|
||||
|
||||
The Meeting Details screen includes a real HTML5 video/audio player.
|
||||
|
||||
The Tauri asset protocol is scoped only to MeetVault's local `meetings` directory.
|
||||
|
||||
This allows locally recorded media to be played without uploading it anywhere.
|
||||
|
||||
---
|
||||
|
||||
# Settings
|
||||
|
||||
The Settings screen is functional and persisted locally.
|
||||
|
||||
Current settings include:
|
||||
|
||||
### General
|
||||
|
||||
- Launch at startup preference
|
||||
- Keep in tray preference
|
||||
- Global hotkey value
|
||||
- Theme preference
|
||||
|
||||
The startup/tray/hotkey preferences are stored now; OS registration is intentionally left for the packaging phase.
|
||||
|
||||
### Audio
|
||||
|
||||
- Microphone device ID
|
||||
- System audio/screen capture toggle
|
||||
- Microphone level monitoring
|
||||
|
||||
### Transcription
|
||||
|
||||
- Whisper model: `base`, `small`, `medium`
|
||||
- Compute device: auto/CPU/CUDA
|
||||
- Language
|
||||
- Speaker diarization toggle
|
||||
|
||||
### Overlay
|
||||
|
||||
- Enabled
|
||||
- Opacity
|
||||
- Font size
|
||||
- Visible line count
|
||||
|
||||
### Speech engine
|
||||
|
||||
- Real Whisper + diart mode
|
||||
- Mock mode
|
||||
- Local WebSocket URL
|
||||
|
||||
---
|
||||
|
||||
# Backend placeholder
|
||||
|
||||
This demo deliberately has no remote backend implementation.
|
||||
|
||||
Any backend-only action returns/displays:
|
||||
|
||||
```text
|
||||
Processing (The job). Please wait a moment
|
||||
```
|
||||
|
||||
This currently applies to example actions such as:
|
||||
|
||||
- Generate/regenerate summary
|
||||
- Create reminder
|
||||
- Action-item processing
|
||||
|
||||
The UI is structured so a real API can replace `src/lib/backend.ts` later.
|
||||
|
||||
---
|
||||
|
||||
# Current limitations
|
||||
|
||||
## System audio
|
||||
|
||||
System-audio capture depends on the operating system, WebView implementation, selected source, and whether that source permits audio capture.
|
||||
|
||||
Microphone capture is the reliable baseline for the demo.
|
||||
|
||||
## diart model setup
|
||||
|
||||
Real diart execution can require pyannote model terms/authentication. The exact dependency combinations can also be sensitive to Python, PyTorch, torchaudio, and platform versions.
|
||||
|
||||
Use the recommended Python version and validate the speech-engine environment on the target development machine before packaging.
|
||||
|
||||
## Realtime accuracy
|
||||
|
||||
Speaker diarization can temporarily swap speaker labels, especially with:
|
||||
|
||||
- overlapping speakers
|
||||
- short utterances
|
||||
- background noise
|
||||
- similar voices
|
||||
|
||||
The rename function changes the display name but does not perform biometric identity recognition.
|
||||
|
||||
## Long recordings
|
||||
|
||||
The media recorder writes the final Blob when the meeting ends in this initial demo. Before production use, change this to incremental/chunked file writing so multi-hour meetings do not accumulate the entire media Blob in memory.
|
||||
|
||||
---
|
||||
|
||||
# Recommended next development steps
|
||||
|
||||
1. Validate real Whisper + diart on the target Windows machine.
|
||||
2. Replace final-Blob recording with incremental local file writes.
|
||||
3. Add proper microphone device enumeration.
|
||||
4. Add native Windows loopback audio capture if system-audio reliability is required.
|
||||
5. Add session crash recovery.
|
||||
6. Add transcript segment editing/reassignment.
|
||||
7. Package the Python speech engine as a sidecar or replace it with native whisper.cpp + a native/ONNX diarization runtime.
|
||||
8. Implement the real backend jobs/API later.
|
||||
9. Replace local-only media with SeaweedFS when backend/storage work begins.
|
||||
|
||||
---
|
||||
|
||||
# Demo design decisions from the supplied specification
|
||||
|
||||
```text
|
||||
Login / subscription REMOVED
|
||||
Remote backend PLACEHOLDER ONLY
|
||||
Backend message "Processing (The job). Please wait a moment"
|
||||
Media storage LOCAL
|
||||
Video playback INCLUDED
|
||||
Live transcription INCLUDED
|
||||
Speaker diarization INCLUDED
|
||||
Live speaker rename INCLUDED
|
||||
Realtime overlay INCLUDED
|
||||
Settings INCLUDED
|
||||
```
|
||||
Reference in New Issue
Block a user