Files
MeetVault/Wireframes/Desktop/MeetVault_Desktop_README.md
2026-09-30 00:43:24 +07:00

1050 lines
20 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# MeetVault Desktop
MeetVault Desktop is the local desktop client for MeetVault.
Its main purpose is to:
- Record meeting audio/video
- Capture microphone and system audio
- Transcribe speech locally using Whisper
- Detect and separate speakers in real time using diart
- Show a real-time transcription overlay on top of other applications
- Allow users to rename detected speakers during the meeting
- Upload recordings to SeaweedFS
- Send the final transcript to the MeetVault backend
- Receive generated summaries, action items, and reminders from the backend
The desktop application should continue to work locally during a meeting even if the backend is temporarily unavailable.
---
# Core Desktop Architecture
```text
MeetVault Desktop
┌─────────────────┐
│ Desktop UI │
│ React + Tauri │
└────────┬────────┘
│
▼
┌─────────────────────┐
│ Audio Engine │
│ │
│ Microphone capture │
│ System audio capture│
│ Audio mixing │
└─────────┬───────────┘
│
┌───────────────┴───────────────┐
│ │
▼ ▼
┌───────────────┐ ┌────────────────┐
│ Whisper │ │ diart │
│ │ │ │
│ transcription │ │ diarization │
│ + timestamps │ │ speaker labels │
└───────┬───────┘ └───────┬────────┘
│ │
└──────────────┬───────────────┘
▼
┌────────────────────┐
│ Segment Merger │
│ │
│ text + timestamps │
│ speaker timeline │
└─────────┬──────────┘
│
┌─────────────┴──────────────┐
│ │
▼ ▼
┌─────────────────┐ ┌──────────────────┐
│ Live Transcript │ │ Realtime Overlay │
│ Window │ │ Always-on-top │
└─────────────────┘ └──────────────────┘
```
---
# Recommended Desktop Stack
```text
Desktop shell
Tauri
UI
React + TypeScript
Native/local processing
Rust where required
Transcription
Whisper / whisper.cpp
Speaker diarization
diart
Backend API
MeetVault backend
Structured data
PostgreSQL
Media storage
SeaweedFS
```
Tauri is recommended because the application needs:
- Multiple windows
- Transparent overlay windows
- Always-on-top behavior
- Native audio access
- Local file access
- Background processes
- System tray support
- Global shortcuts
- Lower memory usage than a full Electron application
---
# Real-Time Audio Pipeline
The application processes meeting audio locally.
```text
Microphone
│
├────────────────┐
│ │
System Audio │
│ │
└──────┬─────────┘
▼
Audio Mixer
│
▼
Audio Stream
│
┌─────┴─────┐
│ │
▼ ▼
Whisper diart
│ │
│ │
▼ ▼
text + speaker
timestamps segments
│ │
└─────┬─────┘
▼
Segment Merger
│
▼
Live Transcript Event
│
┌─────┴────────┐
│ │
▼ ▼
Main Window Overlay Window
```
---
# Whisper
Whisper is responsible for speech-to-text.
For the first version, a lightweight local runtime such as `whisper.cpp` is recommended.
The transcription engine should provide:
```text
text
start timestamp
end timestamp
confidence if available
partial/final state
```
Example:
```json
{
"start": 18.2,
"end": 21.7,
"text": "We should finish the API before Friday.",
"final": true
}
```
## Suggested Model Strategy
Default:
```text
Whisper small
```
Possible fallback for lower-end hardware:
```text
Whisper base
```
Possible higher-accuracy option:
```text
Whisper medium
```
The selected model should be configurable in Settings.
The goal is to balance:
```text
accuracy
latency
CPU/GPU usage
memory usage
```
---
# diart
diart is responsible for real-time speaker diarization.
It does not need to know a person's real identity.
It only needs to produce stable speaker labels such as:
```text
Speaker 1
Speaker 2
Speaker 3
```
Example diarization result:
```json
{
"speaker": "speaker_1",
"start": 18.0,
"end": 22.0
}
```
The application combines the diart result with the Whisper result using timestamps.
---
# Transcript + Speaker Merge
Whisper may produce:
```text
18.2 - 21.7
"We should finish the API before Friday."
```
diart may produce:
```text
18.0 - 22.0
speaker_1
```
The merger produces:
```json
{
"speaker_id": "speaker_1",
"speaker_name": "Speaker 1",
"start": 18.2,
"end": 21.7,
"text": "We should finish the API before Friday."
}
```
This becomes the core live transcript segment.
---
# Speaker Rename
Detected speakers initially appear as:
```text
Speaker 1
Speaker 2
Speaker 3
```
The user can rename them at any time.
Example:
```text
Speaker 1 → Alice
Speaker 2 → Bob
```
The underlying diarization ID does not change.
```text
speaker_1
↓
display name
↓
Alice
```
Recommended local mapping:
```json
{
"speaker_1": "Alice",
"speaker_2": "Bob"
}
```
Existing transcript lines and future overlay lines should immediately use the new display name.
---
# Real-Time Overlay
The overlay is a dedicated transparent Tauri window.
It should be:
- Always on top
- Transparent or semi-transparent
- Resizable
- Draggable
- Lockable
- Optionally click-through
- Hotkey controlled
- Able to show only the latest few transcript lines
Example:
```text
┌──────────────────────────────────────────────┐
│ Alice │
│ We should finish the API before Friday. │
│ │
│ Bob │
│ I'll prepare the deployment plan tonight. │
└──────────────────────────────────────────────┘
```
The overlay should update continuously while the meeting is running.
---
# Overlay Update Model
The desktop app should distinguish between:
```text
partial transcript
final transcript
```
Example:
```text
Speaker 1:
We should finish the...
```
updates to:
```text
Speaker 1:
We should finish the API before Friday.
```
A partial line may be visually distinguished from finalized text.
Recommended behavior:
```text
partial
→ replace current live line
final
→ commit transcript segment
→ append to transcript history
```
---
# Main UI Screens
## Dashboard
Purpose:
- Start a new meeting
- View recent meetings
- Check local transcription state
- Check SeaweedFS/backend connectivity
- Open Settings
Main actions:
```text
Start Meeting
Open Meeting
Preview Overlay
Settings
```
---
# Live Meeting Screen
The Live Meeting screen is the main control surface while recording.
It should show:
- Recording state
- Meeting timer
- Live transcript
- Speaker labels
- Rename controls
- Microphone state
- System audio state
- Whisper model status
- Diarization status
- Overlay toggle
- Pause
- Finish meeting
Example:
```text
Weekly Product Sync REC 00:18:42
Live transcript
Alice
We should finalize the API contract.
Bob
I'll update the database migration.
Alice
Great. I'll prepare the release notes.
Speakers
Alice Rename
Bob Rename
Microphone ON
System Audio ON
Whisper Healthy
diart Healthy
[Pause] [Show Overlay] [Finish Meeting]
```
---
# Speaker Management
The speaker panel should display the diarization speakers detected in the current meeting.
Example:
```text
speaker_1 Alice
speaker_2 Bob
speaker_3 Speaker 3
```
The user may rename a speaker at any time.
The rename should affect:
- Current live transcript
- Overlay
- Final transcript
- Meeting review
---
# Meeting Details Screen
After a meeting ends, the Meeting Details screen provides:
- Video/audio playback
- Meeting information
- Transcript
- Speaker names
- Summary
- Action items
- Reminders
- Export options
Example layout:
```text
Weekly Product Sync
Sep 27, 2026 • 48 min
┌──────────────────────────┐
│ Video Player │
└──────────────────────────┘
Transcript | Summary | Action Items | Files
Alice 00:18:41
We should finalize the API contract.
Bob 00:18:49
I'll update the database migration.
```
The transcript timestamp should be clickable so playback can jump to that position.
---
# Settings
Settings should include the following sections.
## General
```text
Launch at startup
Keep running in tray
Theme
Global hotkey
```
## Audio
```text
Microphone device
System audio device
Input level
System audio level
Audio monitoring
```
## Transcription
```text
Whisper model
Compute device
Language
Chunk size
Voice activity detection
```
Suggested model options:
```text
base
small
medium
```
## Diarization
```text
Enable speaker diarization
Maximum expected speakers
Speaker clustering sensitivity
```
## Overlay
```text
Enable overlay
Always on top
Click-through mode
Opacity
Font size
Maximum visible lines
Show speaker labels
Screen position
```
## Storage & Sync
```text
Local recording folder
Automatic upload
Upload after meeting
Delete local media after successful upload
Backend connection
SeaweedFS upload status
```
## Account
```text
Signed-in user
Subscription
Sign out
```
---
# Local Meeting Session Model
During a live meeting, the desktop application should maintain a local session object.
Example:
```json
{
"meeting_id": "uuid",
"started_at": "2026-09-27T09:00:00Z",
"status": "recording",
"speakers": {
"speaker_1": "Alice",
"speaker_2": "Bob"
},
"segments": [
{
"id": "segment-1",
"speaker_id": "speaker_1",
"start": 18.2,
"end": 21.7,
"text": "We should finish the API before Friday."
}
]
}
```
This state should be stored locally during recording so a temporary network failure does not lose the transcript.
---
# Transcript Segment Model
A transcript should internally be represented as segments rather than only one large text string.
Recommended structure:
```json
{
"id": "segment-uuid",
"speaker_id": "speaker_1",
"speaker_name": "Alice",
"start": 18.2,
"end": 21.7,
"text": "We should finish the API before Friday.",
"final": true
}
```
Benefits:
- Real-time rendering
- Speaker rename
- Timestamp navigation
- Transcript editing
- Summary generation
- Export
- Search
---
# Recording Flow
```text
User clicks Start Meeting
│
▼
Create local meeting session
│
▼
Start microphone capture
│
├──── Start system audio capture
│
▼
Start recording
│
├──── Start Whisper
│
└──── Start diart
│
▼
Live transcript events
│
┌──────┴──────┐
▼ ▼
Main UI Overlay
```
---
# Meeting Finish Flow
```text
User clicks Finish Meeting
│
▼
Stop audio capture
│
▼
Finalize Whisper segments
│
▼
Finalize speaker mapping
│
▼
Save local transcript
│
├──── upload media ─────► SeaweedFS
│
└──── send transcript ──► Backend API
│
▼
PostgreSQL
│
▼
processing_jobs
│
┌───────┴────────┐
▼ ▼
Summarizer Action Items
│
▼
Reminders
```
---
# Upload Strategy
Large media should not pass through the backend API.
Recommended:
```text
Desktop
│
│ request presigned URL
▼
Backend
│
│ returns temporary upload URL
▼
Desktop ───────────────────────► SeaweedFS
media file
```
The transcript is small and can be sent directly through the backend API.
```text
Desktop
│
│ transcript JSON
▼
Backend API
│
▼
PostgreSQL
```
---
# Offline / Network Failure Behavior
The desktop application should not require continuous backend connectivity during a meeting.
If the backend becomes unavailable:
```text
recording continues
transcription continues
diarization continues
overlay continues
local saving continues
```
After the connection returns:
```text
upload recording
send transcript
sync meeting metadata
```
This is an important design requirement for MeetVault Desktop.
---
# Background Components
The desktop application can be separated into these logical modules:
```text
MeetVault Desktop
├── UI
│ ├── Dashboard
│ ├── Live Meeting
│ ├── Meeting Details
│ └── Settings
│
├── Overlay
│ └── Realtime subtitle window
│
├── Audio Engine
│ ├── Microphone
│ ├── System audio
│ └── Mixer
│
├── Speech Engine
│ ├── Whisper
│ └── diart
│
├── Transcript Engine
│ ├── Segment merger
│ ├── Speaker mapping
│ └── Transcript state
│
├── Recorder
│ └── Media recording
│
├── Local Storage
│ └── Session recovery/cache
│
└── Sync Client
├── Backend API
└── SeaweedFS upload
```
---
# Realtime Event Flow
Internally, the UI should consume events rather than directly controlling the models.
Example events:
```text
audio.started
audio.stopped
transcript.partial
transcript.final
speaker.detected
speaker.changed
speaker.renamed
recording.started
recording.stopped
upload.started
upload.progress
upload.completed
upload.failed
```
Example event:
```json
{
"type": "transcript.final",
"payload": {
"speaker_id": "speaker_1",
"start": 18.2,
"end": 21.7,
"text": "We should finish the API before Friday."
}
}
```
Both the main transcript screen and overlay can subscribe to the same event stream.
---
# Realtime Performance Goals
Initial targets:
```text
Overlay update latency
< 2 seconds
Partial transcript refresh
~ 0.5–1.5 seconds
Final segment delay
~ 1–3 seconds
Overlay frame impact
Minimal
Meeting duration
Several hours without restart
```
Actual latency depends on:
- Whisper model
- CPU/GPU
- Audio chunk size
- Number of speakers
- diart processing speed
---
# Speaker Diarization Limitations
Speaker diarization is not perfect.
Expected difficult cases include:
- Two people speaking simultaneously
- Very similar voices
- Poor microphone quality
- Strong background noise
- Echo
- Speakers far from the microphone
- Very short speech segments
The application should allow users to correct speaker names and, later, reassign individual transcript segments if necessary.
No biometric speaker identification is required.
MeetVault only needs:
```text
who spoke when
```
not:
```text
who is this person in real life
```
---
# Security
Permanent SeaweedFS credentials must never be stored in the renderer/frontend.
The desktop app should use:
```text
Backend authentication
+
temporary presigned upload URLs
```
Sensitive local files should be stored only in the configured MeetVault application data directory.
Authentication tokens should use the operating system's secure credential storage when possible.
---
# MVP Scope
The first desktop release should focus on:
- Start/stop meeting recording
- Microphone capture
- System audio capture
- Local Whisper transcription
- Real-time diart speaker separation
- Speaker 1 / Speaker 2 labels
- Rename speakers
- Real-time overlay
- Local transcript storage
- Meeting playback
- Transcript review
- Upload recording to SeaweedFS
- Send transcript to backend
- Receive generated summary
- View action items
- View reminders
Not required for the first release:
- Speaker biometric identification
- Automatic real-name recognition
- Cloud transcription
- Kafka
- Elasticsearch
- Kubernetes
- Vector database
- Complex microservice architecture
---
# Final Desktop Architecture
```text
MeetVault Desktop
Audio Sources
/ \
Microphone System Audio
\ /
\ /
Audio Mixer
│
┌─────────┴─────────┐
│ │
▼ ▼
Whisper diart
│ │
transcript speakers
│ │
└─────────┬─────────┘
▼
Segment Merger
│
┌─────────────┴──────────────┐
│ │
▼ ▼
Live Meeting UI Realtime Overlay
│
▼
Local Session Store
│
meeting completes
│
┌───────┴────────┐
│ │
▼ ▼
SeaweedFS Backend API
media transcript
│
▼
PostgreSQL
│
▼
processing_jobs
│
┌────────┴─────────┐
▼ ▼
Summary Action Items
│
▼
Reminders
```