20 KiB
MeetVault Desktop
MeetVault Desktop is the local desktop client for MeetVault.
Its main purpose is to:
- Record meeting audio/video
- Capture microphone and system audio
- Transcribe speech locally using Whisper
- Detect and separate speakers in real time using diart
- Show a real-time transcription overlay on top of other applications
- Allow users to rename detected speakers during the meeting
- Upload recordings to SeaweedFS
- Send the final transcript to the MeetVault backend
- Receive generated summaries, action items, and reminders from the backend
The desktop application should continue to work locally during a meeting even if the backend is temporarily unavailable.
Core Desktop Architecture
MeetVault Desktop
┌─────────────────┐
│ Desktop UI │
│ React + Tauri │
└────────┬────────┘
│
▼
┌─────────────────────┐
│ Audio Engine │
│ │
│ Microphone capture │
│ System audio capture│
│ Audio mixing │
└─────────┬───────────┘
│
┌───────────────┴───────────────┐
│ │
▼ ▼
┌───────────────┐ ┌────────────────┐
│ Whisper │ │ diart │
│ │ │ │
│ transcription │ │ diarization │
│ + timestamps │ │ speaker labels │
└───────┬───────┘ └───────┬────────┘
│ │
└──────────────┬───────────────┘
▼
┌────────────────────┐
│ Segment Merger │
│ │
│ text + timestamps │
│ speaker timeline │
└─────────┬──────────┘
│
┌─────────────┴──────────────┐
│ │
▼ ▼
┌─────────────────┐ ┌──────────────────┐
│ Live Transcript │ │ Realtime Overlay │
│ Window │ │ Always-on-top │
└─────────────────┘ └──────────────────┘
Recommended Desktop Stack
Desktop shell
Tauri
UI
React + TypeScript
Native/local processing
Rust where required
Transcription
Whisper / whisper.cpp
Speaker diarization
diart
Backend API
MeetVault backend
Structured data
PostgreSQL
Media storage
SeaweedFS
Tauri is recommended because the application needs:
- Multiple windows
- Transparent overlay windows
- Always-on-top behavior
- Native audio access
- Local file access
- Background processes
- System tray support
- Global shortcuts
- Lower memory usage than a full Electron application
Real-Time Audio Pipeline
The application processes meeting audio locally.
Microphone
│
├────────────────┐
│ │
System Audio │
│ │
└──────┬─────────┘
▼
Audio Mixer
│
▼
Audio Stream
│
┌─────┴─────┐
│ │
▼ ▼
Whisper diart
│ │
│ │
▼ ▼
text + speaker
timestamps segments
│ │
└─────┬─────┘
▼
Segment Merger
│
▼
Live Transcript Event
│
┌─────┴────────┐
│ │
▼ ▼
Main Window Overlay Window
Whisper
Whisper is responsible for speech-to-text.
For the first version, a lightweight local runtime such as whisper.cpp is recommended.
The transcription engine should provide:
text
start timestamp
end timestamp
confidence if available
partial/final state
Example:
{
"start": 18.2,
"end": 21.7,
"text": "We should finish the API before Friday.",
"final": true
}
Suggested Model Strategy
Default:
Whisper small
Possible fallback for lower-end hardware:
Whisper base
Possible higher-accuracy option:
Whisper medium
The selected model should be configurable in Settings.
The goal is to balance:
accuracy
latency
CPU/GPU usage
memory usage
diart
diart is responsible for real-time speaker diarization.
It does not need to know a person's real identity.
It only needs to produce stable speaker labels such as:
Speaker 1
Speaker 2
Speaker 3
Example diarization result:
{
"speaker": "speaker_1",
"start": 18.0,
"end": 22.0
}
The application combines the diart result with the Whisper result using timestamps.
Transcript + Speaker Merge
Whisper may produce:
18.2 - 21.7
"We should finish the API before Friday."
diart may produce:
18.0 - 22.0
speaker_1
The merger produces:
{
"speaker_id": "speaker_1",
"speaker_name": "Speaker 1",
"start": 18.2,
"end": 21.7,
"text": "We should finish the API before Friday."
}
This becomes the core live transcript segment.
Speaker Rename
Detected speakers initially appear as:
Speaker 1
Speaker 2
Speaker 3
The user can rename them at any time.
Example:
Speaker 1 → Alice
Speaker 2 → Bob
The underlying diarization ID does not change.
speaker_1
↓
display name
↓
Alice
Recommended local mapping:
{
"speaker_1": "Alice",
"speaker_2": "Bob"
}
Existing transcript lines and future overlay lines should immediately use the new display name.
Real-Time Overlay
The overlay is a dedicated transparent Tauri window.
It should be:
- Always on top
- Transparent or semi-transparent
- Resizable
- Draggable
- Lockable
- Optionally click-through
- Hotkey controlled
- Able to show only the latest few transcript lines
Example:
┌──────────────────────────────────────────────┐
│ Alice │
│ We should finish the API before Friday. │
│ │
│ Bob │
│ I'll prepare the deployment plan tonight. │
└──────────────────────────────────────────────┘
The overlay should update continuously while the meeting is running.
Overlay Update Model
The desktop app should distinguish between:
partial transcript
final transcript
Example:
Speaker 1:
We should finish the...
updates to:
Speaker 1:
We should finish the API before Friday.
A partial line may be visually distinguished from finalized text.
Recommended behavior:
partial
→ replace current live line
final
→ commit transcript segment
→ append to transcript history
Main UI Screens
Dashboard
Purpose:
- Start a new meeting
- View recent meetings
- Check local transcription state
- Check SeaweedFS/backend connectivity
- Open Settings
Main actions:
Start Meeting
Open Meeting
Preview Overlay
Settings
Live Meeting Screen
The Live Meeting screen is the main control surface while recording.
It should show:
- Recording state
- Meeting timer
- Live transcript
- Speaker labels
- Rename controls
- Microphone state
- System audio state
- Whisper model status
- Diarization status
- Overlay toggle
- Pause
- Finish meeting
Example:
Weekly Product Sync REC 00:18:42
Live transcript
Alice
We should finalize the API contract.
Bob
I'll update the database migration.
Alice
Great. I'll prepare the release notes.
Speakers
Alice Rename
Bob Rename
Microphone ON
System Audio ON
Whisper Healthy
diart Healthy
[Pause] [Show Overlay] [Finish Meeting]
Speaker Management
The speaker panel should display the diarization speakers detected in the current meeting.
Example:
speaker_1 Alice
speaker_2 Bob
speaker_3 Speaker 3
The user may rename a speaker at any time.
The rename should affect:
- Current live transcript
- Overlay
- Final transcript
- Meeting review
Meeting Details Screen
After a meeting ends, the Meeting Details screen provides:
- Video/audio playback
- Meeting information
- Transcript
- Speaker names
- Summary
- Action items
- Reminders
- Export options
Example layout:
Weekly Product Sync
Sep 27, 2026 • 48 min
┌──────────────────────────┐
│ Video Player │
└──────────────────────────┘
Transcript | Summary | Action Items | Files
Alice 00:18:41
We should finalize the API contract.
Bob 00:18:49
I'll update the database migration.
The transcript timestamp should be clickable so playback can jump to that position.
Settings
Settings should include the following sections.
General
Launch at startup
Keep running in tray
Theme
Global hotkey
Audio
Microphone device
System audio device
Input level
System audio level
Audio monitoring
Transcription
Whisper model
Compute device
Language
Chunk size
Voice activity detection
Suggested model options:
base
small
medium
Diarization
Enable speaker diarization
Maximum expected speakers
Speaker clustering sensitivity
Overlay
Enable overlay
Always on top
Click-through mode
Opacity
Font size
Maximum visible lines
Show speaker labels
Screen position
Storage & Sync
Local recording folder
Automatic upload
Upload after meeting
Delete local media after successful upload
Backend connection
SeaweedFS upload status
Account
Signed-in user
Subscription
Sign out
Local Meeting Session Model
During a live meeting, the desktop application should maintain a local session object.
Example:
{
"meeting_id": "uuid",
"started_at": "2026-09-27T09:00:00Z",
"status": "recording",
"speakers": {
"speaker_1": "Alice",
"speaker_2": "Bob"
},
"segments": [
{
"id": "segment-1",
"speaker_id": "speaker_1",
"start": 18.2,
"end": 21.7,
"text": "We should finish the API before Friday."
}
]
}
This state should be stored locally during recording so a temporary network failure does not lose the transcript.
Transcript Segment Model
A transcript should internally be represented as segments rather than only one large text string.
Recommended structure:
{
"id": "segment-uuid",
"speaker_id": "speaker_1",
"speaker_name": "Alice",
"start": 18.2,
"end": 21.7,
"text": "We should finish the API before Friday.",
"final": true
}
Benefits:
- Real-time rendering
- Speaker rename
- Timestamp navigation
- Transcript editing
- Summary generation
- Export
- Search
Recording Flow
User clicks Start Meeting
│
▼
Create local meeting session
│
▼
Start microphone capture
│
├──── Start system audio capture
│
▼
Start recording
│
├──── Start Whisper
│
└──── Start diart
│
▼
Live transcript events
│
┌──────┴──────┐
▼ ▼
Main UI Overlay
Meeting Finish Flow
User clicks Finish Meeting
│
▼
Stop audio capture
│
▼
Finalize Whisper segments
│
▼
Finalize speaker mapping
│
▼
Save local transcript
│
├──── upload media ─────► SeaweedFS
│
└──── send transcript ──► Backend API
│
▼
PostgreSQL
│
▼
processing_jobs
│
┌───────┴────────┐
▼ ▼
Summarizer Action Items
│
▼
Reminders
Upload Strategy
Large media should not pass through the backend API.
Recommended:
Desktop
│
│ request presigned URL
▼
Backend
│
│ returns temporary upload URL
▼
Desktop ───────────────────────► SeaweedFS
media file
The transcript is small and can be sent directly through the backend API.
Desktop
│
│ transcript JSON
▼
Backend API
│
▼
PostgreSQL
Offline / Network Failure Behavior
The desktop application should not require continuous backend connectivity during a meeting.
If the backend becomes unavailable:
recording continues
transcription continues
diarization continues
overlay continues
local saving continues
After the connection returns:
upload recording
send transcript
sync meeting metadata
This is an important design requirement for MeetVault Desktop.
Background Components
The desktop application can be separated into these logical modules:
MeetVault Desktop
├── UI
│ ├── Dashboard
│ ├── Live Meeting
│ ├── Meeting Details
│ └── Settings
│
├── Overlay
│ └── Realtime subtitle window
│
├── Audio Engine
│ ├── Microphone
│ ├── System audio
│ └── Mixer
│
├── Speech Engine
│ ├── Whisper
│ └── diart
│
├── Transcript Engine
│ ├── Segment merger
│ ├── Speaker mapping
│ └── Transcript state
│
├── Recorder
│ └── Media recording
│
├── Local Storage
│ └── Session recovery/cache
│
└── Sync Client
├── Backend API
└── SeaweedFS upload
Realtime Event Flow
Internally, the UI should consume events rather than directly controlling the models.
Example events:
audio.started
audio.stopped
transcript.partial
transcript.final
speaker.detected
speaker.changed
speaker.renamed
recording.started
recording.stopped
upload.started
upload.progress
upload.completed
upload.failed
Example event:
{
"type": "transcript.final",
"payload": {
"speaker_id": "speaker_1",
"start": 18.2,
"end": 21.7,
"text": "We should finish the API before Friday."
}
}
Both the main transcript screen and overlay can subscribe to the same event stream.
Realtime Performance Goals
Initial targets:
Overlay update latency
< 2 seconds
Partial transcript refresh
~ 0.5–1.5 seconds
Final segment delay
~ 1–3 seconds
Overlay frame impact
Minimal
Meeting duration
Several hours without restart
Actual latency depends on:
- Whisper model
- CPU/GPU
- Audio chunk size
- Number of speakers
- diart processing speed
Speaker Diarization Limitations
Speaker diarization is not perfect.
Expected difficult cases include:
- Two people speaking simultaneously
- Very similar voices
- Poor microphone quality
- Strong background noise
- Echo
- Speakers far from the microphone
- Very short speech segments
The application should allow users to correct speaker names and, later, reassign individual transcript segments if necessary.
No biometric speaker identification is required.
MeetVault only needs:
who spoke when
not:
who is this person in real life
Security
Permanent SeaweedFS credentials must never be stored in the renderer/frontend.
The desktop app should use:
Backend authentication
+
temporary presigned upload URLs
Sensitive local files should be stored only in the configured MeetVault application data directory.
Authentication tokens should use the operating system's secure credential storage when possible.
MVP Scope
The first desktop release should focus on:
- Start/stop meeting recording
- Microphone capture
- System audio capture
- Local Whisper transcription
- Real-time diart speaker separation
- Speaker 1 / Speaker 2 labels
- Rename speakers
- Real-time overlay
- Local transcript storage
- Meeting playback
- Transcript review
- Upload recording to SeaweedFS
- Send transcript to backend
- Receive generated summary
- View action items
- View reminders
Not required for the first release:
- Speaker biometric identification
- Automatic real-name recognition
- Cloud transcription
- Kafka
- Elasticsearch
- Kubernetes
- Vector database
- Complex microservice architecture
Final Desktop Architecture
MeetVault Desktop
Audio Sources
/ \
Microphone System Audio
\ /
\ /
Audio Mixer
│
┌─────────┴─────────┐
│ │
▼ ▼
Whisper diart
│ │
transcript speakers
│ │
└─────────┬─────────┘
▼
Segment Merger
│
┌─────────────┴──────────────┐
│ │
▼ ▼
Live Meeting UI Realtime Overlay
│
▼
Local Session Store
│
meeting completes
│
┌───────┴────────┐
│ │
▼ ▼
SeaweedFS Backend API
media transcript
│
▼
PostgreSQL
│
▼
processing_jobs
│
┌────────┴─────────┐
▼ ▼
Summary Action Items
│
▼
Reminders