Files
MeetVault/Wireframes/Desktop/MeetVault-Desktop-Demo/docs/MeetVault_Desktop_README_original.md
2026-09-30 00:50:44 +07:00

20 KiB
Raw Blame History

MeetVault Desktop

MeetVault Desktop is the local desktop client for MeetVault.

Its main purpose is to:

  • Record meeting audio/video
  • Capture microphone and system audio
  • Transcribe speech locally using Whisper
  • Detect and separate speakers in real time using diart
  • Show a real-time transcription overlay on top of other applications
  • Allow users to rename detected speakers during the meeting
  • Upload recordings to SeaweedFS
  • Send the final transcript to the MeetVault backend
  • Receive generated summaries, action items, and reminders from the backend

The desktop application should continue to work locally during a meeting even if the backend is temporarily unavailable.


Core Desktop Architecture

                         MeetVault Desktop

                        ┌─────────────────┐
                        │   Desktop UI    │
                        │ React + Tauri   │
                        └────────┬────────┘
                                 │
                                 ▼
                      ┌─────────────────────┐
                      │    Audio Engine     │
                      │                     │
                      │ Microphone capture  │
                      │ System audio capture│
                      │ Audio mixing        │
                      └─────────┬───────────┘
                                │
                ┌───────────────┴───────────────┐
                │                               │
                ▼                               ▼
        ┌───────────────┐              ┌────────────────┐
        │    Whisper    │              │     diart      │
        │               │              │                │
        │ transcription │              │ diarization    │
        │ + timestamps  │              │ speaker labels │
        └───────┬───────┘              └───────┬────────┘
                │                              │
                └──────────────┬───────────────┘
                               ▼
                    ┌────────────────────┐
                    │ Segment Merger     │
                    │                    │
                    │ text + timestamps  │
                    │ speaker timeline   │
                    └─────────┬──────────┘
                              │
                ┌─────────────┴──────────────┐
                │                            │
                ▼                            ▼
      ┌─────────────────┐          ┌──────────────────┐
      │ Live Transcript │          │ Realtime Overlay │
      │ Window          │          │ Always-on-top    │
      └─────────────────┘          └──────────────────┘

Recommended Desktop Stack

Desktop shell
Tauri

UI
React + TypeScript

Native/local processing
Rust where required

Transcription
Whisper / whisper.cpp

Speaker diarization
diart

Backend API
MeetVault backend

Structured data
PostgreSQL

Media storage
SeaweedFS

Tauri is recommended because the application needs:

  • Multiple windows
  • Transparent overlay windows
  • Always-on-top behavior
  • Native audio access
  • Local file access
  • Background processes
  • System tray support
  • Global shortcuts
  • Lower memory usage than a full Electron application

Real-Time Audio Pipeline

The application processes meeting audio locally.

Microphone
     │
     ├────────────────┐
     │                │
System Audio          │
     │                │
     └──────┬─────────┘
            ▼
       Audio Mixer
            │
            ▼
      Audio Stream
            │
      ┌─────┴─────┐
      │           │
      ▼           ▼
   Whisper      diart
      │           │
      │           │
      ▼           ▼
   text +       speaker
 timestamps     segments
      │           │
      └─────┬─────┘
            ▼
      Segment Merger
            │
            ▼
    Live Transcript Event
            │
      ┌─────┴────────┐
      │              │
      ▼              ▼
 Main Window     Overlay Window

Whisper

Whisper is responsible for speech-to-text.

For the first version, a lightweight local runtime such as whisper.cpp is recommended.

The transcription engine should provide:

text
start timestamp
end timestamp
confidence if available
partial/final state

Example:

{
  "start": 18.2,
  "end": 21.7,
  "text": "We should finish the API before Friday.",
  "final": true
}

Suggested Model Strategy

Default:

Whisper small

Possible fallback for lower-end hardware:

Whisper base

Possible higher-accuracy option:

Whisper medium

The selected model should be configurable in Settings.

The goal is to balance:

accuracy
latency
CPU/GPU usage
memory usage

diart

diart is responsible for real-time speaker diarization.

It does not need to know a person's real identity.

It only needs to produce stable speaker labels such as:

Speaker 1
Speaker 2
Speaker 3

Example diarization result:

{
  "speaker": "speaker_1",
  "start": 18.0,
  "end": 22.0
}

The application combines the diart result with the Whisper result using timestamps.


Transcript + Speaker Merge

Whisper may produce:

18.2 - 21.7

"We should finish the API before Friday."

diart may produce:

18.0 - 22.0

speaker_1

The merger produces:

{
  "speaker_id": "speaker_1",
  "speaker_name": "Speaker 1",
  "start": 18.2,
  "end": 21.7,
  "text": "We should finish the API before Friday."
}

This becomes the core live transcript segment.


Speaker Rename

Detected speakers initially appear as:

Speaker 1
Speaker 2
Speaker 3

The user can rename them at any time.

Example:

Speaker 1 → Alice
Speaker 2 → Bob

The underlying diarization ID does not change.

speaker_1
    ↓
display name
    ↓
Alice

Recommended local mapping:

{
  "speaker_1": "Alice",
  "speaker_2": "Bob"
}

Existing transcript lines and future overlay lines should immediately use the new display name.


Real-Time Overlay

The overlay is a dedicated transparent Tauri window.

It should be:

  • Always on top
  • Transparent or semi-transparent
  • Resizable
  • Draggable
  • Lockable
  • Optionally click-through
  • Hotkey controlled
  • Able to show only the latest few transcript lines

Example:

┌──────────────────────────────────────────────┐
│ Alice                                        │
│ We should finish the API before Friday.      │
│                                              │
│ Bob                                          │
│ I'll prepare the deployment plan tonight.   │
└──────────────────────────────────────────────┘

The overlay should update continuously while the meeting is running.


Overlay Update Model

The desktop app should distinguish between:

partial transcript
final transcript

Example:

Speaker 1:
We should finish the...

updates to:

Speaker 1:
We should finish the API before Friday.

A partial line may be visually distinguished from finalized text.

Recommended behavior:

partial
→ replace current live line

final
→ commit transcript segment
→ append to transcript history

Main UI Screens

Dashboard

Purpose:

  • Start a new meeting
  • View recent meetings
  • Check local transcription state
  • Check SeaweedFS/backend connectivity
  • Open Settings

Main actions:

Start Meeting
Open Meeting
Preview Overlay
Settings

Live Meeting Screen

The Live Meeting screen is the main control surface while recording.

It should show:

  • Recording state
  • Meeting timer
  • Live transcript
  • Speaker labels
  • Rename controls
  • Microphone state
  • System audio state
  • Whisper model status
  • Diarization status
  • Overlay toggle
  • Pause
  • Finish meeting

Example:

Weekly Product Sync                     REC 00:18:42

Live transcript

Alice
We should finalize the API contract.

Bob
I'll update the database migration.

Alice
Great. I'll prepare the release notes.


Speakers

Alice       Rename
Bob         Rename


Microphone       ON
System Audio     ON
Whisper          Healthy
diart            Healthy

[Pause] [Show Overlay] [Finish Meeting]

Speaker Management

The speaker panel should display the diarization speakers detected in the current meeting.

Example:

speaker_1    Alice
speaker_2    Bob
speaker_3    Speaker 3

The user may rename a speaker at any time.

The rename should affect:

  • Current live transcript
  • Overlay
  • Final transcript
  • Meeting review

Meeting Details Screen

After a meeting ends, the Meeting Details screen provides:

  • Video/audio playback
  • Meeting information
  • Transcript
  • Speaker names
  • Summary
  • Action items
  • Reminders
  • Export options

Example layout:

Weekly Product Sync
Sep 27, 2026 • 48 min

┌──────────────────────────┐
│      Video Player        │
└──────────────────────────┘

Transcript | Summary | Action Items | Files

Alice  00:18:41
We should finalize the API contract.

Bob    00:18:49
I'll update the database migration.

The transcript timestamp should be clickable so playback can jump to that position.


Settings

Settings should include the following sections.

General

Launch at startup
Keep running in tray
Theme
Global hotkey

Audio

Microphone device
System audio device
Input level
System audio level
Audio monitoring

Transcription

Whisper model
Compute device
Language
Chunk size
Voice activity detection

Suggested model options:

base
small
medium

Diarization

Enable speaker diarization
Maximum expected speakers
Speaker clustering sensitivity

Overlay

Enable overlay
Always on top
Click-through mode
Opacity
Font size
Maximum visible lines
Show speaker labels
Screen position

Storage & Sync

Local recording folder
Automatic upload
Upload after meeting
Delete local media after successful upload
Backend connection
SeaweedFS upload status

Account

Signed-in user
Subscription
Sign out

Local Meeting Session Model

During a live meeting, the desktop application should maintain a local session object.

Example:

{
  "meeting_id": "uuid",
  "started_at": "2026-09-27T09:00:00Z",
  "status": "recording",

  "speakers": {
    "speaker_1": "Alice",
    "speaker_2": "Bob"
  },

  "segments": [
    {
      "id": "segment-1",
      "speaker_id": "speaker_1",
      "start": 18.2,
      "end": 21.7,
      "text": "We should finish the API before Friday."
    }
  ]
}

This state should be stored locally during recording so a temporary network failure does not lose the transcript.


Transcript Segment Model

A transcript should internally be represented as segments rather than only one large text string.

Recommended structure:

{
  "id": "segment-uuid",
  "speaker_id": "speaker_1",
  "speaker_name": "Alice",
  "start": 18.2,
  "end": 21.7,
  "text": "We should finish the API before Friday.",
  "final": true
}

Benefits:

  • Real-time rendering
  • Speaker rename
  • Timestamp navigation
  • Transcript editing
  • Summary generation
  • Export
  • Search

Recording Flow

User clicks Start Meeting
        │
        ▼
Create local meeting session
        │
        ▼
Start microphone capture
        │
        ├──── Start system audio capture
        │
        ▼
Start recording
        │
        ├──── Start Whisper
        │
        └──── Start diart
                 │
                 ▼
        Live transcript events
                 │
          ┌──────┴──────┐
          ▼             ▼
     Main UI        Overlay

Meeting Finish Flow

User clicks Finish Meeting
        │
        ▼
Stop audio capture
        │
        ▼
Finalize Whisper segments
        │
        ▼
Finalize speaker mapping
        │
        ▼
Save local transcript
        │
        ├──── upload media ─────► SeaweedFS
        │
        └──── send transcript ──► Backend API
                                      │
                                      ▼
                               PostgreSQL
                                      │
                                      ▼
                               processing_jobs
                                      │
                              ┌───────┴────────┐
                              ▼                ▼
                         Summarizer       Action Items
                              │
                              ▼
                         Reminders

Upload Strategy

Large media should not pass through the backend API.

Recommended:

Desktop
    │
    │ request presigned URL
    ▼
Backend
    │
    │ returns temporary upload URL
    ▼
Desktop ───────────────────────► SeaweedFS
             media file

The transcript is small and can be sent directly through the backend API.

Desktop
    │
    │ transcript JSON
    ▼
Backend API
    │
    ▼
PostgreSQL

Offline / Network Failure Behavior

The desktop application should not require continuous backend connectivity during a meeting.

If the backend becomes unavailable:

recording       continues
transcription   continues
diarization     continues
overlay         continues
local saving    continues

After the connection returns:

upload recording
send transcript
sync meeting metadata

This is an important design requirement for MeetVault Desktop.


Background Components

The desktop application can be separated into these logical modules:

MeetVault Desktop

├── UI
│   ├── Dashboard
│   ├── Live Meeting
│   ├── Meeting Details
│   └── Settings
│
├── Overlay
│   └── Realtime subtitle window
│
├── Audio Engine
│   ├── Microphone
│   ├── System audio
│   └── Mixer
│
├── Speech Engine
│   ├── Whisper
│   └── diart
│
├── Transcript Engine
│   ├── Segment merger
│   ├── Speaker mapping
│   └── Transcript state
│
├── Recorder
│   └── Media recording
│
├── Local Storage
│   └── Session recovery/cache
│
└── Sync Client
    ├── Backend API
    └── SeaweedFS upload

Realtime Event Flow

Internally, the UI should consume events rather than directly controlling the models.

Example events:

audio.started
audio.stopped

transcript.partial
transcript.final

speaker.detected
speaker.changed

speaker.renamed

recording.started
recording.stopped

upload.started
upload.progress
upload.completed
upload.failed

Example event:

{
  "type": "transcript.final",
  "payload": {
    "speaker_id": "speaker_1",
    "start": 18.2,
    "end": 21.7,
    "text": "We should finish the API before Friday."
  }
}

Both the main transcript screen and overlay can subscribe to the same event stream.


Realtime Performance Goals

Initial targets:

Overlay update latency
< 2 seconds

Partial transcript refresh
~ 0.5–1.5 seconds

Final segment delay
~ 1–3 seconds

Overlay frame impact
Minimal

Meeting duration
Several hours without restart

Actual latency depends on:

  • Whisper model
  • CPU/GPU
  • Audio chunk size
  • Number of speakers
  • diart processing speed

Speaker Diarization Limitations

Speaker diarization is not perfect.

Expected difficult cases include:

  • Two people speaking simultaneously
  • Very similar voices
  • Poor microphone quality
  • Strong background noise
  • Echo
  • Speakers far from the microphone
  • Very short speech segments

The application should allow users to correct speaker names and, later, reassign individual transcript segments if necessary.

No biometric speaker identification is required.

MeetVault only needs:

who spoke when

not:

who is this person in real life

Security

Permanent SeaweedFS credentials must never be stored in the renderer/frontend.

The desktop app should use:

Backend authentication
+
temporary presigned upload URLs

Sensitive local files should be stored only in the configured MeetVault application data directory.

Authentication tokens should use the operating system's secure credential storage when possible.


MVP Scope

The first desktop release should focus on:

  • Start/stop meeting recording
  • Microphone capture
  • System audio capture
  • Local Whisper transcription
  • Real-time diart speaker separation
  • Speaker 1 / Speaker 2 labels
  • Rename speakers
  • Real-time overlay
  • Local transcript storage
  • Meeting playback
  • Transcript review
  • Upload recording to SeaweedFS
  • Send transcript to backend
  • Receive generated summary
  • View action items
  • View reminders

Not required for the first release:

  • Speaker biometric identification
  • Automatic real-name recognition
  • Cloud transcription
  • Kafka
  • Elasticsearch
  • Kubernetes
  • Vector database
  • Complex microservice architecture

Final Desktop Architecture

                       MeetVault Desktop

                         Audio Sources
                      /                \
                Microphone          System Audio
                      \                /
                       \              /
                          Audio Mixer
                              │
                    ┌─────────┴─────────┐
                    │                   │
                    ▼                   ▼
                 Whisper              diart
                    │                   │
                 transcript          speakers
                    │                   │
                    └─────────┬─────────┘
                              ▼
                       Segment Merger
                              │
                ┌─────────────┴──────────────┐
                │                            │
                ▼                            ▼
          Live Meeting UI             Realtime Overlay
                │
                ▼
          Local Session Store
                │
         meeting completes
                │
        ┌───────┴────────┐
        │                │
        ▼                ▼
     SeaweedFS       Backend API
      media           transcript
                         │
                         ▼
                     PostgreSQL
                         │
                         ▼
                  processing_jobs
                         │
                ┌────────┴─────────┐
                ▼                  ▼
             Summary           Action Items
                                   │
                                   ▼
                               Reminders