829 lines
17 KiB
Markdown
829 lines
17 KiB
Markdown
# MeetVault
|
|
|
|
MeetVault is a self-hosted meeting recording and intelligence platform.
|
|
|
|
The user machine handles meeting recording and local transcription. The backend handles authentication, meeting metadata, transcript ingestion, summarization jobs, task extraction, reminder scheduling, and persistent storage.
|
|
|
|
Large media files are stored in SeaweedFS through its S3-compatible API, while PostgreSQL stores structured application data.
|
|
|
|
---
|
|
|
|
## Architecture Overview
|
|
|
|
### User Machine
|
|
|
|
The browser or desktop application is responsible for:
|
|
|
|
- Recording meeting audio/video
|
|
- Running local speech-to-text using Whisper, whisper.cpp, or another local STT engine
|
|
- Sending the generated transcript to the backend
|
|
- Uploading large media files directly to SeaweedFS using presigned URLs
|
|
|
|
Local transcription keeps transcription compute and cost on the user's device rather than on the backend.
|
|
|
|
### Backend Infrastructure
|
|
|
|
The backend API is responsible for:
|
|
|
|
- Authentication
|
|
- User management
|
|
- Meeting management
|
|
- Presigned upload URL generation
|
|
- Transcript ingestion
|
|
- Summarization job creation
|
|
- Action-item extraction
|
|
- Reminder scheduling
|
|
- Job status tracking
|
|
|
|
PostgreSQL stores structured application data such as users, meetings, transcripts, summaries, tasks, schedules, and processing jobs.
|
|
|
|
SeaweedFS stores the actual large binary files such as:
|
|
|
|
- Video
|
|
- Audio
|
|
- Thumbnails
|
|
- Attachments
|
|
- Exported files
|
|
|
|
---
|
|
|
|
## Data Flow
|
|
|
|
```text
|
|
User machine
|
|
│
|
|
├── video ─────────────────────────────► SeaweedFS
|
|
│
|
|
└── local transcript ────────────────► PostgreSQL
|
|
│
|
|
▼
|
|
processing_jobs
|
|
│
|
|
Summarizer
|
|
│
|
|
┌─────────────┴─────────────┐
|
|
▼ ▼
|
|
meeting_tasks meeting_summaries
|
|
│
|
|
▼
|
|
meeting_schedules
|
|
│
|
|
▼
|
|
Reminder Worker
|
|
```
|
|
|
|
---
|
|
|
|
## Database Design
|
|
|
|
```text
|
|
users
|
|
│
|
|
├──── auth_identities
|
|
│
|
|
├──── subscriptions
|
|
│
|
|
├──── meetings
|
|
│ │
|
|
│ ├──── meeting_files ─────────► SeaweedFS
|
|
│ │
|
|
│ ├──── transcripts
|
|
│ │
|
|
│ ├──── meeting_summaries
|
|
│ │
|
|
│ ├──── meeting_tasks
|
|
│ │ │
|
|
│ │ └──── meeting_schedules
|
|
│ │
|
|
│ └──── processing_jobs
|
|
│
|
|
└──── meeting_schedules
|
|
```
|
|
|
|
---
|
|
|
|
# PostgreSQL Schema
|
|
|
|
## Enable UUID Generation
|
|
|
|
```sql
|
|
CREATE EXTENSION IF NOT EXISTS pgcrypto;
|
|
```
|
|
|
|
---
|
|
|
|
## Users
|
|
|
|
```sql
|
|
CREATE TABLE users (
|
|
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
|
|
|
|
email VARCHAR(320) UNIQUE NOT NULL,
|
|
display_name VARCHAR(255),
|
|
|
|
identification_status VARCHAR(30)
|
|
NOT NULL DEFAULT 'unidentified'
|
|
CHECK (
|
|
identification_status IN (
|
|
'unidentified',
|
|
'pending',
|
|
'identified',
|
|
'rejected'
|
|
)
|
|
),
|
|
|
|
account_status VARCHAR(30)
|
|
NOT NULL DEFAULT 'active'
|
|
CHECK (
|
|
account_status IN (
|
|
'active',
|
|
'suspended',
|
|
'disabled'
|
|
)
|
|
),
|
|
|
|
created_at TIMESTAMPTZ NOT NULL DEFAULT NOW(),
|
|
updated_at TIMESTAMPTZ NOT NULL DEFAULT NOW()
|
|
);
|
|
```
|
|
|
|
---
|
|
|
|
## Authentication Identities
|
|
|
|
Supports local login and external providers such as Google or Microsoft.
|
|
|
|
```sql
|
|
CREATE TABLE auth_identities (
|
|
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
|
|
|
|
user_id UUID NOT NULL
|
|
REFERENCES users(id)
|
|
ON DELETE CASCADE,
|
|
|
|
provider VARCHAR(50) NOT NULL,
|
|
|
|
provider_user_id VARCHAR(255),
|
|
password_hash TEXT,
|
|
|
|
created_at TIMESTAMPTZ NOT NULL DEFAULT NOW(),
|
|
|
|
UNIQUE(provider, provider_user_id)
|
|
);
|
|
```
|
|
|
|
Typical providers:
|
|
|
|
```text
|
|
local
|
|
google
|
|
microsoft
|
|
```
|
|
|
|
---
|
|
|
|
## Subscriptions
|
|
|
|
```sql
|
|
CREATE TABLE subscriptions (
|
|
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
|
|
|
|
user_id UUID NOT NULL
|
|
REFERENCES users(id)
|
|
ON DELETE CASCADE,
|
|
|
|
plan VARCHAR(50)
|
|
NOT NULL DEFAULT 'free',
|
|
|
|
status VARCHAR(30)
|
|
NOT NULL DEFAULT 'active'
|
|
CHECK (
|
|
status IN (
|
|
'active',
|
|
'trialing',
|
|
'past_due',
|
|
'cancelled',
|
|
'expired'
|
|
)
|
|
),
|
|
|
|
provider VARCHAR(50),
|
|
|
|
provider_customer_id VARCHAR(255),
|
|
provider_subscription_id VARCHAR(255),
|
|
|
|
current_period_start TIMESTAMPTZ,
|
|
current_period_end TIMESTAMPTZ,
|
|
|
|
cancel_at_period_end BOOLEAN
|
|
NOT NULL DEFAULT FALSE,
|
|
|
|
created_at TIMESTAMPTZ NOT NULL DEFAULT NOW(),
|
|
updated_at TIMESTAMPTZ NOT NULL DEFAULT NOW()
|
|
);
|
|
```
|
|
|
|
---
|
|
|
|
## Meetings
|
|
|
|
```sql
|
|
CREATE TABLE meetings (
|
|
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
|
|
|
|
owner_user_id UUID NOT NULL
|
|
REFERENCES users(id)
|
|
ON DELETE CASCADE,
|
|
|
|
title VARCHAR(500),
|
|
description TEXT,
|
|
|
|
started_at TIMESTAMPTZ,
|
|
ended_at TIMESTAMPTZ,
|
|
|
|
timezone VARCHAR(100),
|
|
|
|
status VARCHAR(30)
|
|
NOT NULL DEFAULT 'created'
|
|
CHECK (
|
|
status IN (
|
|
'created',
|
|
'recording',
|
|
'uploading',
|
|
'transcribing',
|
|
'processing',
|
|
'ready',
|
|
'failed',
|
|
'archived',
|
|
'deleted'
|
|
)
|
|
),
|
|
|
|
metadata JSONB NOT NULL DEFAULT '{}',
|
|
|
|
deleted_at TIMESTAMPTZ,
|
|
|
|
created_at TIMESTAMPTZ NOT NULL DEFAULT NOW(),
|
|
updated_at TIMESTAMPTZ NOT NULL DEFAULT NOW(),
|
|
|
|
CHECK (
|
|
ended_at IS NULL
|
|
OR started_at IS NULL
|
|
OR ended_at >= started_at
|
|
)
|
|
);
|
|
```
|
|
|
|
Summaries are stored separately in `meeting_summaries`.
|
|
|
|
---
|
|
|
|
## Meeting Files
|
|
|
|
This table stores file metadata only. The actual media files are stored in SeaweedFS.
|
|
|
|
```sql
|
|
CREATE TABLE meeting_files (
|
|
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
|
|
|
|
meeting_id UUID NOT NULL
|
|
REFERENCES meetings(id)
|
|
ON DELETE CASCADE,
|
|
|
|
uploaded_by UUID
|
|
REFERENCES users(id)
|
|
ON DELETE SET NULL,
|
|
|
|
file_type VARCHAR(50)
|
|
NOT NULL
|
|
CHECK (
|
|
file_type IN (
|
|
'video',
|
|
'audio',
|
|
'thumbnail',
|
|
'attachment',
|
|
'export'
|
|
)
|
|
),
|
|
|
|
original_filename VARCHAR(500),
|
|
mime_type VARCHAR(150),
|
|
|
|
storage_provider VARCHAR(30)
|
|
NOT NULL DEFAULT 'seaweedfs',
|
|
|
|
storage_bucket VARCHAR(255) NOT NULL,
|
|
storage_key TEXT NOT NULL,
|
|
|
|
size_bytes BIGINT
|
|
CHECK (
|
|
size_bytes IS NULL
|
|
OR size_bytes >= 0
|
|
),
|
|
|
|
checksum VARCHAR(255),
|
|
|
|
upload_status VARCHAR(30)
|
|
NOT NULL DEFAULT 'pending'
|
|
CHECK (
|
|
upload_status IN (
|
|
'pending',
|
|
'uploading',
|
|
'uploaded',
|
|
'failed',
|
|
'deleted'
|
|
)
|
|
),
|
|
|
|
created_at TIMESTAMPTZ NOT NULL DEFAULT NOW(),
|
|
|
|
UNIQUE(storage_provider, storage_bucket, storage_key)
|
|
);
|
|
```
|
|
|
|
Example SeaweedFS object path:
|
|
|
|
```text
|
|
meeting-storage/
|
|
users/{user_id}/
|
|
meetings/{meeting_id}/
|
|
original/
|
|
recording.webm
|
|
```
|
|
|
|
---
|
|
|
|
## Transcripts
|
|
|
|
Transcription happens locally on the user's machine and is uploaded to the backend afterward.
|
|
|
|
```sql
|
|
CREATE TABLE transcripts (
|
|
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
|
|
|
|
meeting_id UUID NOT NULL
|
|
REFERENCES meetings(id)
|
|
ON DELETE CASCADE,
|
|
|
|
source VARCHAR(30)
|
|
NOT NULL DEFAULT 'local'
|
|
CHECK (
|
|
source IN (
|
|
'local',
|
|
'imported'
|
|
)
|
|
),
|
|
|
|
engine VARCHAR(100),
|
|
language VARCHAR(20),
|
|
|
|
version INTEGER
|
|
NOT NULL DEFAULT 1
|
|
CHECK (version > 0),
|
|
|
|
is_active BOOLEAN
|
|
NOT NULL DEFAULT TRUE,
|
|
|
|
plain_text TEXT,
|
|
transcript_json JSONB,
|
|
|
|
duration_seconds NUMERIC(12,2)
|
|
CHECK (
|
|
duration_seconds IS NULL
|
|
OR duration_seconds >= 0
|
|
),
|
|
|
|
status VARCHAR(30)
|
|
NOT NULL DEFAULT 'uploaded'
|
|
CHECK (
|
|
status IN (
|
|
'uploaded',
|
|
'processing',
|
|
'ready',
|
|
'failed'
|
|
)
|
|
),
|
|
|
|
created_at TIMESTAMPTZ NOT NULL DEFAULT NOW(),
|
|
updated_at TIMESTAMPTZ NOT NULL DEFAULT NOW(),
|
|
|
|
UNIQUE(meeting_id, version)
|
|
);
|
|
```
|
|
|
|
Example transcript JSON:
|
|
|
|
```json
|
|
{
|
|
"segments": [
|
|
{
|
|
"speaker": "Speaker 1",
|
|
"start": 4.2,
|
|
"end": 9.8,
|
|
"text": "Let's review the deployment."
|
|
}
|
|
]
|
|
}
|
|
```
|
|
|
|
---
|
|
|
|
## Meeting Summaries
|
|
|
|
```sql
|
|
CREATE TABLE meeting_summaries (
|
|
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
|
|
|
|
meeting_id UUID NOT NULL
|
|
REFERENCES meetings(id)
|
|
ON DELETE CASCADE,
|
|
|
|
summary_type VARCHAR(50)
|
|
NOT NULL DEFAULT 'general',
|
|
|
|
content TEXT NOT NULL,
|
|
|
|
model VARCHAR(100),
|
|
|
|
metadata JSONB NOT NULL DEFAULT '{}',
|
|
|
|
created_at TIMESTAMPTZ NOT NULL DEFAULT NOW()
|
|
);
|
|
```
|
|
|
|
Possible summary types:
|
|
|
|
```text
|
|
general
|
|
short
|
|
executive
|
|
technical
|
|
action-focused
|
|
```
|
|
|
|
---
|
|
|
|
## Meeting Tasks
|
|
|
|
```sql
|
|
CREATE TABLE meeting_tasks (
|
|
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
|
|
|
|
meeting_id UUID NOT NULL
|
|
REFERENCES meetings(id)
|
|
ON DELETE CASCADE,
|
|
|
|
assigned_user_id UUID
|
|
REFERENCES users(id)
|
|
ON DELETE SET NULL,
|
|
|
|
title VARCHAR(500) NOT NULL,
|
|
description TEXT,
|
|
|
|
due_at TIMESTAMPTZ,
|
|
|
|
status VARCHAR(30)
|
|
NOT NULL DEFAULT 'open'
|
|
CHECK (
|
|
status IN (
|
|
'open',
|
|
'in_progress',
|
|
'completed',
|
|
'cancelled'
|
|
)
|
|
),
|
|
|
|
source VARCHAR(30)
|
|
NOT NULL DEFAULT 'ai'
|
|
CHECK (
|
|
source IN (
|
|
'ai',
|
|
'manual',
|
|
'imported'
|
|
)
|
|
),
|
|
|
|
source_metadata JSONB,
|
|
|
|
created_at TIMESTAMPTZ NOT NULL DEFAULT NOW(),
|
|
updated_at TIMESTAMPTZ NOT NULL DEFAULT NOW()
|
|
);
|
|
```
|
|
|
|
---
|
|
|
|
## Meeting Schedules / Reminders
|
|
|
|
```sql
|
|
CREATE TABLE meeting_schedules (
|
|
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
|
|
|
|
meeting_id UUID NOT NULL
|
|
REFERENCES meetings(id)
|
|
ON DELETE CASCADE,
|
|
|
|
task_id UUID
|
|
REFERENCES meeting_tasks(id)
|
|
ON DELETE CASCADE,
|
|
|
|
user_id UUID NOT NULL
|
|
REFERENCES users(id)
|
|
ON DELETE CASCADE,
|
|
|
|
schedule_type VARCHAR(50) NOT NULL,
|
|
|
|
scheduled_at TIMESTAMPTZ NOT NULL,
|
|
|
|
status VARCHAR(30)
|
|
NOT NULL DEFAULT 'pending'
|
|
CHECK (
|
|
status IN (
|
|
'pending',
|
|
'processing',
|
|
'executed',
|
|
'failed',
|
|
'cancelled'
|
|
)
|
|
),
|
|
|
|
payload JSONB NOT NULL DEFAULT '{}',
|
|
|
|
executed_at TIMESTAMPTZ,
|
|
|
|
attempts INTEGER
|
|
NOT NULL DEFAULT 0,
|
|
|
|
last_error TEXT,
|
|
|
|
created_at TIMESTAMPTZ NOT NULL DEFAULT NOW(),
|
|
updated_at TIMESTAMPTZ NOT NULL DEFAULT NOW()
|
|
);
|
|
```
|
|
|
|
Example reminder payload:
|
|
|
|
```json
|
|
{
|
|
"channel": "email",
|
|
"message": "Prepare the deployment report"
|
|
}
|
|
```
|
|
|
|
---
|
|
|
|
## Processing Jobs
|
|
|
|
MeetVault can use PostgreSQL itself as the first job queue.
|
|
|
|
```sql
|
|
CREATE TABLE processing_jobs (
|
|
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
|
|
|
|
meeting_id UUID
|
|
REFERENCES meetings(id)
|
|
ON DELETE CASCADE,
|
|
|
|
job_type VARCHAR(50)
|
|
NOT NULL
|
|
CHECK (
|
|
job_type IN (
|
|
'summarize_meeting',
|
|
'extract_action_items',
|
|
'generate_reminders'
|
|
)
|
|
),
|
|
|
|
status VARCHAR(30)
|
|
NOT NULL DEFAULT 'queued'
|
|
CHECK (
|
|
status IN (
|
|
'queued',
|
|
'running',
|
|
'completed',
|
|
'failed',
|
|
'cancelled'
|
|
)
|
|
),
|
|
|
|
priority SMALLINT
|
|
NOT NULL DEFAULT 0,
|
|
|
|
attempts INTEGER
|
|
NOT NULL DEFAULT 0,
|
|
|
|
max_attempts INTEGER
|
|
NOT NULL DEFAULT 3
|
|
CHECK (max_attempts > 0),
|
|
|
|
input_data JSONB,
|
|
output_data JSONB,
|
|
|
|
error_message TEXT,
|
|
|
|
available_at TIMESTAMPTZ
|
|
NOT NULL DEFAULT NOW(),
|
|
|
|
locked_at TIMESTAMPTZ,
|
|
locked_by VARCHAR(255),
|
|
|
|
started_at TIMESTAMPTZ,
|
|
completed_at TIMESTAMPTZ,
|
|
|
|
created_at TIMESTAMPTZ NOT NULL DEFAULT NOW()
|
|
);
|
|
```
|
|
|
|
Example job pickup query:
|
|
|
|
```sql
|
|
SELECT *
|
|
FROM processing_jobs
|
|
WHERE status = 'queued'
|
|
AND available_at <= NOW()
|
|
ORDER BY priority DESC, created_at
|
|
FOR UPDATE SKIP LOCKED
|
|
LIMIT 1;
|
|
```
|
|
|
|
This allows multiple workers without two workers processing the same job.
|
|
|
|
Redis or another dedicated queue can be added later if required.
|
|
|
|
---
|
|
|
|
# Indexes
|
|
|
|
```sql
|
|
CREATE INDEX idx_meetings_owner
|
|
ON meetings(owner_user_id);
|
|
|
|
CREATE INDEX idx_meetings_owner_started
|
|
ON meetings(owner_user_id, started_at DESC);
|
|
|
|
CREATE INDEX idx_meetings_status
|
|
ON meetings(status);
|
|
|
|
CREATE INDEX idx_meeting_files_meeting
|
|
ON meeting_files(meeting_id);
|
|
|
|
CREATE INDEX idx_transcripts_meeting
|
|
ON transcripts(meeting_id);
|
|
|
|
CREATE INDEX idx_transcripts_active
|
|
ON transcripts(meeting_id, is_active);
|
|
|
|
CREATE INDEX idx_summaries_meeting
|
|
ON meeting_summaries(meeting_id);
|
|
|
|
CREATE INDEX idx_tasks_meeting
|
|
ON meeting_tasks(meeting_id);
|
|
|
|
CREATE INDEX idx_tasks_assigned_user
|
|
ON meeting_tasks(assigned_user_id);
|
|
|
|
CREATE INDEX idx_schedules_due
|
|
ON meeting_schedules(status, scheduled_at);
|
|
|
|
CREATE INDEX idx_processing_jobs_queue
|
|
ON processing_jobs(status, available_at, priority DESC);
|
|
```
|
|
|
|
The most important worker indexes are:
|
|
|
|
```text
|
|
meeting_schedules(status, scheduled_at)
|
|
processing_jobs(status, available_at, priority)
|
|
```
|
|
|
|
---
|
|
|
|
# Automatic `updated_at`
|
|
|
|
PostgreSQL does not automatically modify `updated_at` when a row changes.
|
|
|
|
Create a shared trigger function:
|
|
|
|
```sql
|
|
CREATE OR REPLACE FUNCTION set_updated_at()
|
|
RETURNS TRIGGER AS $$
|
|
BEGIN
|
|
NEW.updated_at = NOW();
|
|
RETURN NEW;
|
|
END;
|
|
$$ LANGUAGE plpgsql;
|
|
```
|
|
|
|
Apply it to relevant tables:
|
|
|
|
```sql
|
|
CREATE TRIGGER trg_users_updated_at
|
|
BEFORE UPDATE ON users
|
|
FOR EACH ROW
|
|
EXECUTE FUNCTION set_updated_at();
|
|
|
|
CREATE TRIGGER trg_subscriptions_updated_at
|
|
BEFORE UPDATE ON subscriptions
|
|
FOR EACH ROW
|
|
EXECUTE FUNCTION set_updated_at();
|
|
|
|
CREATE TRIGGER trg_meetings_updated_at
|
|
BEFORE UPDATE ON meetings
|
|
FOR EACH ROW
|
|
EXECUTE FUNCTION set_updated_at();
|
|
|
|
CREATE TRIGGER trg_transcripts_updated_at
|
|
BEFORE UPDATE ON transcripts
|
|
FOR EACH ROW
|
|
EXECUTE FUNCTION set_updated_at();
|
|
|
|
CREATE TRIGGER trg_tasks_updated_at
|
|
BEFORE UPDATE ON meeting_tasks
|
|
FOR EACH ROW
|
|
EXECUTE FUNCTION set_updated_at();
|
|
|
|
CREATE TRIGGER trg_schedules_updated_at
|
|
BEFORE UPDATE ON meeting_schedules
|
|
FOR EACH ROW
|
|
EXECUTE FUNCTION set_updated_at();
|
|
```
|
|
|
|
---
|
|
|
|
# Final Architecture
|
|
|
|
```text
|
|
MeetVault
|
|
|
|
USER MACHINE
|
|
│
|
|
┌──────────┴──────────┐
|
|
│ │
|
|
Recording Local STT
|
|
│ Whisper
|
|
│ │
|
|
▼ ▼
|
|
Video/audio Transcript
|
|
│ │
|
|
│ ▼
|
|
│ Backend API
|
|
│ │
|
|
│ ▼
|
|
│ PostgreSQL
|
|
│ │
|
|
│ processing_jobs
|
|
│ │
|
|
│ ▼
|
|
│ Summary Worker
|
|
│ / \
|
|
│ / \
|
|
▼ ▼ ▼
|
|
SeaweedFS meeting_tasks summaries
|
|
│
|
|
▼
|
|
meeting_schedules
|
|
│
|
|
▼
|
|
Reminder Worker
|
|
```
|
|
|
|
---
|
|
|
|
# Recommended MVP Stack
|
|
|
|
A practical first implementation could use:
|
|
|
|
```text
|
|
Frontend
|
|
React / Next.js
|
|
|
|
Backend
|
|
FastAPI or Node.js
|
|
|
|
Database
|
|
PostgreSQL
|
|
|
|
Media Storage
|
|
SeaweedFS
|
|
|
|
Local Transcription
|
|
Whisper / whisper.cpp
|
|
|
|
Backend Jobs
|
|
PostgreSQL processing_jobs
|
|
|
|
Workers
|
|
Summarization Worker
|
|
Reminder Worker
|
|
```
|
|
|
|
For the first version, MeetVault does not require:
|
|
|
|
- Kafka
|
|
- Elasticsearch
|
|
- Kubernetes
|
|
- A vector database
|
|
- A separate microservice architecture
|
|
- Redis, unless job volume grows enough to justify it
|
|
|
|
PostgreSQL + SeaweedFS + the backend API + workers is sufficient for the MVP and leaves room to scale later.
|