Add Phase 0 test harness (vitest + ESLint) and fix build config

- Add vitest unit tests for Zustand app store state transitions (6 tests)
- Add mock speech-engine WebSocket protocol integration test (spawns server.py --mock on a free port)
- Add ESLint 10 flat config with typescript-eslint; add typecheck/lint/test scripts to package.json
- Fix pre-existing tsc -b breakage via noEmit in tsconfig.node.json; add dedicated tsconfig.tests.json project for Node-side tests
- Track src-tauri/Cargo.lock for reproducible Rust builds; enable protocol-asset Tauri feature
- Ignore generated artifacts (*.tsbuildinfo, src-tauri/gen/, emitted vite.config.js/.d.ts); remove stale emitted vite config copies
- Update AGENTS.md agent guide (demo scope lock, roadmap, verification commands)
- Add HANDOFF.md documenting Phase 0 state and remaining steps for the next agent
This commit is contained in:
2026-10-03 22:33:47 +07:00
parent 9eff88215c
commit 654033fba3
15 changed files with 7824 additions and 61 deletions

970
AGENTS.md
View File

@@ -1,82 +1,944 @@
# MeetVault — Agent Guide
# MeetVault Desktop Demo — Agent Guide
## Repo layout
This file is the operating guide for coding agents working on this repository. Read it before making changes.
## 1. Mission
MeetVault Desktop is a **local-first meeting recorder and live transcription desktop app**. The current repository is a runnable demo/prototype, not the full production MeetVault platform.
The demo must be able to:
- run as a Tauri desktop application;
- record microphone audio and optionally screen/system audio;
- stream 16 kHz mono PCM to a local speech engine;
- transcribe locally with Whisper;
- diarize speakers locally as stable session IDs (`speaker_1`, `speaker_2`, ...);
- let users rename speaker labels without changing the underlying speaker ID;
- show live transcript text in the main window and a Tauri always-on-top overlay;
- save recording/transcript data locally;
- play the saved media in Meeting Details;
- keep backend-only actions as demo placeholders until a real backend contract is provided.
Do not turn this repository into a cloud-dependent application during the demo phase.
---
## 2. Source-of-truth order
When documents disagree, use this order:
1. **`PROJECT_SCOPE.md`** — current demo scope lock. Highest authority.
2. **This `AGENT.md`** — implementation rules and roadmap.
3. **`README.md`** — current runnable-demo documentation.
4. **`docs/MeetVault_Desktop_README_original.md`** — intended production desktop architecture.
5. **`docs/MeetVault_UI_Wireframes_v2.drawio`** — product/UI reference, including features intentionally removed from the current demo.
Important: the original wireframes contain account, subscription, SeaweedFS, summary, action-item, and other production concepts. **Do not re-add them merely because they appear in the wireframes.** `PROJECT_SCOPE.md` deliberately removes or stubs some of them for this demo.
---
## 3. Non-negotiable demo scope
Unless the user explicitly changes the scope, preserve all of the following:
- No login UI or authentication flow.
- No subscription UI or billing logic.
- Recordings and transcript files remain local.
- Meeting Details supports local audio/video playback.
- Live transcription is local Whisper.
- Speaker identification means **session diarization**, not biometric identity recognition.
- Speaker labels can be renamed live.
- The realtime overlay remains part of the demo.
- Settings remain local.
- Backend-only actions remain placeholders.
- The placeholder text must remain exactly:
```text
MeetVault/
├── Overview/ # Architecture docs, schema, data flow
Processing (The job). Please wait a moment
```
Do not invent backend endpoints, credentials, authentication behavior, billing rules, or cloud schemas that are not supplied by the user or an actual backend repository/API contract.
---
## 4. Current stack
### Desktop/UI
- Tauri 2
- React 18
- TypeScript
- Vite
- React Router
- Zustand
- Lucide React
### Tauri/Rust shell
Current Rust code is intentionally thin. It initializes:
- `tauri-plugin-fs`
- `tauri-plugin-dialog`
The overlay window is currently created from TypeScript.
### Local speech engine
`speech-engine/server.py` is a local FastAPI/WebSocket process.
Real mode:
- `faster-whisper` for speech-to-text
- `diart` for streaming speaker diarization
- NumPy for PCM processing
Mock mode:
- returns deterministic speaker/transcript events for demo testing without loading Whisper or diart models.
### Current local persistence
- Zustand persist stores settings and meeting state in renderer storage.
- Tauri filesystem APIs save completed meeting files under AppLocalData.
- Media is stored as `recording.webm`.
- `meeting.json` and `transcript.json` are also written.
This persistence model is prototype-grade and should be improved; see the roadmap below.
---
## 5. Repository map
```text
MeetVault-Desktop-Demo/
├── AGENT.md Agent instructions
├── PROJECT_SCOPE.md Current demo scope lock
├── README.md Current demo/run documentation
├── package.json Frontend/Tauri scripts
├── src/
│ ├── components/ Reusable React UI
│ ├── lib/
│ │ ├── backend.ts Backend placeholder only
│ │ ├── capture.ts WebView mic/display capture + MediaRecorder
│ │ ├── liveBus.ts Main window ↔ overlay live-state events
│ │ ├── localFiles.ts Current local file persistence
│ │ ├── speechClient.ts WebSocket client for speech engine
│ │ └── tauri.ts Runtime detection
│ ├── pages/
│ │ ├── Dashboard.tsx
│ │ ├── LiveMeeting.tsx Current orchestration hotspot
│ │ ├── MeetingDetails.tsx
│ │ ├── Meetings.tsx
│ │ ├── Overlay.tsx
│ │ └── Settings.tsx
│ ├── store/useAppStore.ts Current Zustand domain/UI state
│ └── types.ts Shared frontend types
├── src-tauri/
│ ├── capabilities/default.json Tauri permissions
│ ├── src/lib.rs Tauri plugin/bootstrap code
│ └── tauri.conf.json Desktop/bundle configuration
├── speech-engine/
│ ├── server.py Local Whisper + diart WebSocket server
│ ├── requirements.txt
│ └── README.md
├── Wireframes/ # Desktop client design notes
│ └── Desktop/
│ ├── Desktop_README.md
│ └── Desktop-Demo/ # Runnable Tauri + React demo
│ └── speech-engine/ # Local Whisper transcription engine
├── scripts/
│ ├── start-demo.ps1
│ ├── start-demo-mock.ps1
│ └── start-demo.sh
└── docs/
├── MeetVault_Desktop_README_original.md
└── MeetVault_UI_Wireframes_v2.drawio
```
## Core architecture (remember this)
Do not edit generated/dependency directories such as `node_modules/`, `speech-engine/.venv/`, `dist/`, or `src-tauri/target/`.
- **User machine**: records audio/video, runs local Whisper for STT, sends transcript to backend. Keeps compute and cost on the device.
- **Backend API**: auth, meeting metadata, presigned upload URLs, transcript ingestion, summarization jobs, task extraction, reminder scheduling.
- **PostgreSQL**: structured app data + `processing_jobs` table used as a job queue (no external broker needed for MVP).
- **SeaweedFS**: S3-compatible object storage for large media files.
---
## PostgreSQL gotchas
## 6. How to run the demo
### Enable UUID generation first
### 6.1 Prerequisites
```sql
CREATE EXTENSION IF NOT EXISTS pgcrypto;
For the desktop app:
- Node.js 20+
- npm
- Rust toolchain compatible with Tauri 2
- Tauri platform prerequisites for the host OS
For the speech engine:
- Python 3.12 is the environment used by the archived demo (`3.12.10`).
- `requirements.txt` declares `diart==0.9.2` and NumPy `<2`.
- The current README still mentions Python 3.10/3.11 in places. Treat that as documentation drift and normalize the docs when dependency compatibility has been verified on the target machine.
For real diarization, additional platform/audio/model requirements may be needed by PyTorch/diart/pyannote.
### 6.2 Clean frontend install
Use the lockfile for reproducible installs:
```bash
npm ci
```
### Always set `updated_at` manually
Use `npm install` only when intentionally changing dependencies and updating `package-lock.json`.
PostgreSQL does not auto-update it. Create a shared trigger function and apply to every table that tracks changes:
### 6.3 Fastest useful demo: desktop + mock speech
```sql
CREATE OR REPLACE FUNCTION set_updated_at()
RETURNS TRIGGER AS $$
BEGIN
NEW.updated_at = NOW();
RETURN NEW;
END;
$$ LANGUAGE plpgsql;
This is the preferred smoke-test path because it validates the Tauri window, microphone pipeline, overlay, local persistence, and meeting flow without requiring ML model setup.
-- then attach per-table triggers, e.g.:
CREATE TRIGGER trg_meetings_updated_at
BEFORE UPDATE ON meetings FOR EACH ROW EXECUTE FUNCTION set_updated_at();
Create the Python environment once:
```bash
cd speech-engine
python -m venv .venv
```
### Job queue pattern (no external broker)
Activate it and install dependencies. The full environment is:
Workers pick jobs with this exact query:
```sql
SELECT *
FROM processing_jobs
WHERE status = 'queued'
AND available_at <= NOW()
ORDER BY priority DESC, created_at
FOR UPDATE SKIP LOCKED
LIMIT 1;
```bash
pip install -r requirements.txt
```
Key indexes for workers:
For a lightweight mock-only environment, `server.py --mock` only needs the packages imported at process start (FastAPI/Uvicorn/NumPy and their dependencies); it does not initialize Whisper or diart.
```sql
CREATE INDEX idx_schedules_due ON meeting_schedules(status, scheduled_at);
CREATE INDEX idx_processing_jobs_queue ON processing_jobs(status, available_at, priority DESC);
Then return to the repository root.
Windows convenience command:
```powershell
scripts\start-demo-mock.ps1
```
### SeaweedFS object paths
Equivalent manual flow in two terminals:
Files live under: `meeting-storage/users/{user_id}/meetings/{meeting_id}/original/`
```bash
python speech-engine/server.py --mock
```
## Testing flow
and:
1. Run the desktop demo locally (Tauri + React) to verify UI and local transcription without backend dependencies.
2. Spin up PostgreSQL, run the schema SQL from `Overview/README.md`, then start the backend API.
3. End-to-end: record a short clip on the desktop client → upload via presigned URL → ingest transcript → trigger summarization job → verify summary appears in DB.
```bash
npm run desktop
```
## What is intentionally out of scope for MVP
The desktop connects to:
- Kafka, Elasticsearch, Kubernetes, vector DBs, Redis (unless job volume grows), microservices. PostgreSQL + SeaweedFS + workers is sufficient.
```text
ws://127.0.0.1:8765/ws/live
```
### 6.4 Real speech mode
Install the complete speech environment:
```bash
cd speech-engine
python -m venv .venv
# activate the venv
pip install -r requirements.txt
```
Then start:
```bash
python speech-engine/server.py
```
and in another terminal:
```bash
npm run desktop
```
On Windows the convenience script is:
```powershell
scripts\start-demo.ps1
```
Real diarization may require model access/terms/authentication required by the current diart/pyannote stack. Do not hard-code personal model tokens into this repository.
### 6.5 Browser-only UI development
```bash
npm run dev
```
Open:
```text
http://127.0.0.1:1420
```
Browser mode is **not** a full acceptance test. Native overlay behavior, AppLocalData filesystem behavior, and other Tauri-only functionality are reduced or unavailable.
### 6.6 Current verification commands
Run at least:
```bash
npm run build
cargo check --manifest-path src-tauri/Cargo.toml
python -m py_compile speech-engine/server.py
```
The repository currently has **no automated test script and no lint script**. Adding them is an early roadmap item. Until then, do not claim tests/lint passed when those commands do not exist.
---
## 7. Demo acceptance walkthrough
After a meaningful change, verify this flow in Tauri, preferably with mock speech first:
1. Launch the speech engine or select client-side Mock mode.
2. Launch the desktop app.
3. Open Settings and confirm settings persist after navigation/restart.
4. Start a new meeting.
5. Grant microphone permission.
6. Confirm the microphone level responds.
7. Confirm transcript segments begin appearing.
8. Confirm multiple speaker labels appear in mock/diarization mode.
9. Rename a speaker and verify existing and new transcript rows use the new display name.
10. Open the overlay.
11. Verify the overlay receives live transcript updates.
12. Rename a speaker from the overlay and verify the main window updates.
13. Pause/resume and verify the final intended pause semantics once pause behavior is fixed.
14. Finish the meeting.
15. Confirm Meeting Details opens.
16. Confirm the local recording is playable.
17. Confirm the transcript is present.
18. Confirm backend buttons show the exact processing placeholder rather than calling a fabricated backend.
19. Restart the app and verify completed meeting data survives using the current persistence path.
When system/screen audio work is touched, also run a Windows-specific capture test; browser/WebView behavior alone is insufficient.
---
## 8. Current architecture
```text
React/Tauri WebView
│
├─ LiveMeeting.tsx
│ ├─ startCapture()
│ │ ├─ getUserMedia(mic)
│ │ ├─ optional getDisplayMedia(screen/system audio)
│ │ ├─ Web Audio mixing/downsampling
│ │ └─ MediaRecorder
│ │
│ ├─ SpeechClient
│ │ └─ WebSocket → 127.0.0.1:8765
│ │
│ ├─ Zustand meeting state
│ ├─ liveBus → overlay
│ └─ saveMeetingLocally()
│
└─ Tauri shell
├─ filesystem plugin
├─ dialog plugin
└─ overlay WebviewWindow
Python speech-engine/server.py
├─ faster-whisper
├─ diart
└─ timestamp-based transcript/speaker merge
```
This architecture is suitable for a demo, but too much session orchestration currently lives inside the React renderer.
---
## 9. Known problems and technical debt
Treat these as known facts before adding features.
### P0 — long recording and data-integrity problems
#### Entire recording is buffered in renderer memory
`capture.ts` keeps every `MediaRecorder` chunk in an array and creates one final Blob on stop. `localFiles.ts` then converts that Blob to an ArrayBuffer before writing it.
This can use very large amounts of memory for multi-hour meetings and may duplicate the media in memory.
**Target:** incremental/chunked recording writes outside the renderer.
#### `meeting.json` is written before final meeting state is committed
`LiveMeeting.finish()` currently calls `saveMeetingLocally(current, media)` and only afterwards calls `finishMeeting(...)`.
Therefore the saved `meeting.json` can contain stale values such as:
- `status: "recording"`
- no `endedAt`
- no final duration
- no media path
Fix the persistence transaction so final metadata and file state are consistent.
#### Persisted meeting files are not the source of truth
The UI primarily restores meetings from Zustand/localStorage. Completed local meeting folders are not scanned/reloaded into the application state on startup.
This prevents robust crash recovery and can create disagreement between localStorage and disk.
**Target:** a real local repository layer; renderer state should be a cache/view of persisted domain data, not the only authoritative copy.
### P0 — pause semantics are wrong
`CaptureController.pause()` currently pauses only `MediaRecorder`. The Web Audio processor continues producing PCM and `SpeechClient` continues receiving it.
The UI therefore can say “Paused” while transcription continues.
Define one semantic and enforce it consistently:
- pause recording + transcription, or
- pause recording only and make that explicit in UI.
For this product, pausing the meeting should normally pause both recording and speech processing.
### P0/P1 — speech timestamps and long sessions
The Python speech engine retains at most 20 minutes of audio and trims older samples. Whisper timestamps are then calculated from the remaining rolling buffer while diart has its own continuing timeline.
For meetings beyond the rolling window, verify/fix absolute meeting timestamps; do not let transcript time reset or drift after samples are dropped.
Also avoid repeated `np.concatenate` of an ever-growing array for every incoming PCM frame. Use a bounded/ring-buffer or chunk queue.
### P1 — speech readiness/protocol
`SpeechClient.connect()` resolves when the WebSocket opens and the config is sent; it does not wait for the engine's `ready` response.
Add:
- protocol version;
- explicit ready handshake;
- startup timeout;
- engine/model status;
- clean shutdown acknowledgement;
- reconnect/error state where appropriate.
### P1 — `computeDevice: auto` is not truly automatic
The Python server currently maps `auto` to CPU. If automatic CUDA/GPU selection is promised by the UI, implement actual capability detection or rename the option.
### P1 — system audio is prototype-only
Current system audio uses `getDisplayMedia`. Reliability depends on OS/WebView/source selection.
For a Windows-quality product, put system audio behind a native audio abstraction and implement a Windows WASAPI loopback path rather than making WebView display capture the core audio backend.
### P1 — settings are partly cosmetic
The store contains settings for launch at startup, tray, global hotkey, theme, and overlay click-through, but several are not wired to OS behavior/UI.
Do not mark a setting “implemented” until its native behavior is actually connected.
### P1 — no test/lint harness
There is no `test` or `lint` script today. Add automated coverage before substantial refactors.
### P2 — security hardening
Current Tauri CSP is `null` and the local FastAPI CORS configuration allows `*`.
That is acceptable only as a local demo shortcut. Before production packaging:
- add a restrictive CSP;
- minimize Tauri capabilities;
- bind sidecar services to loopback only;
- use a per-launch secret/token or private IPC mechanism if a local HTTP/WebSocket service remains;
- never expose long-lived cloud credentials to the renderer;
- never ship personal Hugging Face/API tokens in source or config.
---
## 10. Target production architecture
Do **not** rewrite everything at once. Evolve the current demo toward this boundary:
```text
┌──────────────────────────────────────────────┐
│ React UI │
│ │
│ pages/components │
│ view state only │
└───────────────────┬──────────────────────────┘
│ typed commands/events
▼
┌──────────────────────────────────────────────┐
│ Tauri application/core layer │
│ │
│ SessionService │
│ AudioService │
│ RecordingService │
│ TranscriptService │
│ LocalMeetingRepository │
│ SpeechEngineManager │
│ OverlayService │
│ SyncClient (future) │
└───────┬──────────────┬──────────────┬────────┘
│ │ │
▼ ▼ ▼
Native audio SQLite/local Speech sidecar
+ recording meeting files or native runtime
│ │
│ Whisper + diarization
│
└──────────────┬──────────────────────────────
▼
AppLocalData
meetings/{id}/media...
Future network boundary after the backend contract exists:
Desktop ── transcript/metadata ──► MeetVault Backend ──► PostgreSQL
Desktop ── presigned upload ─────► SeaweedFS
```
### Core rule
The React renderer should **request actions and render state**. It should not eventually own long-running recording, crash recovery, secrets, backend credentials, or critical persistence orchestration.
---
## 11. Existing projects/components to reuse instead of reinventing
### Keep Tauri 2
The current Tauri choice fits the product: multi-window overlay, native integration, filesystem access, background behavior, lower footprint than a full Electron runtime, and Rust escape hatches for native capture.
Do not migrate to Electron unless a concrete missing Tauri capability is demonstrated.
### Short-term speech packaging: keep the Python engine as a Tauri sidecar
The lowest-risk path from demo to distributable app is:
1. keep `speech-engine/server.py` while its behavior is being validated;
2. package it into a standalone executable (for example with PyInstaller or equivalent);
3. bundle it as a Tauri 2 **sidecar/external binary**;
4. let Tauri start/stop and monitor the process;
5. remove the requirement that end users manually install Python.
Tauri 2 officially supports bundled external binaries/sidecars, including Python CLI/API programs packaged as executables.
### Medium-term transcription option: evaluate `whisper.cpp`
The original architecture already recommends `whisper.cpp`. It is a strong candidate for replacing the Python Whisper portion because it is designed for local C/C++ inference and provides real-time streaming examples.
Do not replace `faster-whisper` simply for architectural purity. Benchmark on target Windows hardware first for:
- word accuracy;
- end-to-end latency;
- CPU/GPU usage;
- memory;
- model download/packaging size;
- cold startup time.
If the Python engine meets product targets after sidecar packaging, keeping it is valid.
### Diarization
Keep diart for the demo while validating quality. Treat diarization behind an interface so it can later be replaced by an ONNX/native runtime without changing UI/domain code.
The product needs session speaker separation, **not biometric speaker recognition**.
### Local structured storage: SQLite
Move meeting metadata, segments, speaker mappings, job/sync state, and crash-recovery state to SQLite rather than renderer localStorage.
Tauri has an official SQL plugin with SQLite support, or the same repository boundary can be implemented in Rust directly. Whichever approach is chosen, keep SQL access behind `LocalMeetingRepository` so UI code does not scatter SQL statements through components.
Large media remains in files, not in SQLite blobs.
### Native desktop settings
For production behavior, prefer official/native Tauri integrations for features such as:
- global shortcut;
- autostart;
- tray/window lifecycle;
- sidecar process control.
Do not simulate OS features in React state.
### Windows audio
For reliable Windows system-audio capture, create a native Windows capture implementation using WASAPI loopback behind `AudioService`.
Keep the existing WebView capture path as a demo/fallback until the native path is proven.
---
## 12. Recommended local data model
Use SQLite for structured state and the filesystem for media.
Suggested logical tables/entities:
```text
meetings
- id
- title
- started_at
- ended_at
- duration_ms
- status
- media_path
- media_mime
- local_only
- sync_status
- created_at
- updated_at
speakers
- meeting_id
- speaker_id
- display_name
transcript_segments
- id
- meeting_id
- speaker_id
- start_ms
- end_ms
- text
- is_final
- created_at
- updated_at
settings
- key
- value
sync_jobs # future
- id
- meeting_id
- type
- state
- attempts
- last_error
- updated_at
```
Suggested files:
```text
AppLocalData/
└── meetings/
└── {meeting_id}/
├── recording.webm # demo/current format
├── recording.part-* # optional temporary chunk files
└── exports/ # future user exports only
```
A JSON transcript may still be generated as an **export/interchange file**, but should not become a second unsynchronized database.
---
## 13. Event and service boundaries
Prefer typed domain events over components reaching directly into engines.
Recommended events:
```text
session.started
session.paused
session.resumed
session.finishing
session.finished
capture.started
capture.stopped
capture.level
capture.error
speech.starting
speech.ready
speech.partial
speech.final
speech.error
speech.stopped
speaker.detected
speaker.renamed
segment.reassigned
recording.chunk-written
recording.finalized
recording.error
sync.queued # future
sync.progress # future
sync.completed # future
sync.failed # future
```
Define event payloads in one shared contract rather than duplicating ad-hoc JSON shapes across the React app and Python process.
---
## 14. Development roadmap — do this in order
### Phase 0 — protect the current demo
Before major refactoring:
1. Add a reproducible smoke-test checklist.
2. Add `typecheck`, `lint`, and test scripts.
3. Add unit tests for state/domain transformations.
4. Add at least one integration test for the speech WebSocket protocol using mock mode.
5. Keep the mock demo runnable after every step.
### Phase 1 — fix persistence and long-session safety
1. Introduce `LocalMeetingRepository`.
2. Make persisted meeting state authoritative.
3. Fix finalization ordering so `endedAt`, duration, status, transcript, and media path commit consistently.
4. Add startup recovery for unfinished sessions.
5. Replace final-Blob-only media persistence with incremental/chunked writing.
6. Make pause semantics consistent across recording and transcription.
This phase has higher priority than adding more UI features.
### Phase 2 — separate application logic from React
Refactor `LiveMeeting.tsx` so it does not directly coordinate every subsystem.
Create services/controllers such as:
```text
MeetingSessionController
CaptureService
SpeechEngineClient
RecordingService
OverlayService
```
React should subscribe to state/events and issue commands.
### Phase 3 — productionize the speech process
1. Version the speech protocol.
2. Wait for explicit `ready` before treating the engine as usable.
3. Add health/startup timeout/shutdown behavior.
4. Fix long-session timestamp handling.
5. Replace repeated large NumPy concatenations with a bounded buffer.
6. Benchmark `faster-whisper` + diart on target Windows hardware.
7. Package the current Python runtime as a Tauri sidecar.
8. Separately benchmark `whisper.cpp`; migrate only if the measured tradeoff is better.
### Phase 4 — native Windows audio
1. Define an `AudioService` interface.
2. Keep WebView microphone/display capture as a fallback.
3. Add native microphone enumeration.
4. Add WASAPI loopback system-audio capture.
5. Add device-change/error recovery.
6. Verify echo/mix behavior when microphone and system audio are both active.
### Phase 5 — transcript correction features
Implement product features already anticipated by the original design:
- edit transcript segment text;
- reassign a segment to a different speaker;
- rename speakers globally within the meeting;
- search transcript;
- jump playback to transcript timestamp;
- export transcript.
Keep speaker corrections local unless a future backend contract specifies sync behavior.
### Phase 6 — finish desktop integration
Wire the existing settings to real behavior:
- autostart;
- tray behavior;
- global hotkey;
- theme;
- overlay click-through;
- overlay lock/position persistence;
- proper device selectors.
### Phase 7 — backend/sync, only when the backend contract exists
The intended architecture from the supplied design is:
- structured backend data in PostgreSQL;
- large media in SeaweedFS;
- desktop asks backend for temporary/presigned upload authorization;
- desktop uploads media directly to storage;
- transcript/metadata goes to the backend API;
- recording/transcription continues offline;
- failed uploads/sync are queued locally and retried.
Do **not** send large media through the application API merely for convenience.
Do **not** put permanent SeaweedFS credentials in the renderer.
Until the backend repository/OpenAPI contract is supplied, keep `src/lib/backend.ts` as a clearly isolated placeholder.
### Phase 8 — packaging and release hardening
- bundle the speech runtime/models or implement managed model download;
- code-sign installers;
- enforce restrictive Tauri capabilities/CSP;
- handle sidecar startup/crash/update behavior;
- validate upgrade/migration paths for SQLite;
- add crash-safe session recovery;
- test multi-hour meetings;
- test CPU-only and GPU-capable Windows machines;
- test no-network operation during a meeting.
---
## 15. Rules for changes made by agents
### Preserve a runnable vertical slice
Do not leave the repository halfway between architectures. A refactor should keep at least mock-mode desktop operation working at each logical checkpoint.
### Prefer small interfaces over large rewrites
Introduce adapters/services around current behavior first, then move implementations behind them.
Example:
```text
bad: rewrite capture + speech + storage + UI simultaneously
good:
1. define RecordingService
2. wrap current MediaRecorder implementation
3. add tests/contracts
4. replace implementation with chunked/native writer
```
### Do not silently broaden scope
Do not add login, payments, cloud storage, biometric speaker identity, Kafka, Kubernetes, Elasticsearch, vector databases, or a microservice fleet unless explicitly required.
### Keep secrets out of the renderer and repository
Never commit:
- API keys;
- Hugging Face tokens;
- storage credentials;
- access tokens;
- passwords.
Use environment/secure OS credential storage when such features are actually introduced.
### Do not rely on checked-in dependency folders
The supplied archive contains development artifacts such as `node_modules` and a Windows `.venv`, but `.gitignore` correctly excludes them.
A clean checkout must be reproducible from:
- `package-lock.json`;
- `speech-engine/requirements.txt`;
- Rust manifests/lockfile when present.
Do not design workflows that depend on copying someone else's `.venv` or `node_modules` directory.
### Keep frontend types strict
`tsconfig.app.json` has `strict: true`. Do not weaken TypeScript strictness to make an error disappear.
### Update documentation with behavior
If a command, prerequisite, port, protocol, storage location, or architecture boundary changes, update the relevant README/this guide in the same change.
### Do not claim completion without verification
When reporting a task complete, state which checks actually ran and which could not run.
---
## 16. Suggested testing strategy
Add a lightweight stack rather than a large test framework rollout.
### Frontend/domain
Use unit tests for:
- meeting creation/finalization;
- speaker rename;
- transcript append/update;
- segment reassignment;
- formatting/mapping helpers;
- repository adapters.
### Speech protocol
Test the WebSocket contract in mock mode:
- config accepted;
- ready event;
- PCM input;
- final transcript event;
- stop/shutdown;
- malformed config;
- unsupported sample rate.
### Tauri integration
Manually/integration-test:
- AppLocalData permissions;
- overlay creation/events;
- local media playback;
- sidecar start/stop;
- global shortcut/autostart when introduced.
### Long-session tests
Add synthetic tests that simulate durations greater than 20 minutes so timestamp-reset/drift bugs are caught automatically.
Also stress recording for multi-hour memory growth before calling the recording path production-ready.
---
## 17. Definition of “good” for the next milestone
The project is ready to move beyond “demo scaffold” when all of the following are true:
- A clean checkout installs and starts predictably.
- Mock mode is one-command or nearly one-command.
- The real speech engine is packaged or automatically managed by the desktop app.
- Recording does not accumulate the whole meeting in renderer RAM.
- Meeting state survives crashes/restarts from an authoritative local repository.
- 60+ minute timestamps remain correct.
- Pause/resume behavior is semantically correct.
- Windows system audio is reliable enough for the stated product requirement.
- Transcript text and speaker assignment can be corrected.
- Overlay behavior/settings are fully wired.
- Automated tests cover domain state and the mock speech protocol.
- Tauri permissions/CSP are hardened for distribution.
- No user is required to manually install Python for a packaged release.
After that milestone, backend/sync integration can be added against a real contract without destabilizing the core meeting-recording experience.
---
## 18. First actions for a new agent
When starting work in this repository:
1. Read `PROJECT_SCOPE.md`.
2. Read this file.
3. Read the relevant section of `README.md`.
4. If the task affects architecture, read `docs/MeetVault_Desktop_README_original.md`.
5. Inspect the existing implementation before proposing a rewrite.
6. Run the current build/check commands that are available.
7. Prefer mock-mode smoke testing first.
8. Make the smallest coherent change.
9. Re-run verification.
10. Document any command/architecture change.
If the requested task conflicts with `PROJECT_SCOPE.md`, do not silently choose one. Treat the user's newest explicit request as authority and update the scope documentation accordingly.