Files
MeetVault/AGENTS.md
supersanta 654033fba3 Add Phase 0 test harness (vitest + ESLint) and fix build config
- Add vitest unit tests for Zustand app store state transitions (6 tests)
- Add mock speech-engine WebSocket protocol integration test (spawns server.py --mock on a free port)
- Add ESLint 10 flat config with typescript-eslint; add typecheck/lint/test scripts to package.json
- Fix pre-existing tsc -b breakage via noEmit in tsconfig.node.json; add dedicated tsconfig.tests.json project for Node-side tests
- Track src-tauri/Cargo.lock for reproducible Rust builds; enable protocol-asset Tauri feature
- Ignore generated artifacts (*.tsbuildinfo, src-tauri/gen/, emitted vite.config.js/.d.ts); remove stale emitted vite config copies
- Update AGENTS.md agent guide (demo scope lock, roadmap, verification commands)
- Add HANDOFF.md documenting Phase 0 state and remaining steps for the next agent
2026-10-03 22:33:47 +07:00

30 KiB

MeetVault Desktop Demo — Agent Guide

This file is the operating guide for coding agents working on this repository. Read it before making changes.

1. Mission

MeetVault Desktop is a local-first meeting recorder and live transcription desktop app. The current repository is a runnable demo/prototype, not the full production MeetVault platform.

The demo must be able to:

  • run as a Tauri desktop application;
  • record microphone audio and optionally screen/system audio;
  • stream 16 kHz mono PCM to a local speech engine;
  • transcribe locally with Whisper;
  • diarize speakers locally as stable session IDs (speaker_1, speaker_2, ...);
  • let users rename speaker labels without changing the underlying speaker ID;
  • show live transcript text in the main window and a Tauri always-on-top overlay;
  • save recording/transcript data locally;
  • play the saved media in Meeting Details;
  • keep backend-only actions as demo placeholders until a real backend contract is provided.

Do not turn this repository into a cloud-dependent application during the demo phase.


2. Source-of-truth order

When documents disagree, use this order:

  1. PROJECT_SCOPE.md — current demo scope lock. Highest authority.
  2. This AGENT.md — implementation rules and roadmap.
  3. README.md — current runnable-demo documentation.
  4. docs/MeetVault_Desktop_README_original.md — intended production desktop architecture.
  5. docs/MeetVault_UI_Wireframes_v2.drawio — product/UI reference, including features intentionally removed from the current demo.

Important: the original wireframes contain account, subscription, SeaweedFS, summary, action-item, and other production concepts. Do not re-add them merely because they appear in the wireframes. PROJECT_SCOPE.md deliberately removes or stubs some of them for this demo.


3. Non-negotiable demo scope

Unless the user explicitly changes the scope, preserve all of the following:

  • No login UI or authentication flow.
  • No subscription UI or billing logic.
  • Recordings and transcript files remain local.
  • Meeting Details supports local audio/video playback.
  • Live transcription is local Whisper.
  • Speaker identification means session diarization, not biometric identity recognition.
  • Speaker labels can be renamed live.
  • The realtime overlay remains part of the demo.
  • Settings remain local.
  • Backend-only actions remain placeholders.
  • The placeholder text must remain exactly:
Processing (The job). Please wait a moment

Do not invent backend endpoints, credentials, authentication behavior, billing rules, or cloud schemas that are not supplied by the user or an actual backend repository/API contract.


4. Current stack

Desktop/UI

  • Tauri 2
  • React 18
  • TypeScript
  • Vite
  • React Router
  • Zustand
  • Lucide React

Tauri/Rust shell

Current Rust code is intentionally thin. It initializes:

  • tauri-plugin-fs
  • tauri-plugin-dialog

The overlay window is currently created from TypeScript.

Local speech engine

speech-engine/server.py is a local FastAPI/WebSocket process.

Real mode:

  • faster-whisper for speech-to-text
  • diart for streaming speaker diarization
  • NumPy for PCM processing

Mock mode:

  • returns deterministic speaker/transcript events for demo testing without loading Whisper or diart models.

Current local persistence

  • Zustand persist stores settings and meeting state in renderer storage.
  • Tauri filesystem APIs save completed meeting files under AppLocalData.
  • Media is stored as recording.webm.
  • meeting.json and transcript.json are also written.

This persistence model is prototype-grade and should be improved; see the roadmap below.


5. Repository map

MeetVault-Desktop-Demo/
├── AGENT.md                       Agent instructions
├── PROJECT_SCOPE.md               Current demo scope lock
├── README.md                      Current demo/run documentation
├── package.json                   Frontend/Tauri scripts
├── src/
│   ├── components/                Reusable React UI
│   ├── lib/
│   │   ├── backend.ts             Backend placeholder only
│   │   ├── capture.ts             WebView mic/display capture + MediaRecorder
│   │   ├── liveBus.ts             Main window ↔ overlay live-state events
│   │   ├── localFiles.ts          Current local file persistence
│   │   ├── speechClient.ts        WebSocket client for speech engine
│   │   └── tauri.ts               Runtime detection
│   ├── pages/
│   │   ├── Dashboard.tsx
│   │   ├── LiveMeeting.tsx        Current orchestration hotspot
│   │   ├── MeetingDetails.tsx
│   │   ├── Meetings.tsx
│   │   ├── Overlay.tsx
│   │   └── Settings.tsx
│   ├── store/useAppStore.ts       Current Zustand domain/UI state
│   └── types.ts                   Shared frontend types
├── src-tauri/
│   ├── capabilities/default.json  Tauri permissions
│   ├── src/lib.rs                 Tauri plugin/bootstrap code
│   └── tauri.conf.json            Desktop/bundle configuration
├── speech-engine/
│   ├── server.py                  Local Whisper + diart WebSocket server
│   ├── requirements.txt
│   └── README.md
├── scripts/
│   ├── start-demo.ps1
│   ├── start-demo-mock.ps1
│   └── start-demo.sh
└── docs/
    ├── MeetVault_Desktop_README_original.md
    └── MeetVault_UI_Wireframes_v2.drawio

Do not edit generated/dependency directories such as node_modules/, speech-engine/.venv/, dist/, or src-tauri/target/.


6. How to run the demo

6.1 Prerequisites

For the desktop app:

  • Node.js 20+
  • npm
  • Rust toolchain compatible with Tauri 2
  • Tauri platform prerequisites for the host OS

For the speech engine:

  • Python 3.12 is the environment used by the archived demo (3.12.10).
  • requirements.txt declares diart==0.9.2 and NumPy <2.
  • The current README still mentions Python 3.10/3.11 in places. Treat that as documentation drift and normalize the docs when dependency compatibility has been verified on the target machine.

For real diarization, additional platform/audio/model requirements may be needed by PyTorch/diart/pyannote.

6.2 Clean frontend install

Use the lockfile for reproducible installs:

npm ci

Use npm install only when intentionally changing dependencies and updating package-lock.json.

6.3 Fastest useful demo: desktop + mock speech

This is the preferred smoke-test path because it validates the Tauri window, microphone pipeline, overlay, local persistence, and meeting flow without requiring ML model setup.

Create the Python environment once:

cd speech-engine
python -m venv .venv

Activate it and install dependencies. The full environment is:

pip install -r requirements.txt

For a lightweight mock-only environment, server.py --mock only needs the packages imported at process start (FastAPI/Uvicorn/NumPy and their dependencies); it does not initialize Whisper or diart.

Then return to the repository root.

Windows convenience command:

scripts\start-demo-mock.ps1

Equivalent manual flow in two terminals:

python speech-engine/server.py --mock

and:

npm run desktop

The desktop connects to:

ws://127.0.0.1:8765/ws/live

6.4 Real speech mode

Install the complete speech environment:

cd speech-engine
python -m venv .venv
# activate the venv
pip install -r requirements.txt

Then start:

python speech-engine/server.py

and in another terminal:

npm run desktop

On Windows the convenience script is:

scripts\start-demo.ps1

Real diarization may require model access/terms/authentication required by the current diart/pyannote stack. Do not hard-code personal model tokens into this repository.

6.5 Browser-only UI development

npm run dev

Open:

http://127.0.0.1:1420

Browser mode is not a full acceptance test. Native overlay behavior, AppLocalData filesystem behavior, and other Tauri-only functionality are reduced or unavailable.

6.6 Current verification commands

Run at least:

npm run build
cargo check --manifest-path src-tauri/Cargo.toml
python -m py_compile speech-engine/server.py

The repository currently has no automated test script and no lint script. Adding them is an early roadmap item. Until then, do not claim tests/lint passed when those commands do not exist.


7. Demo acceptance walkthrough

After a meaningful change, verify this flow in Tauri, preferably with mock speech first:

  1. Launch the speech engine or select client-side Mock mode.
  2. Launch the desktop app.
  3. Open Settings and confirm settings persist after navigation/restart.
  4. Start a new meeting.
  5. Grant microphone permission.
  6. Confirm the microphone level responds.
  7. Confirm transcript segments begin appearing.
  8. Confirm multiple speaker labels appear in mock/diarization mode.
  9. Rename a speaker and verify existing and new transcript rows use the new display name.
  10. Open the overlay.
  11. Verify the overlay receives live transcript updates.
  12. Rename a speaker from the overlay and verify the main window updates.
  13. Pause/resume and verify the final intended pause semantics once pause behavior is fixed.
  14. Finish the meeting.
  15. Confirm Meeting Details opens.
  16. Confirm the local recording is playable.
  17. Confirm the transcript is present.
  18. Confirm backend buttons show the exact processing placeholder rather than calling a fabricated backend.
  19. Restart the app and verify completed meeting data survives using the current persistence path.

When system/screen audio work is touched, also run a Windows-specific capture test; browser/WebView behavior alone is insufficient.


8. Current architecture

React/Tauri WebView
│
├─ LiveMeeting.tsx
│  ├─ startCapture()
│  │  ├─ getUserMedia(mic)
│  │  ├─ optional getDisplayMedia(screen/system audio)
│  │  ├─ Web Audio mixing/downsampling
│  │  └─ MediaRecorder
│  │
│  ├─ SpeechClient
│  │  └─ WebSocket → 127.0.0.1:8765
│  │
│  ├─ Zustand meeting state
│  ├─ liveBus → overlay
│  └─ saveMeetingLocally()
│
└─ Tauri shell
   ├─ filesystem plugin
   ├─ dialog plugin
   └─ overlay WebviewWindow

Python speech-engine/server.py
├─ faster-whisper
├─ diart
└─ timestamp-based transcript/speaker merge

This architecture is suitable for a demo, but too much session orchestration currently lives inside the React renderer.


9. Known problems and technical debt

Treat these as known facts before adding features.

P0 — long recording and data-integrity problems

Entire recording is buffered in renderer memory

capture.ts keeps every MediaRecorder chunk in an array and creates one final Blob on stop. localFiles.ts then converts that Blob to an ArrayBuffer before writing it.

This can use very large amounts of memory for multi-hour meetings and may duplicate the media in memory.

Target: incremental/chunked recording writes outside the renderer.

meeting.json is written before final meeting state is committed

LiveMeeting.finish() currently calls saveMeetingLocally(current, media) and only afterwards calls finishMeeting(...).

Therefore the saved meeting.json can contain stale values such as:

  • status: "recording"
  • no endedAt
  • no final duration
  • no media path

Fix the persistence transaction so final metadata and file state are consistent.

Persisted meeting files are not the source of truth

The UI primarily restores meetings from Zustand/localStorage. Completed local meeting folders are not scanned/reloaded into the application state on startup.

This prevents robust crash recovery and can create disagreement between localStorage and disk.

Target: a real local repository layer; renderer state should be a cache/view of persisted domain data, not the only authoritative copy.

P0 — pause semantics are wrong

CaptureController.pause() currently pauses only MediaRecorder. The Web Audio processor continues producing PCM and SpeechClient continues receiving it.

The UI therefore can say “Paused” while transcription continues.

Define one semantic and enforce it consistently:

  • pause recording + transcription, or
  • pause recording only and make that explicit in UI.

For this product, pausing the meeting should normally pause both recording and speech processing.

P0/P1 — speech timestamps and long sessions

The Python speech engine retains at most 20 minutes of audio and trims older samples. Whisper timestamps are then calculated from the remaining rolling buffer while diart has its own continuing timeline.

For meetings beyond the rolling window, verify/fix absolute meeting timestamps; do not let transcript time reset or drift after samples are dropped.

Also avoid repeated np.concatenate of an ever-growing array for every incoming PCM frame. Use a bounded/ring-buffer or chunk queue.

P1 — speech readiness/protocol

SpeechClient.connect() resolves when the WebSocket opens and the config is sent; it does not wait for the engine's ready response.

Add:

  • protocol version;
  • explicit ready handshake;
  • startup timeout;
  • engine/model status;
  • clean shutdown acknowledgement;
  • reconnect/error state where appropriate.

P1 — computeDevice: auto is not truly automatic

The Python server currently maps auto to CPU. If automatic CUDA/GPU selection is promised by the UI, implement actual capability detection or rename the option.

P1 — system audio is prototype-only

Current system audio uses getDisplayMedia. Reliability depends on OS/WebView/source selection.

For a Windows-quality product, put system audio behind a native audio abstraction and implement a Windows WASAPI loopback path rather than making WebView display capture the core audio backend.

P1 — settings are partly cosmetic

The store contains settings for launch at startup, tray, global hotkey, theme, and overlay click-through, but several are not wired to OS behavior/UI.

Do not mark a setting “implemented” until its native behavior is actually connected.

P1 — no test/lint harness

There is no test or lint script today. Add automated coverage before substantial refactors.

P2 — security hardening

Current Tauri CSP is null and the local FastAPI CORS configuration allows *.

That is acceptable only as a local demo shortcut. Before production packaging:

  • add a restrictive CSP;
  • minimize Tauri capabilities;
  • bind sidecar services to loopback only;
  • use a per-launch secret/token or private IPC mechanism if a local HTTP/WebSocket service remains;
  • never expose long-lived cloud credentials to the renderer;
  • never ship personal Hugging Face/API tokens in source or config.

10. Target production architecture

Do not rewrite everything at once. Evolve the current demo toward this boundary:

┌──────────────────────────────────────────────┐
│ React UI                                     │
│                                              │
│ pages/components                             │
│ view state only                              │
└───────────────────┬──────────────────────────┘
                    │ typed commands/events
                    ▼
┌──────────────────────────────────────────────┐
│ Tauri application/core layer                │
│                                              │
│ SessionService                               │
│ AudioService                                 │
│ RecordingService                             │
│ TranscriptService                            │
│ LocalMeetingRepository                       │
│ SpeechEngineManager                          │
│ OverlayService                               │
│ SyncClient (future)                          │
└───────┬──────────────┬──────────────┬────────┘
        │              │              │
        ▼              ▼              ▼
 Native audio     SQLite/local     Speech sidecar
 + recording      meeting files    or native runtime
        │                             │
        │                       Whisper + diarization
        │
        └──────────────┬──────────────────────────────
                       ▼
                AppLocalData
                meetings/{id}/media...

Future network boundary after the backend contract exists:

Desktop ── transcript/metadata ──► MeetVault Backend ──► PostgreSQL
Desktop ── presigned upload ─────► SeaweedFS

Core rule

The React renderer should request actions and render state. It should not eventually own long-running recording, crash recovery, secrets, backend credentials, or critical persistence orchestration.


11. Existing projects/components to reuse instead of reinventing

Keep Tauri 2

The current Tauri choice fits the product: multi-window overlay, native integration, filesystem access, background behavior, lower footprint than a full Electron runtime, and Rust escape hatches for native capture.

Do not migrate to Electron unless a concrete missing Tauri capability is demonstrated.

Short-term speech packaging: keep the Python engine as a Tauri sidecar

The lowest-risk path from demo to distributable app is:

  1. keep speech-engine/server.py while its behavior is being validated;
  2. package it into a standalone executable (for example with PyInstaller or equivalent);
  3. bundle it as a Tauri 2 sidecar/external binary;
  4. let Tauri start/stop and monitor the process;
  5. remove the requirement that end users manually install Python.

Tauri 2 officially supports bundled external binaries/sidecars, including Python CLI/API programs packaged as executables.

Medium-term transcription option: evaluate whisper.cpp

The original architecture already recommends whisper.cpp. It is a strong candidate for replacing the Python Whisper portion because it is designed for local C/C++ inference and provides real-time streaming examples.

Do not replace faster-whisper simply for architectural purity. Benchmark on target Windows hardware first for:

  • word accuracy;
  • end-to-end latency;
  • CPU/GPU usage;
  • memory;
  • model download/packaging size;
  • cold startup time.

If the Python engine meets product targets after sidecar packaging, keeping it is valid.

Diarization

Keep diart for the demo while validating quality. Treat diarization behind an interface so it can later be replaced by an ONNX/native runtime without changing UI/domain code.

The product needs session speaker separation, not biometric speaker recognition.

Local structured storage: SQLite

Move meeting metadata, segments, speaker mappings, job/sync state, and crash-recovery state to SQLite rather than renderer localStorage.

Tauri has an official SQL plugin with SQLite support, or the same repository boundary can be implemented in Rust directly. Whichever approach is chosen, keep SQL access behind LocalMeetingRepository so UI code does not scatter SQL statements through components.

Large media remains in files, not in SQLite blobs.

Native desktop settings

For production behavior, prefer official/native Tauri integrations for features such as:

  • global shortcut;
  • autostart;
  • tray/window lifecycle;
  • sidecar process control.

Do not simulate OS features in React state.

Windows audio

For reliable Windows system-audio capture, create a native Windows capture implementation using WASAPI loopback behind AudioService.

Keep the existing WebView capture path as a demo/fallback until the native path is proven.


Use SQLite for structured state and the filesystem for media.

Suggested logical tables/entities:

meetings
- id
- title
- started_at
- ended_at
- duration_ms
- status
- media_path
- media_mime
- local_only
- sync_status
- created_at
- updated_at

speakers
- meeting_id
- speaker_id
- display_name

transcript_segments
- id
- meeting_id
- speaker_id
- start_ms
- end_ms
- text
- is_final
- created_at
- updated_at

settings
- key
- value

sync_jobs                 # future
- id
- meeting_id
- type
- state
- attempts
- last_error
- updated_at

Suggested files:

AppLocalData/
└── meetings/
    └── {meeting_id}/
        ├── recording.webm        # demo/current format
        ├── recording.part-*      # optional temporary chunk files
        └── exports/              # future user exports only

A JSON transcript may still be generated as an export/interchange file, but should not become a second unsynchronized database.


13. Event and service boundaries

Prefer typed domain events over components reaching directly into engines.

Recommended events:

session.started
session.paused
session.resumed
session.finishing
session.finished

capture.started
capture.stopped
capture.level
capture.error

speech.starting
speech.ready
speech.partial
speech.final
speech.error
speech.stopped

speaker.detected
speaker.renamed
segment.reassigned

recording.chunk-written
recording.finalized
recording.error

sync.queued             # future
sync.progress           # future
sync.completed          # future
sync.failed             # future

Define event payloads in one shared contract rather than duplicating ad-hoc JSON shapes across the React app and Python process.


14. Development roadmap — do this in order

Phase 0 — protect the current demo

Before major refactoring:

  1. Add a reproducible smoke-test checklist.
  2. Add typecheck, lint, and test scripts.
  3. Add unit tests for state/domain transformations.
  4. Add at least one integration test for the speech WebSocket protocol using mock mode.
  5. Keep the mock demo runnable after every step.

Phase 1 — fix persistence and long-session safety

  1. Introduce LocalMeetingRepository.
  2. Make persisted meeting state authoritative.
  3. Fix finalization ordering so endedAt, duration, status, transcript, and media path commit consistently.
  4. Add startup recovery for unfinished sessions.
  5. Replace final-Blob-only media persistence with incremental/chunked writing.
  6. Make pause semantics consistent across recording and transcription.

This phase has higher priority than adding more UI features.

Phase 2 — separate application logic from React

Refactor LiveMeeting.tsx so it does not directly coordinate every subsystem.

Create services/controllers such as:

MeetingSessionController
CaptureService
SpeechEngineClient
RecordingService
OverlayService

React should subscribe to state/events and issue commands.

Phase 3 — productionize the speech process

  1. Version the speech protocol.
  2. Wait for explicit ready before treating the engine as usable.
  3. Add health/startup timeout/shutdown behavior.
  4. Fix long-session timestamp handling.
  5. Replace repeated large NumPy concatenations with a bounded buffer.
  6. Benchmark faster-whisper + diart on target Windows hardware.
  7. Package the current Python runtime as a Tauri sidecar.
  8. Separately benchmark whisper.cpp; migrate only if the measured tradeoff is better.

Phase 4 — native Windows audio

  1. Define an AudioService interface.
  2. Keep WebView microphone/display capture as a fallback.
  3. Add native microphone enumeration.
  4. Add WASAPI loopback system-audio capture.
  5. Add device-change/error recovery.
  6. Verify echo/mix behavior when microphone and system audio are both active.

Phase 5 — transcript correction features

Implement product features already anticipated by the original design:

  • edit transcript segment text;
  • reassign a segment to a different speaker;
  • rename speakers globally within the meeting;
  • search transcript;
  • jump playback to transcript timestamp;
  • export transcript.

Keep speaker corrections local unless a future backend contract specifies sync behavior.

Phase 6 — finish desktop integration

Wire the existing settings to real behavior:

  • autostart;
  • tray behavior;
  • global hotkey;
  • theme;
  • overlay click-through;
  • overlay lock/position persistence;
  • proper device selectors.

Phase 7 — backend/sync, only when the backend contract exists

The intended architecture from the supplied design is:

  • structured backend data in PostgreSQL;
  • large media in SeaweedFS;
  • desktop asks backend for temporary/presigned upload authorization;
  • desktop uploads media directly to storage;
  • transcript/metadata goes to the backend API;
  • recording/transcription continues offline;
  • failed uploads/sync are queued locally and retried.

Do not send large media through the application API merely for convenience.

Do not put permanent SeaweedFS credentials in the renderer.

Until the backend repository/OpenAPI contract is supplied, keep src/lib/backend.ts as a clearly isolated placeholder.

Phase 8 — packaging and release hardening

  • bundle the speech runtime/models or implement managed model download;
  • code-sign installers;
  • enforce restrictive Tauri capabilities/CSP;
  • handle sidecar startup/crash/update behavior;
  • validate upgrade/migration paths for SQLite;
  • add crash-safe session recovery;
  • test multi-hour meetings;
  • test CPU-only and GPU-capable Windows machines;
  • test no-network operation during a meeting.

15. Rules for changes made by agents

Preserve a runnable vertical slice

Do not leave the repository halfway between architectures. A refactor should keep at least mock-mode desktop operation working at each logical checkpoint.

Prefer small interfaces over large rewrites

Introduce adapters/services around current behavior first, then move implementations behind them.

Example:

bad: rewrite capture + speech + storage + UI simultaneously

good:
1. define RecordingService
2. wrap current MediaRecorder implementation
3. add tests/contracts
4. replace implementation with chunked/native writer

Do not silently broaden scope

Do not add login, payments, cloud storage, biometric speaker identity, Kafka, Kubernetes, Elasticsearch, vector databases, or a microservice fleet unless explicitly required.

Keep secrets out of the renderer and repository

Never commit:

  • API keys;
  • Hugging Face tokens;
  • storage credentials;
  • access tokens;
  • passwords.

Use environment/secure OS credential storage when such features are actually introduced.

Do not rely on checked-in dependency folders

The supplied archive contains development artifacts such as node_modules and a Windows .venv, but .gitignore correctly excludes them.

A clean checkout must be reproducible from:

  • package-lock.json;
  • speech-engine/requirements.txt;
  • Rust manifests/lockfile when present.

Do not design workflows that depend on copying someone else's .venv or node_modules directory.

Keep frontend types strict

tsconfig.app.json has strict: true. Do not weaken TypeScript strictness to make an error disappear.

Update documentation with behavior

If a command, prerequisite, port, protocol, storage location, or architecture boundary changes, update the relevant README/this guide in the same change.

Do not claim completion without verification

When reporting a task complete, state which checks actually ran and which could not run.


16. Suggested testing strategy

Add a lightweight stack rather than a large test framework rollout.

Frontend/domain

Use unit tests for:

  • meeting creation/finalization;
  • speaker rename;
  • transcript append/update;
  • segment reassignment;
  • formatting/mapping helpers;
  • repository adapters.

Speech protocol

Test the WebSocket contract in mock mode:

  • config accepted;
  • ready event;
  • PCM input;
  • final transcript event;
  • stop/shutdown;
  • malformed config;
  • unsupported sample rate.

Tauri integration

Manually/integration-test:

  • AppLocalData permissions;
  • overlay creation/events;
  • local media playback;
  • sidecar start/stop;
  • global shortcut/autostart when introduced.

Long-session tests

Add synthetic tests that simulate durations greater than 20 minutes so timestamp-reset/drift bugs are caught automatically.

Also stress recording for multi-hour memory growth before calling the recording path production-ready.


17. Definition of “good” for the next milestone

The project is ready to move beyond “demo scaffold” when all of the following are true:

  • A clean checkout installs and starts predictably.
  • Mock mode is one-command or nearly one-command.
  • The real speech engine is packaged or automatically managed by the desktop app.
  • Recording does not accumulate the whole meeting in renderer RAM.
  • Meeting state survives crashes/restarts from an authoritative local repository.
  • 60+ minute timestamps remain correct.
  • Pause/resume behavior is semantically correct.
  • Windows system audio is reliable enough for the stated product requirement.
  • Transcript text and speaker assignment can be corrected.
  • Overlay behavior/settings are fully wired.
  • Automated tests cover domain state and the mock speech protocol.
  • Tauri permissions/CSP are hardened for distribution.
  • No user is required to manually install Python for a packaged release.

After that milestone, backend/sync integration can be added against a real contract without destabilizing the core meeting-recording experience.


18. First actions for a new agent

When starting work in this repository:

  1. Read PROJECT_SCOPE.md.
  2. Read this file.
  3. Read the relevant section of README.md.
  4. If the task affects architecture, read docs/MeetVault_Desktop_README_original.md.
  5. Inspect the existing implementation before proposing a rewrite.
  6. Run the current build/check commands that are available.
  7. Prefer mock-mode smoke testing first.
  8. Make the smallest coherent change.
  9. Re-run verification.
  10. Document any command/architecture change.

If the requested task conflicts with PROJECT_SCOPE.md, do not silently choose one. Treat the user's newest explicit request as authority and update the scope documentation accordingly.