Add Phase 0 test harness (vitest + ESLint) and fix build config
- Add vitest unit tests for Zustand app store state transitions (6 tests) - Add mock speech-engine WebSocket protocol integration test (spawns server.py --mock on a free port) - Add ESLint 10 flat config with typescript-eslint; add typecheck/lint/test scripts to package.json - Fix pre-existing tsc -b breakage via noEmit in tsconfig.node.json; add dedicated tsconfig.tests.json project for Node-side tests - Track src-tauri/Cargo.lock for reproducible Rust builds; enable protocol-asset Tauri feature - Ignore generated artifacts (*.tsbuildinfo, src-tauri/gen/, emitted vite.config.js/.d.ts); remove stale emitted vite config copies - Update AGENTS.md agent guide (demo scope lock, roadmap, verification commands) - Add HANDOFF.md documenting Phase 0 state and remaining steps for the next agent
This commit is contained in:
970
AGENTS.md
970
AGENTS.md
@@ -1,82 +1,944 @@
|
|||||||
# MeetVault — Agent Guide
|
# MeetVault Desktop Demo — Agent Guide
|
||||||
|
|
||||||
## Repo layout
|
This file is the operating guide for coding agents working on this repository. Read it before making changes.
|
||||||
|
|
||||||
|
## 1. Mission
|
||||||
|
|
||||||
|
MeetVault Desktop is a **local-first meeting recorder and live transcription desktop app**. The current repository is a runnable demo/prototype, not the full production MeetVault platform.
|
||||||
|
|
||||||
|
The demo must be able to:
|
||||||
|
|
||||||
|
- run as a Tauri desktop application;
|
||||||
|
- record microphone audio and optionally screen/system audio;
|
||||||
|
- stream 16 kHz mono PCM to a local speech engine;
|
||||||
|
- transcribe locally with Whisper;
|
||||||
|
- diarize speakers locally as stable session IDs (`speaker_1`, `speaker_2`, ...);
|
||||||
|
- let users rename speaker labels without changing the underlying speaker ID;
|
||||||
|
- show live transcript text in the main window and a Tauri always-on-top overlay;
|
||||||
|
- save recording/transcript data locally;
|
||||||
|
- play the saved media in Meeting Details;
|
||||||
|
- keep backend-only actions as demo placeholders until a real backend contract is provided.
|
||||||
|
|
||||||
|
Do not turn this repository into a cloud-dependent application during the demo phase.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2. Source-of-truth order
|
||||||
|
|
||||||
|
When documents disagree, use this order:
|
||||||
|
|
||||||
|
1. **`PROJECT_SCOPE.md`** — current demo scope lock. Highest authority.
|
||||||
|
2. **This `AGENT.md`** — implementation rules and roadmap.
|
||||||
|
3. **`README.md`** — current runnable-demo documentation.
|
||||||
|
4. **`docs/MeetVault_Desktop_README_original.md`** — intended production desktop architecture.
|
||||||
|
5. **`docs/MeetVault_UI_Wireframes_v2.drawio`** — product/UI reference, including features intentionally removed from the current demo.
|
||||||
|
|
||||||
|
Important: the original wireframes contain account, subscription, SeaweedFS, summary, action-item, and other production concepts. **Do not re-add them merely because they appear in the wireframes.** `PROJECT_SCOPE.md` deliberately removes or stubs some of them for this demo.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 3. Non-negotiable demo scope
|
||||||
|
|
||||||
|
Unless the user explicitly changes the scope, preserve all of the following:
|
||||||
|
|
||||||
|
- No login UI or authentication flow.
|
||||||
|
- No subscription UI or billing logic.
|
||||||
|
- Recordings and transcript files remain local.
|
||||||
|
- Meeting Details supports local audio/video playback.
|
||||||
|
- Live transcription is local Whisper.
|
||||||
|
- Speaker identification means **session diarization**, not biometric identity recognition.
|
||||||
|
- Speaker labels can be renamed live.
|
||||||
|
- The realtime overlay remains part of the demo.
|
||||||
|
- Settings remain local.
|
||||||
|
- Backend-only actions remain placeholders.
|
||||||
|
- The placeholder text must remain exactly:
|
||||||
|
|
||||||
```text
|
```text
|
||||||
MeetVault/
|
Processing (The job). Please wait a moment
|
||||||
├── Overview/ # Architecture docs, schema, data flow
|
```
|
||||||
|
|
||||||
|
Do not invent backend endpoints, credentials, authentication behavior, billing rules, or cloud schemas that are not supplied by the user or an actual backend repository/API contract.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 4. Current stack
|
||||||
|
|
||||||
|
### Desktop/UI
|
||||||
|
|
||||||
|
- Tauri 2
|
||||||
|
- React 18
|
||||||
|
- TypeScript
|
||||||
|
- Vite
|
||||||
|
- React Router
|
||||||
|
- Zustand
|
||||||
|
- Lucide React
|
||||||
|
|
||||||
|
### Tauri/Rust shell
|
||||||
|
|
||||||
|
Current Rust code is intentionally thin. It initializes:
|
||||||
|
|
||||||
|
- `tauri-plugin-fs`
|
||||||
|
- `tauri-plugin-dialog`
|
||||||
|
|
||||||
|
The overlay window is currently created from TypeScript.
|
||||||
|
|
||||||
|
### Local speech engine
|
||||||
|
|
||||||
|
`speech-engine/server.py` is a local FastAPI/WebSocket process.
|
||||||
|
|
||||||
|
Real mode:
|
||||||
|
|
||||||
|
- `faster-whisper` for speech-to-text
|
||||||
|
- `diart` for streaming speaker diarization
|
||||||
|
- NumPy for PCM processing
|
||||||
|
|
||||||
|
Mock mode:
|
||||||
|
|
||||||
|
- returns deterministic speaker/transcript events for demo testing without loading Whisper or diart models.
|
||||||
|
|
||||||
|
### Current local persistence
|
||||||
|
|
||||||
|
- Zustand persist stores settings and meeting state in renderer storage.
|
||||||
|
- Tauri filesystem APIs save completed meeting files under AppLocalData.
|
||||||
|
- Media is stored as `recording.webm`.
|
||||||
|
- `meeting.json` and `transcript.json` are also written.
|
||||||
|
|
||||||
|
This persistence model is prototype-grade and should be improved; see the roadmap below.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 5. Repository map
|
||||||
|
|
||||||
|
```text
|
||||||
|
MeetVault-Desktop-Demo/
|
||||||
|
├── AGENT.md Agent instructions
|
||||||
|
├── PROJECT_SCOPE.md Current demo scope lock
|
||||||
|
├── README.md Current demo/run documentation
|
||||||
|
├── package.json Frontend/Tauri scripts
|
||||||
|
├── src/
|
||||||
|
│ ├── components/ Reusable React UI
|
||||||
|
│ ├── lib/
|
||||||
|
│ │ ├── backend.ts Backend placeholder only
|
||||||
|
│ │ ├── capture.ts WebView mic/display capture + MediaRecorder
|
||||||
|
│ │ ├── liveBus.ts Main window ↔ overlay live-state events
|
||||||
|
│ │ ├── localFiles.ts Current local file persistence
|
||||||
|
│ │ ├── speechClient.ts WebSocket client for speech engine
|
||||||
|
│ │ └── tauri.ts Runtime detection
|
||||||
|
│ ├── pages/
|
||||||
|
│ │ ├── Dashboard.tsx
|
||||||
|
│ │ ├── LiveMeeting.tsx Current orchestration hotspot
|
||||||
|
│ │ ├── MeetingDetails.tsx
|
||||||
|
│ │ ├── Meetings.tsx
|
||||||
|
│ │ ├── Overlay.tsx
|
||||||
|
│ │ └── Settings.tsx
|
||||||
|
│ ├── store/useAppStore.ts Current Zustand domain/UI state
|
||||||
|
│ └── types.ts Shared frontend types
|
||||||
|
├── src-tauri/
|
||||||
|
│ ├── capabilities/default.json Tauri permissions
|
||||||
|
│ ├── src/lib.rs Tauri plugin/bootstrap code
|
||||||
|
│ └── tauri.conf.json Desktop/bundle configuration
|
||||||
|
├── speech-engine/
|
||||||
|
│ ├── server.py Local Whisper + diart WebSocket server
|
||||||
|
│ ├── requirements.txt
|
||||||
│ └── README.md
|
│ └── README.md
|
||||||
├── Wireframes/ # Desktop client design notes
|
├── scripts/
|
||||||
│ └── Desktop/
|
│ ├── start-demo.ps1
|
||||||
│ ├── Desktop_README.md
|
│ ├── start-demo-mock.ps1
|
||||||
│ └── Desktop-Demo/ # Runnable Tauri + React demo
|
│ └── start-demo.sh
|
||||||
│ └── speech-engine/ # Local Whisper transcription engine
|
└── docs/
|
||||||
|
├── MeetVault_Desktop_README_original.md
|
||||||
|
└── MeetVault_UI_Wireframes_v2.drawio
|
||||||
```
|
```
|
||||||
|
|
||||||
## Core architecture (remember this)
|
Do not edit generated/dependency directories such as `node_modules/`, `speech-engine/.venv/`, `dist/`, or `src-tauri/target/`.
|
||||||
|
|
||||||
- **User machine**: records audio/video, runs local Whisper for STT, sends transcript to backend. Keeps compute and cost on the device.
|
---
|
||||||
- **Backend API**: auth, meeting metadata, presigned upload URLs, transcript ingestion, summarization jobs, task extraction, reminder scheduling.
|
|
||||||
- **PostgreSQL**: structured app data + `processing_jobs` table used as a job queue (no external broker needed for MVP).
|
|
||||||
- **SeaweedFS**: S3-compatible object storage for large media files.
|
|
||||||
|
|
||||||
## PostgreSQL gotchas
|
## 6. How to run the demo
|
||||||
|
|
||||||
### Enable UUID generation first
|
### 6.1 Prerequisites
|
||||||
|
|
||||||
```sql
|
For the desktop app:
|
||||||
CREATE EXTENSION IF NOT EXISTS pgcrypto;
|
|
||||||
|
- Node.js 20+
|
||||||
|
- npm
|
||||||
|
- Rust toolchain compatible with Tauri 2
|
||||||
|
- Tauri platform prerequisites for the host OS
|
||||||
|
|
||||||
|
For the speech engine:
|
||||||
|
|
||||||
|
- Python 3.12 is the environment used by the archived demo (`3.12.10`).
|
||||||
|
- `requirements.txt` declares `diart==0.9.2` and NumPy `<2`.
|
||||||
|
- The current README still mentions Python 3.10/3.11 in places. Treat that as documentation drift and normalize the docs when dependency compatibility has been verified on the target machine.
|
||||||
|
|
||||||
|
For real diarization, additional platform/audio/model requirements may be needed by PyTorch/diart/pyannote.
|
||||||
|
|
||||||
|
### 6.2 Clean frontend install
|
||||||
|
|
||||||
|
Use the lockfile for reproducible installs:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
npm ci
|
||||||
```
|
```
|
||||||
|
|
||||||
### Always set `updated_at` manually
|
Use `npm install` only when intentionally changing dependencies and updating `package-lock.json`.
|
||||||
|
|
||||||
PostgreSQL does not auto-update it. Create a shared trigger function and apply to every table that tracks changes:
|
### 6.3 Fastest useful demo: desktop + mock speech
|
||||||
|
|
||||||
```sql
|
This is the preferred smoke-test path because it validates the Tauri window, microphone pipeline, overlay, local persistence, and meeting flow without requiring ML model setup.
|
||||||
CREATE OR REPLACE FUNCTION set_updated_at()
|
|
||||||
RETURNS TRIGGER AS $$
|
|
||||||
BEGIN
|
|
||||||
NEW.updated_at = NOW();
|
|
||||||
RETURN NEW;
|
|
||||||
END;
|
|
||||||
$$ LANGUAGE plpgsql;
|
|
||||||
|
|
||||||
-- then attach per-table triggers, e.g.:
|
Create the Python environment once:
|
||||||
CREATE TRIGGER trg_meetings_updated_at
|
|
||||||
BEFORE UPDATE ON meetings FOR EACH ROW EXECUTE FUNCTION set_updated_at();
|
```bash
|
||||||
|
cd speech-engine
|
||||||
|
python -m venv .venv
|
||||||
```
|
```
|
||||||
|
|
||||||
### Job queue pattern (no external broker)
|
Activate it and install dependencies. The full environment is:
|
||||||
|
|
||||||
Workers pick jobs with this exact query:
|
```bash
|
||||||
|
pip install -r requirements.txt
|
||||||
```sql
|
|
||||||
SELECT *
|
|
||||||
FROM processing_jobs
|
|
||||||
WHERE status = 'queued'
|
|
||||||
AND available_at <= NOW()
|
|
||||||
ORDER BY priority DESC, created_at
|
|
||||||
FOR UPDATE SKIP LOCKED
|
|
||||||
LIMIT 1;
|
|
||||||
```
|
```
|
||||||
|
|
||||||
Key indexes for workers:
|
For a lightweight mock-only environment, `server.py --mock` only needs the packages imported at process start (FastAPI/Uvicorn/NumPy and their dependencies); it does not initialize Whisper or diart.
|
||||||
|
|
||||||
```sql
|
Then return to the repository root.
|
||||||
CREATE INDEX idx_schedules_due ON meeting_schedules(status, scheduled_at);
|
|
||||||
CREATE INDEX idx_processing_jobs_queue ON processing_jobs(status, available_at, priority DESC);
|
Windows convenience command:
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
scripts\start-demo-mock.ps1
|
||||||
```
|
```
|
||||||
|
|
||||||
### SeaweedFS object paths
|
Equivalent manual flow in two terminals:
|
||||||
|
|
||||||
Files live under: `meeting-storage/users/{user_id}/meetings/{meeting_id}/original/`
|
```bash
|
||||||
|
python speech-engine/server.py --mock
|
||||||
|
```
|
||||||
|
|
||||||
## Testing flow
|
and:
|
||||||
|
|
||||||
1. Run the desktop demo locally (Tauri + React) to verify UI and local transcription without backend dependencies.
|
```bash
|
||||||
2. Spin up PostgreSQL, run the schema SQL from `Overview/README.md`, then start the backend API.
|
npm run desktop
|
||||||
3. End-to-end: record a short clip on the desktop client → upload via presigned URL → ingest transcript → trigger summarization job → verify summary appears in DB.
|
```
|
||||||
|
|
||||||
## What is intentionally out of scope for MVP
|
The desktop connects to:
|
||||||
|
|
||||||
- Kafka, Elasticsearch, Kubernetes, vector DBs, Redis (unless job volume grows), microservices. PostgreSQL + SeaweedFS + workers is sufficient.
|
```text
|
||||||
|
ws://127.0.0.1:8765/ws/live
|
||||||
|
```
|
||||||
|
|
||||||
|
### 6.4 Real speech mode
|
||||||
|
|
||||||
|
Install the complete speech environment:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
cd speech-engine
|
||||||
|
python -m venv .venv
|
||||||
|
# activate the venv
|
||||||
|
pip install -r requirements.txt
|
||||||
|
```
|
||||||
|
|
||||||
|
Then start:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
python speech-engine/server.py
|
||||||
|
```
|
||||||
|
|
||||||
|
and in another terminal:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
npm run desktop
|
||||||
|
```
|
||||||
|
|
||||||
|
On Windows the convenience script is:
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
scripts\start-demo.ps1
|
||||||
|
```
|
||||||
|
|
||||||
|
Real diarization may require model access/terms/authentication required by the current diart/pyannote stack. Do not hard-code personal model tokens into this repository.
|
||||||
|
|
||||||
|
### 6.5 Browser-only UI development
|
||||||
|
|
||||||
|
```bash
|
||||||
|
npm run dev
|
||||||
|
```
|
||||||
|
|
||||||
|
Open:
|
||||||
|
|
||||||
|
```text
|
||||||
|
http://127.0.0.1:1420
|
||||||
|
```
|
||||||
|
|
||||||
|
Browser mode is **not** a full acceptance test. Native overlay behavior, AppLocalData filesystem behavior, and other Tauri-only functionality are reduced or unavailable.
|
||||||
|
|
||||||
|
### 6.6 Current verification commands
|
||||||
|
|
||||||
|
Run at least:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
npm run build
|
||||||
|
cargo check --manifest-path src-tauri/Cargo.toml
|
||||||
|
python -m py_compile speech-engine/server.py
|
||||||
|
```
|
||||||
|
|
||||||
|
The repository currently has **no automated test script and no lint script**. Adding them is an early roadmap item. Until then, do not claim tests/lint passed when those commands do not exist.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 7. Demo acceptance walkthrough
|
||||||
|
|
||||||
|
After a meaningful change, verify this flow in Tauri, preferably with mock speech first:
|
||||||
|
|
||||||
|
1. Launch the speech engine or select client-side Mock mode.
|
||||||
|
2. Launch the desktop app.
|
||||||
|
3. Open Settings and confirm settings persist after navigation/restart.
|
||||||
|
4. Start a new meeting.
|
||||||
|
5. Grant microphone permission.
|
||||||
|
6. Confirm the microphone level responds.
|
||||||
|
7. Confirm transcript segments begin appearing.
|
||||||
|
8. Confirm multiple speaker labels appear in mock/diarization mode.
|
||||||
|
9. Rename a speaker and verify existing and new transcript rows use the new display name.
|
||||||
|
10. Open the overlay.
|
||||||
|
11. Verify the overlay receives live transcript updates.
|
||||||
|
12. Rename a speaker from the overlay and verify the main window updates.
|
||||||
|
13. Pause/resume and verify the final intended pause semantics once pause behavior is fixed.
|
||||||
|
14. Finish the meeting.
|
||||||
|
15. Confirm Meeting Details opens.
|
||||||
|
16. Confirm the local recording is playable.
|
||||||
|
17. Confirm the transcript is present.
|
||||||
|
18. Confirm backend buttons show the exact processing placeholder rather than calling a fabricated backend.
|
||||||
|
19. Restart the app and verify completed meeting data survives using the current persistence path.
|
||||||
|
|
||||||
|
When system/screen audio work is touched, also run a Windows-specific capture test; browser/WebView behavior alone is insufficient.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 8. Current architecture
|
||||||
|
|
||||||
|
```text
|
||||||
|
React/Tauri WebView
|
||||||
|
│
|
||||||
|
├─ LiveMeeting.tsx
|
||||||
|
│ ├─ startCapture()
|
||||||
|
│ │ ├─ getUserMedia(mic)
|
||||||
|
│ │ ├─ optional getDisplayMedia(screen/system audio)
|
||||||
|
│ │ ├─ Web Audio mixing/downsampling
|
||||||
|
│ │ └─ MediaRecorder
|
||||||
|
│ │
|
||||||
|
│ ├─ SpeechClient
|
||||||
|
│ │ └─ WebSocket → 127.0.0.1:8765
|
||||||
|
│ │
|
||||||
|
│ ├─ Zustand meeting state
|
||||||
|
│ ├─ liveBus → overlay
|
||||||
|
│ └─ saveMeetingLocally()
|
||||||
|
│
|
||||||
|
└─ Tauri shell
|
||||||
|
├─ filesystem plugin
|
||||||
|
├─ dialog plugin
|
||||||
|
└─ overlay WebviewWindow
|
||||||
|
|
||||||
|
Python speech-engine/server.py
|
||||||
|
├─ faster-whisper
|
||||||
|
├─ diart
|
||||||
|
└─ timestamp-based transcript/speaker merge
|
||||||
|
```
|
||||||
|
|
||||||
|
This architecture is suitable for a demo, but too much session orchestration currently lives inside the React renderer.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 9. Known problems and technical debt
|
||||||
|
|
||||||
|
Treat these as known facts before adding features.
|
||||||
|
|
||||||
|
### P0 — long recording and data-integrity problems
|
||||||
|
|
||||||
|
#### Entire recording is buffered in renderer memory
|
||||||
|
|
||||||
|
`capture.ts` keeps every `MediaRecorder` chunk in an array and creates one final Blob on stop. `localFiles.ts` then converts that Blob to an ArrayBuffer before writing it.
|
||||||
|
|
||||||
|
This can use very large amounts of memory for multi-hour meetings and may duplicate the media in memory.
|
||||||
|
|
||||||
|
**Target:** incremental/chunked recording writes outside the renderer.
|
||||||
|
|
||||||
|
#### `meeting.json` is written before final meeting state is committed
|
||||||
|
|
||||||
|
`LiveMeeting.finish()` currently calls `saveMeetingLocally(current, media)` and only afterwards calls `finishMeeting(...)`.
|
||||||
|
|
||||||
|
Therefore the saved `meeting.json` can contain stale values such as:
|
||||||
|
|
||||||
|
- `status: "recording"`
|
||||||
|
- no `endedAt`
|
||||||
|
- no final duration
|
||||||
|
- no media path
|
||||||
|
|
||||||
|
Fix the persistence transaction so final metadata and file state are consistent.
|
||||||
|
|
||||||
|
#### Persisted meeting files are not the source of truth
|
||||||
|
|
||||||
|
The UI primarily restores meetings from Zustand/localStorage. Completed local meeting folders are not scanned/reloaded into the application state on startup.
|
||||||
|
|
||||||
|
This prevents robust crash recovery and can create disagreement between localStorage and disk.
|
||||||
|
|
||||||
|
**Target:** a real local repository layer; renderer state should be a cache/view of persisted domain data, not the only authoritative copy.
|
||||||
|
|
||||||
|
### P0 — pause semantics are wrong
|
||||||
|
|
||||||
|
`CaptureController.pause()` currently pauses only `MediaRecorder`. The Web Audio processor continues producing PCM and `SpeechClient` continues receiving it.
|
||||||
|
|
||||||
|
The UI therefore can say “Paused” while transcription continues.
|
||||||
|
|
||||||
|
Define one semantic and enforce it consistently:
|
||||||
|
|
||||||
|
- pause recording + transcription, or
|
||||||
|
- pause recording only and make that explicit in UI.
|
||||||
|
|
||||||
|
For this product, pausing the meeting should normally pause both recording and speech processing.
|
||||||
|
|
||||||
|
### P0/P1 — speech timestamps and long sessions
|
||||||
|
|
||||||
|
The Python speech engine retains at most 20 minutes of audio and trims older samples. Whisper timestamps are then calculated from the remaining rolling buffer while diart has its own continuing timeline.
|
||||||
|
|
||||||
|
For meetings beyond the rolling window, verify/fix absolute meeting timestamps; do not let transcript time reset or drift after samples are dropped.
|
||||||
|
|
||||||
|
Also avoid repeated `np.concatenate` of an ever-growing array for every incoming PCM frame. Use a bounded/ring-buffer or chunk queue.
|
||||||
|
|
||||||
|
### P1 — speech readiness/protocol
|
||||||
|
|
||||||
|
`SpeechClient.connect()` resolves when the WebSocket opens and the config is sent; it does not wait for the engine's `ready` response.
|
||||||
|
|
||||||
|
Add:
|
||||||
|
|
||||||
|
- protocol version;
|
||||||
|
- explicit ready handshake;
|
||||||
|
- startup timeout;
|
||||||
|
- engine/model status;
|
||||||
|
- clean shutdown acknowledgement;
|
||||||
|
- reconnect/error state where appropriate.
|
||||||
|
|
||||||
|
### P1 — `computeDevice: auto` is not truly automatic
|
||||||
|
|
||||||
|
The Python server currently maps `auto` to CPU. If automatic CUDA/GPU selection is promised by the UI, implement actual capability detection or rename the option.
|
||||||
|
|
||||||
|
### P1 — system audio is prototype-only
|
||||||
|
|
||||||
|
Current system audio uses `getDisplayMedia`. Reliability depends on OS/WebView/source selection.
|
||||||
|
|
||||||
|
For a Windows-quality product, put system audio behind a native audio abstraction and implement a Windows WASAPI loopback path rather than making WebView display capture the core audio backend.
|
||||||
|
|
||||||
|
### P1 — settings are partly cosmetic
|
||||||
|
|
||||||
|
The store contains settings for launch at startup, tray, global hotkey, theme, and overlay click-through, but several are not wired to OS behavior/UI.
|
||||||
|
|
||||||
|
Do not mark a setting “implemented” until its native behavior is actually connected.
|
||||||
|
|
||||||
|
### P1 — no test/lint harness
|
||||||
|
|
||||||
|
There is no `test` or `lint` script today. Add automated coverage before substantial refactors.
|
||||||
|
|
||||||
|
### P2 — security hardening
|
||||||
|
|
||||||
|
Current Tauri CSP is `null` and the local FastAPI CORS configuration allows `*`.
|
||||||
|
|
||||||
|
That is acceptable only as a local demo shortcut. Before production packaging:
|
||||||
|
|
||||||
|
- add a restrictive CSP;
|
||||||
|
- minimize Tauri capabilities;
|
||||||
|
- bind sidecar services to loopback only;
|
||||||
|
- use a per-launch secret/token or private IPC mechanism if a local HTTP/WebSocket service remains;
|
||||||
|
- never expose long-lived cloud credentials to the renderer;
|
||||||
|
- never ship personal Hugging Face/API tokens in source or config.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 10. Target production architecture
|
||||||
|
|
||||||
|
Do **not** rewrite everything at once. Evolve the current demo toward this boundary:
|
||||||
|
|
||||||
|
```text
|
||||||
|
┌──────────────────────────────────────────────┐
|
||||||
|
│ React UI │
|
||||||
|
│ │
|
||||||
|
│ pages/components │
|
||||||
|
│ view state only │
|
||||||
|
└───────────────────┬──────────────────────────┘
|
||||||
|
│ typed commands/events
|
||||||
|
▼
|
||||||
|
┌──────────────────────────────────────────────┐
|
||||||
|
│ Tauri application/core layer │
|
||||||
|
│ │
|
||||||
|
│ SessionService │
|
||||||
|
│ AudioService │
|
||||||
|
│ RecordingService │
|
||||||
|
│ TranscriptService │
|
||||||
|
│ LocalMeetingRepository │
|
||||||
|
│ SpeechEngineManager │
|
||||||
|
│ OverlayService │
|
||||||
|
│ SyncClient (future) │
|
||||||
|
└───────┬──────────────┬──────────────┬────────┘
|
||||||
|
│ │ │
|
||||||
|
▼ ▼ ▼
|
||||||
|
Native audio SQLite/local Speech sidecar
|
||||||
|
+ recording meeting files or native runtime
|
||||||
|
│ │
|
||||||
|
│ Whisper + diarization
|
||||||
|
│
|
||||||
|
└──────────────┬──────────────────────────────
|
||||||
|
▼
|
||||||
|
AppLocalData
|
||||||
|
meetings/{id}/media...
|
||||||
|
|
||||||
|
Future network boundary after the backend contract exists:
|
||||||
|
|
||||||
|
Desktop ── transcript/metadata ──► MeetVault Backend ──► PostgreSQL
|
||||||
|
Desktop ── presigned upload ─────► SeaweedFS
|
||||||
|
```
|
||||||
|
|
||||||
|
### Core rule
|
||||||
|
|
||||||
|
The React renderer should **request actions and render state**. It should not eventually own long-running recording, crash recovery, secrets, backend credentials, or critical persistence orchestration.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 11. Existing projects/components to reuse instead of reinventing
|
||||||
|
|
||||||
|
### Keep Tauri 2
|
||||||
|
|
||||||
|
The current Tauri choice fits the product: multi-window overlay, native integration, filesystem access, background behavior, lower footprint than a full Electron runtime, and Rust escape hatches for native capture.
|
||||||
|
|
||||||
|
Do not migrate to Electron unless a concrete missing Tauri capability is demonstrated.
|
||||||
|
|
||||||
|
### Short-term speech packaging: keep the Python engine as a Tauri sidecar
|
||||||
|
|
||||||
|
The lowest-risk path from demo to distributable app is:
|
||||||
|
|
||||||
|
1. keep `speech-engine/server.py` while its behavior is being validated;
|
||||||
|
2. package it into a standalone executable (for example with PyInstaller or equivalent);
|
||||||
|
3. bundle it as a Tauri 2 **sidecar/external binary**;
|
||||||
|
4. let Tauri start/stop and monitor the process;
|
||||||
|
5. remove the requirement that end users manually install Python.
|
||||||
|
|
||||||
|
Tauri 2 officially supports bundled external binaries/sidecars, including Python CLI/API programs packaged as executables.
|
||||||
|
|
||||||
|
### Medium-term transcription option: evaluate `whisper.cpp`
|
||||||
|
|
||||||
|
The original architecture already recommends `whisper.cpp`. It is a strong candidate for replacing the Python Whisper portion because it is designed for local C/C++ inference and provides real-time streaming examples.
|
||||||
|
|
||||||
|
Do not replace `faster-whisper` simply for architectural purity. Benchmark on target Windows hardware first for:
|
||||||
|
|
||||||
|
- word accuracy;
|
||||||
|
- end-to-end latency;
|
||||||
|
- CPU/GPU usage;
|
||||||
|
- memory;
|
||||||
|
- model download/packaging size;
|
||||||
|
- cold startup time.
|
||||||
|
|
||||||
|
If the Python engine meets product targets after sidecar packaging, keeping it is valid.
|
||||||
|
|
||||||
|
### Diarization
|
||||||
|
|
||||||
|
Keep diart for the demo while validating quality. Treat diarization behind an interface so it can later be replaced by an ONNX/native runtime without changing UI/domain code.
|
||||||
|
|
||||||
|
The product needs session speaker separation, **not biometric speaker recognition**.
|
||||||
|
|
||||||
|
### Local structured storage: SQLite
|
||||||
|
|
||||||
|
Move meeting metadata, segments, speaker mappings, job/sync state, and crash-recovery state to SQLite rather than renderer localStorage.
|
||||||
|
|
||||||
|
Tauri has an official SQL plugin with SQLite support, or the same repository boundary can be implemented in Rust directly. Whichever approach is chosen, keep SQL access behind `LocalMeetingRepository` so UI code does not scatter SQL statements through components.
|
||||||
|
|
||||||
|
Large media remains in files, not in SQLite blobs.
|
||||||
|
|
||||||
|
### Native desktop settings
|
||||||
|
|
||||||
|
For production behavior, prefer official/native Tauri integrations for features such as:
|
||||||
|
|
||||||
|
- global shortcut;
|
||||||
|
- autostart;
|
||||||
|
- tray/window lifecycle;
|
||||||
|
- sidecar process control.
|
||||||
|
|
||||||
|
Do not simulate OS features in React state.
|
||||||
|
|
||||||
|
### Windows audio
|
||||||
|
|
||||||
|
For reliable Windows system-audio capture, create a native Windows capture implementation using WASAPI loopback behind `AudioService`.
|
||||||
|
|
||||||
|
Keep the existing WebView capture path as a demo/fallback until the native path is proven.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 12. Recommended local data model
|
||||||
|
|
||||||
|
Use SQLite for structured state and the filesystem for media.
|
||||||
|
|
||||||
|
Suggested logical tables/entities:
|
||||||
|
|
||||||
|
```text
|
||||||
|
meetings
|
||||||
|
- id
|
||||||
|
- title
|
||||||
|
- started_at
|
||||||
|
- ended_at
|
||||||
|
- duration_ms
|
||||||
|
- status
|
||||||
|
- media_path
|
||||||
|
- media_mime
|
||||||
|
- local_only
|
||||||
|
- sync_status
|
||||||
|
- created_at
|
||||||
|
- updated_at
|
||||||
|
|
||||||
|
speakers
|
||||||
|
- meeting_id
|
||||||
|
- speaker_id
|
||||||
|
- display_name
|
||||||
|
|
||||||
|
transcript_segments
|
||||||
|
- id
|
||||||
|
- meeting_id
|
||||||
|
- speaker_id
|
||||||
|
- start_ms
|
||||||
|
- end_ms
|
||||||
|
- text
|
||||||
|
- is_final
|
||||||
|
- created_at
|
||||||
|
- updated_at
|
||||||
|
|
||||||
|
settings
|
||||||
|
- key
|
||||||
|
- value
|
||||||
|
|
||||||
|
sync_jobs # future
|
||||||
|
- id
|
||||||
|
- meeting_id
|
||||||
|
- type
|
||||||
|
- state
|
||||||
|
- attempts
|
||||||
|
- last_error
|
||||||
|
- updated_at
|
||||||
|
```
|
||||||
|
|
||||||
|
Suggested files:
|
||||||
|
|
||||||
|
```text
|
||||||
|
AppLocalData/
|
||||||
|
└── meetings/
|
||||||
|
└── {meeting_id}/
|
||||||
|
├── recording.webm # demo/current format
|
||||||
|
├── recording.part-* # optional temporary chunk files
|
||||||
|
└── exports/ # future user exports only
|
||||||
|
```
|
||||||
|
|
||||||
|
A JSON transcript may still be generated as an **export/interchange file**, but should not become a second unsynchronized database.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 13. Event and service boundaries
|
||||||
|
|
||||||
|
Prefer typed domain events over components reaching directly into engines.
|
||||||
|
|
||||||
|
Recommended events:
|
||||||
|
|
||||||
|
```text
|
||||||
|
session.started
|
||||||
|
session.paused
|
||||||
|
session.resumed
|
||||||
|
session.finishing
|
||||||
|
session.finished
|
||||||
|
|
||||||
|
capture.started
|
||||||
|
capture.stopped
|
||||||
|
capture.level
|
||||||
|
capture.error
|
||||||
|
|
||||||
|
speech.starting
|
||||||
|
speech.ready
|
||||||
|
speech.partial
|
||||||
|
speech.final
|
||||||
|
speech.error
|
||||||
|
speech.stopped
|
||||||
|
|
||||||
|
speaker.detected
|
||||||
|
speaker.renamed
|
||||||
|
segment.reassigned
|
||||||
|
|
||||||
|
recording.chunk-written
|
||||||
|
recording.finalized
|
||||||
|
recording.error
|
||||||
|
|
||||||
|
sync.queued # future
|
||||||
|
sync.progress # future
|
||||||
|
sync.completed # future
|
||||||
|
sync.failed # future
|
||||||
|
```
|
||||||
|
|
||||||
|
Define event payloads in one shared contract rather than duplicating ad-hoc JSON shapes across the React app and Python process.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 14. Development roadmap — do this in order
|
||||||
|
|
||||||
|
### Phase 0 — protect the current demo
|
||||||
|
|
||||||
|
Before major refactoring:
|
||||||
|
|
||||||
|
1. Add a reproducible smoke-test checklist.
|
||||||
|
2. Add `typecheck`, `lint`, and test scripts.
|
||||||
|
3. Add unit tests for state/domain transformations.
|
||||||
|
4. Add at least one integration test for the speech WebSocket protocol using mock mode.
|
||||||
|
5. Keep the mock demo runnable after every step.
|
||||||
|
|
||||||
|
### Phase 1 — fix persistence and long-session safety
|
||||||
|
|
||||||
|
1. Introduce `LocalMeetingRepository`.
|
||||||
|
2. Make persisted meeting state authoritative.
|
||||||
|
3. Fix finalization ordering so `endedAt`, duration, status, transcript, and media path commit consistently.
|
||||||
|
4. Add startup recovery for unfinished sessions.
|
||||||
|
5. Replace final-Blob-only media persistence with incremental/chunked writing.
|
||||||
|
6. Make pause semantics consistent across recording and transcription.
|
||||||
|
|
||||||
|
This phase has higher priority than adding more UI features.
|
||||||
|
|
||||||
|
### Phase 2 — separate application logic from React
|
||||||
|
|
||||||
|
Refactor `LiveMeeting.tsx` so it does not directly coordinate every subsystem.
|
||||||
|
|
||||||
|
Create services/controllers such as:
|
||||||
|
|
||||||
|
```text
|
||||||
|
MeetingSessionController
|
||||||
|
CaptureService
|
||||||
|
SpeechEngineClient
|
||||||
|
RecordingService
|
||||||
|
OverlayService
|
||||||
|
```
|
||||||
|
|
||||||
|
React should subscribe to state/events and issue commands.
|
||||||
|
|
||||||
|
### Phase 3 — productionize the speech process
|
||||||
|
|
||||||
|
1. Version the speech protocol.
|
||||||
|
2. Wait for explicit `ready` before treating the engine as usable.
|
||||||
|
3. Add health/startup timeout/shutdown behavior.
|
||||||
|
4. Fix long-session timestamp handling.
|
||||||
|
5. Replace repeated large NumPy concatenations with a bounded buffer.
|
||||||
|
6. Benchmark `faster-whisper` + diart on target Windows hardware.
|
||||||
|
7. Package the current Python runtime as a Tauri sidecar.
|
||||||
|
8. Separately benchmark `whisper.cpp`; migrate only if the measured tradeoff is better.
|
||||||
|
|
||||||
|
### Phase 4 — native Windows audio
|
||||||
|
|
||||||
|
1. Define an `AudioService` interface.
|
||||||
|
2. Keep WebView microphone/display capture as a fallback.
|
||||||
|
3. Add native microphone enumeration.
|
||||||
|
4. Add WASAPI loopback system-audio capture.
|
||||||
|
5. Add device-change/error recovery.
|
||||||
|
6. Verify echo/mix behavior when microphone and system audio are both active.
|
||||||
|
|
||||||
|
### Phase 5 — transcript correction features
|
||||||
|
|
||||||
|
Implement product features already anticipated by the original design:
|
||||||
|
|
||||||
|
- edit transcript segment text;
|
||||||
|
- reassign a segment to a different speaker;
|
||||||
|
- rename speakers globally within the meeting;
|
||||||
|
- search transcript;
|
||||||
|
- jump playback to transcript timestamp;
|
||||||
|
- export transcript.
|
||||||
|
|
||||||
|
Keep speaker corrections local unless a future backend contract specifies sync behavior.
|
||||||
|
|
||||||
|
### Phase 6 — finish desktop integration
|
||||||
|
|
||||||
|
Wire the existing settings to real behavior:
|
||||||
|
|
||||||
|
- autostart;
|
||||||
|
- tray behavior;
|
||||||
|
- global hotkey;
|
||||||
|
- theme;
|
||||||
|
- overlay click-through;
|
||||||
|
- overlay lock/position persistence;
|
||||||
|
- proper device selectors.
|
||||||
|
|
||||||
|
### Phase 7 — backend/sync, only when the backend contract exists
|
||||||
|
|
||||||
|
The intended architecture from the supplied design is:
|
||||||
|
|
||||||
|
- structured backend data in PostgreSQL;
|
||||||
|
- large media in SeaweedFS;
|
||||||
|
- desktop asks backend for temporary/presigned upload authorization;
|
||||||
|
- desktop uploads media directly to storage;
|
||||||
|
- transcript/metadata goes to the backend API;
|
||||||
|
- recording/transcription continues offline;
|
||||||
|
- failed uploads/sync are queued locally and retried.
|
||||||
|
|
||||||
|
Do **not** send large media through the application API merely for convenience.
|
||||||
|
|
||||||
|
Do **not** put permanent SeaweedFS credentials in the renderer.
|
||||||
|
|
||||||
|
Until the backend repository/OpenAPI contract is supplied, keep `src/lib/backend.ts` as a clearly isolated placeholder.
|
||||||
|
|
||||||
|
### Phase 8 — packaging and release hardening
|
||||||
|
|
||||||
|
- bundle the speech runtime/models or implement managed model download;
|
||||||
|
- code-sign installers;
|
||||||
|
- enforce restrictive Tauri capabilities/CSP;
|
||||||
|
- handle sidecar startup/crash/update behavior;
|
||||||
|
- validate upgrade/migration paths for SQLite;
|
||||||
|
- add crash-safe session recovery;
|
||||||
|
- test multi-hour meetings;
|
||||||
|
- test CPU-only and GPU-capable Windows machines;
|
||||||
|
- test no-network operation during a meeting.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 15. Rules for changes made by agents
|
||||||
|
|
||||||
|
### Preserve a runnable vertical slice
|
||||||
|
|
||||||
|
Do not leave the repository halfway between architectures. A refactor should keep at least mock-mode desktop operation working at each logical checkpoint.
|
||||||
|
|
||||||
|
### Prefer small interfaces over large rewrites
|
||||||
|
|
||||||
|
Introduce adapters/services around current behavior first, then move implementations behind them.
|
||||||
|
|
||||||
|
Example:
|
||||||
|
|
||||||
|
```text
|
||||||
|
bad: rewrite capture + speech + storage + UI simultaneously
|
||||||
|
|
||||||
|
good:
|
||||||
|
1. define RecordingService
|
||||||
|
2. wrap current MediaRecorder implementation
|
||||||
|
3. add tests/contracts
|
||||||
|
4. replace implementation with chunked/native writer
|
||||||
|
```
|
||||||
|
|
||||||
|
### Do not silently broaden scope
|
||||||
|
|
||||||
|
Do not add login, payments, cloud storage, biometric speaker identity, Kafka, Kubernetes, Elasticsearch, vector databases, or a microservice fleet unless explicitly required.
|
||||||
|
|
||||||
|
### Keep secrets out of the renderer and repository
|
||||||
|
|
||||||
|
Never commit:
|
||||||
|
|
||||||
|
- API keys;
|
||||||
|
- Hugging Face tokens;
|
||||||
|
- storage credentials;
|
||||||
|
- access tokens;
|
||||||
|
- passwords.
|
||||||
|
|
||||||
|
Use environment/secure OS credential storage when such features are actually introduced.
|
||||||
|
|
||||||
|
### Do not rely on checked-in dependency folders
|
||||||
|
|
||||||
|
The supplied archive contains development artifacts such as `node_modules` and a Windows `.venv`, but `.gitignore` correctly excludes them.
|
||||||
|
|
||||||
|
A clean checkout must be reproducible from:
|
||||||
|
|
||||||
|
- `package-lock.json`;
|
||||||
|
- `speech-engine/requirements.txt`;
|
||||||
|
- Rust manifests/lockfile when present.
|
||||||
|
|
||||||
|
Do not design workflows that depend on copying someone else's `.venv` or `node_modules` directory.
|
||||||
|
|
||||||
|
### Keep frontend types strict
|
||||||
|
|
||||||
|
`tsconfig.app.json` has `strict: true`. Do not weaken TypeScript strictness to make an error disappear.
|
||||||
|
|
||||||
|
### Update documentation with behavior
|
||||||
|
|
||||||
|
If a command, prerequisite, port, protocol, storage location, or architecture boundary changes, update the relevant README/this guide in the same change.
|
||||||
|
|
||||||
|
### Do not claim completion without verification
|
||||||
|
|
||||||
|
When reporting a task complete, state which checks actually ran and which could not run.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 16. Suggested testing strategy
|
||||||
|
|
||||||
|
Add a lightweight stack rather than a large test framework rollout.
|
||||||
|
|
||||||
|
### Frontend/domain
|
||||||
|
|
||||||
|
Use unit tests for:
|
||||||
|
|
||||||
|
- meeting creation/finalization;
|
||||||
|
- speaker rename;
|
||||||
|
- transcript append/update;
|
||||||
|
- segment reassignment;
|
||||||
|
- formatting/mapping helpers;
|
||||||
|
- repository adapters.
|
||||||
|
|
||||||
|
### Speech protocol
|
||||||
|
|
||||||
|
Test the WebSocket contract in mock mode:
|
||||||
|
|
||||||
|
- config accepted;
|
||||||
|
- ready event;
|
||||||
|
- PCM input;
|
||||||
|
- final transcript event;
|
||||||
|
- stop/shutdown;
|
||||||
|
- malformed config;
|
||||||
|
- unsupported sample rate.
|
||||||
|
|
||||||
|
### Tauri integration
|
||||||
|
|
||||||
|
Manually/integration-test:
|
||||||
|
|
||||||
|
- AppLocalData permissions;
|
||||||
|
- overlay creation/events;
|
||||||
|
- local media playback;
|
||||||
|
- sidecar start/stop;
|
||||||
|
- global shortcut/autostart when introduced.
|
||||||
|
|
||||||
|
### Long-session tests
|
||||||
|
|
||||||
|
Add synthetic tests that simulate durations greater than 20 minutes so timestamp-reset/drift bugs are caught automatically.
|
||||||
|
|
||||||
|
Also stress recording for multi-hour memory growth before calling the recording path production-ready.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 17. Definition of “good” for the next milestone
|
||||||
|
|
||||||
|
The project is ready to move beyond “demo scaffold” when all of the following are true:
|
||||||
|
|
||||||
|
- A clean checkout installs and starts predictably.
|
||||||
|
- Mock mode is one-command or nearly one-command.
|
||||||
|
- The real speech engine is packaged or automatically managed by the desktop app.
|
||||||
|
- Recording does not accumulate the whole meeting in renderer RAM.
|
||||||
|
- Meeting state survives crashes/restarts from an authoritative local repository.
|
||||||
|
- 60+ minute timestamps remain correct.
|
||||||
|
- Pause/resume behavior is semantically correct.
|
||||||
|
- Windows system audio is reliable enough for the stated product requirement.
|
||||||
|
- Transcript text and speaker assignment can be corrected.
|
||||||
|
- Overlay behavior/settings are fully wired.
|
||||||
|
- Automated tests cover domain state and the mock speech protocol.
|
||||||
|
- Tauri permissions/CSP are hardened for distribution.
|
||||||
|
- No user is required to manually install Python for a packaged release.
|
||||||
|
|
||||||
|
After that milestone, backend/sync integration can be added against a real contract without destabilizing the core meeting-recording experience.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 18. First actions for a new agent
|
||||||
|
|
||||||
|
When starting work in this repository:
|
||||||
|
|
||||||
|
1. Read `PROJECT_SCOPE.md`.
|
||||||
|
2. Read this file.
|
||||||
|
3. Read the relevant section of `README.md`.
|
||||||
|
4. If the task affects architecture, read `docs/MeetVault_Desktop_README_original.md`.
|
||||||
|
5. Inspect the existing implementation before proposing a rewrite.
|
||||||
|
6. Run the current build/check commands that are available.
|
||||||
|
7. Prefer mock-mode smoke testing first.
|
||||||
|
8. Make the smallest coherent change.
|
||||||
|
9. Re-run verification.
|
||||||
|
10. Document any command/architecture change.
|
||||||
|
|
||||||
|
If the requested task conflicts with `PROJECT_SCOPE.md`, do not silently choose one. Treat the user's newest explicit request as authority and update the scope documentation accordingly.
|
||||||
|
|||||||
@@ -1,8 +1,30 @@
|
|||||||
|
# Dependencies (rebuilt via `npm ci` from package-lock.json)
|
||||||
node_modules/
|
node_modules/
|
||||||
|
|
||||||
|
# Build outputs
|
||||||
dist/
|
dist/
|
||||||
src-tauri/target/
|
src-tauri/target/
|
||||||
|
|
||||||
|
# TypeScript incremental build info (`tsc -b`)
|
||||||
|
*.tsbuildinfo
|
||||||
|
|
||||||
|
# Tauri generated files (tauri-build emits schemas/capabilities here on every build)
|
||||||
|
src-tauri/gen/
|
||||||
|
|
||||||
|
# Stale emitted copies of vite.config.ts — the .ts file is the source of truth.
|
||||||
|
# Vite resolves .js before .ts, so an emitted copy would silently shadow it.
|
||||||
|
vite.config.js
|
||||||
|
vite.config.d.ts
|
||||||
|
|
||||||
|
# Python / speech engine (rebuilt via `python -m venv .venv && pip install -r requirements.txt`)
|
||||||
speech-engine/.venv/
|
speech-engine/.venv/
|
||||||
speech-engine/__pycache__/
|
speech-engine/__pycache__/
|
||||||
.env
|
__pycache__/
|
||||||
*.pyc
|
*.pyc
|
||||||
|
|
||||||
|
# Environment & secrets — never commit
|
||||||
|
.env
|
||||||
|
|
||||||
|
# OS / editor noise
|
||||||
.DS_Store
|
.DS_Store
|
||||||
|
Thumbs.db
|
||||||
|
|||||||
154
Wireframes/Desktop/MeetVault-Desktop-Demo/HANDOFF.md
Normal file
154
Wireframes/Desktop/MeetVault-Desktop-Demo/HANDOFF.md
Normal file
@@ -0,0 +1,154 @@
|
|||||||
|
# Phase 0 Handoff — MeetVault Desktop Demo
|
||||||
|
|
||||||
|
Session handoff for the **Phase 0 (test harness + runtime verification)** work defined in `AGENTS.md` §14.
|
||||||
|
Read `PROJECT_SCOPE.md` and `AGENTS.md` first; this file is a continuation note, not a substitute.
|
||||||
|
|
||||||
|
**Delete or replace this file when Phase 0 is complete.**
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 1. Current verification status (verified 2026-10-03)
|
||||||
|
|
||||||
|
| Check | Command | Status | Detail |
|
||||||
|
|---|---|---|---|
|
||||||
|
| Typecheck | `npm run typecheck` | **RED** | 3 errors, all in `tests/speechProtocol.mock.test.ts`: TS2322 line 111 (`ChildProcessByStdio<null, Readable, Readable>` not assignable to `ChildProcessWithoutNullStreams`), TS18048 lines 115/116 (`server` possibly undefined) |
|
||||||
|
| Lint | `npm run lint` | **RED** | 3 errors: `react-hooks/exhaustive-deps` rule definition not found ×2 (`src/pages/LiveMeeting.tsx` lines 51, 70 — pre-existing disable comments; plugin installed but NOT wired into config), and `preserve-caught-error` at test line ~120 (rethrow without `{ cause }`) |
|
||||||
|
| Tests | `npm test` | **RED** | Unit: `src/store/useAppStore.test.ts` **6/6 pass** (with stderr noise — see §4.3). Integration: `tests/speechProtocol.mock.test.ts` **2/5 pass**; failures: close code got 1006 expected 1000; "rejects unsupported sample rates" and "reports malformed config as an error event" both time out waiting for the error event |
|
||||||
|
| Rust | `cargo check --manifest-path src-tauri/Cargo.toml` | **RED** | All deps compile (MSVC toolchain works). Fails only in tauri-build: `` `src-tauri/icons/icon.ico` not found; required for generating a Windows Resource file during tauri-build`` |
|
||||||
|
| Python | `.venv\Scripts\python.exe -m py_compile speech-engine/server.py` | **GREEN** | No changes to server.py planned in Phase 0 |
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2. Completed work this session (done and verified)
|
||||||
|
|
||||||
|
1. **Removed `speech-engine/fresh_venv/`** — it was Python 3.14, incompatible with diart. The working venv is `speech-engine/.venv` (Python 3.12).
|
||||||
|
2. **Fixed pre-existing build breakage**: added `"noEmit": true` to `tsconfig.node.json`. `tsc -b` now works; the composite+noEmit combo was verified.
|
||||||
|
3. **Installed test/lint devDeps** (lockfile updated): vitest 3.2.7, eslint 10.12, @eslint/js ^10, typescript-eslint 8.71, @types/node ^26, **eslint-plugin-react-hooks ^7.1.1**.
|
||||||
|
4. **Created the harness**:
|
||||||
|
- `eslint.config.js` — ESLint 10 flat config: `@eslint/js` recommended + `typescript-eslint` recommended, scoped to `**/*.{ts,tsx}` via `.map()`; ignores dist/node_modules/target/.venv. (react-hooks plugin NOT yet wired in.)
|
||||||
|
- `vitest.config.ts`, `tsconfig.tests.json` (ES2022, no DOM lib, `types: ["node"]`; includes `tests/**/*.ts` + `vitest.config.ts`) referenced from root `tsconfig.json`.
|
||||||
|
- `src/test/setup.ts` — in-memory localStorage stub for zustand persist under Node. **Not working as intended; see §4.3.**
|
||||||
|
- `src/store/useAppStore.test.ts` — 6 unit tests (createMeeting, appendSegment add/replace, renameSpeaker trim/fallback, finishMeeting commit, setSettings merge). All pass.
|
||||||
|
- `tests/speechProtocol.mock.test.ts` — 5 integration tests; spawns `.venv\Scripts\python.exe speech-engine/server.py --mock --port <free-port>` (skips if venv missing), drives it with Node's native WebSocket + fetch.
|
||||||
|
5. **package.json scripts added**: `typecheck` (`tsc -b`), `lint` (`eslint .`), `test` (`vitest run`), `test:watch` (`vitest`).
|
||||||
|
6. **Toolchain unblocked**: installed VS C++ BuildTools workload; MSVC 14.51.36231 at `C:\Program Files (x86)\Microsoft Visual Studio\18\BuildTools`. cargo check compiles all dependencies successfully — only the icon step fails.
|
||||||
|
7. **Verified `.venv` versions** (within requirements pins): fastapi 0.141.1, starlette **1.7.0**, uvicorn 0.54.0, pydantic 2.13.5.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 3. Remaining work — do in this order
|
||||||
|
|
||||||
|
### Step 1 — Create the placeholder icon (unblocks cargo check)
|
||||||
|
|
||||||
|
The archive shipped without `src-tauri/icons/`; git confirms it was never tracked (`git ls-files src-tauri/icons/*` is empty). `tauri.conf.json` has no explicit icon array, so tauri-build requires the default `icons/icon.ico` on Windows.
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
New-Item -ItemType Directory "src-tauri\icons" # Node writeFileSync does NOT create parent dirs
|
||||||
|
node C:\Users\supersanta\AppData\Local\Temp\opencode\make-ico.cjs "D:\workspace\MeetVault\Wireframes\Desktop\MeetVault-Desktop-Demo\src-tauri\icons\icon.ico"
|
||||||
|
```
|
||||||
|
|
||||||
|
`make-ico.cjs` is a one-off generator in the temp dir (writes a 1150-byte solid-blue 16×16 32bpp ICO; takes output path as `argv[2]`). If it's gone, recreate it or hand-write any valid minimal ICO. **Commit the generated `icon.ico`** — the build requires it and it was never tracked.
|
||||||
|
|
||||||
|
Then re-run `cargo check --manifest-path src-tauri/Cargo.toml` → expect GREEN.
|
||||||
|
|
||||||
|
### Step 2 — Settle ground truth for WS teardown before touching test assertions
|
||||||
|
|
||||||
|
Facts verified by reading `speech-engine/server.py` (ws handler, lines ~318–415):
|
||||||
|
|
||||||
|
- **No explicit `websocket.close()` anywhere in the handler.** Both error paths do `send_json({"type":"error", ...})` then `return`; starlette tears down when the handler returns. Observed close code is **1006** (abnormal, no close frame) — consistent with starlette 1.7.0 behavior.
|
||||||
|
- Bad sample rate → error message exactly: `"MeetVault speech engine currently requires 16 kHz mono PCM."` then return.
|
||||||
|
- Malformed JSON hits the generic `except Exception` (~line 391): tries `send_json({"type":"error","message":str(exc)})`, swallows send failures, then `session.close()`.
|
||||||
|
- Yet in tests those error events **never arrive** (client times out). Hypothesis: starlette drops/loses the pre-close send during teardown. Unconfirmed.
|
||||||
|
|
||||||
|
Do one of these to confirm before editing assertions:
|
||||||
|
|
||||||
|
1. Read `.venv\Lib\site-packages\starlette\websockets.py` (line 39+) for what happens on handler return, or
|
||||||
|
2. Manual probe: start `server.py --mock --port <free>`, connect with a small Node WebSocket script, send (a) bad sample_rate config, (b) malformed JSON, (c) valid config + stop after finals; log every frame and the close code for each.
|
||||||
|
|
||||||
|
Then align the three failing assertions to actual behavior. If 1006/no-close-frame is confirmed, accept it in tests **with a comment** documenting that an explicit close/ready handshake is Phase 3 protocol debt (AGENTS.md §9 P1 "speech readiness/protocol"). Do NOT change `server.py` behavior in Phase 0 — tests only.
|
||||||
|
|
||||||
|
### Step 3 — Fix `tests/speechProtocol.mock.test.ts`
|
||||||
|
|
||||||
|
- Line ~6: import `ChildProcessByStdio` from `node:child_process` and `Readable` from `node:stream`.
|
||||||
|
- Line ~104: `let server: ChildProcessByStdio<null, Readable, Readable> | undefined;` — fixes TS2322 (line 111) and both TS18048s (lines 115/116).
|
||||||
|
- Line ~120: attach `{ cause: error }` to the rethrown `Error(...)` in `beforeAll` (ESLint 10 `preserve-caught-error`).
|
||||||
|
- Close-code assertion per Step 2 findings.
|
||||||
|
|
||||||
|
### Step 4 — Wire eslint-plugin-react-hooks into `eslint.config.js`
|
||||||
|
|
||||||
|
Plugin is installed (`^7.1.1`) but absent from the config, so the pre-existing disable comments in `src/pages/LiveMeeting.tsx` (lines 51, 70) fail with "Definition for rule ... was not found". Check the v7 export shape first (`node_modules/eslint-plugin-react-hooks/package.json` + main entry — v7 is ESM), then add a flat-config entry registering the plugin with `rules-of-hooks: error` and `exhaustive-deps: warn` (or error). Keep the disable comments in LiveMeeting.tsx as-is; once the rule exists they become valid.
|
||||||
|
|
||||||
|
### Step 5 — Fix `src/test/setup.ts` localStorage stub
|
||||||
|
|
||||||
|
Current state: gated on `typeof globalThis.localStorage === 'undefined'`. Under vitest workers the gate passes (Node 26's getter returns undefined while emitting an ExperimentalWarning) yet zustand persist still logs `[zustand persist middleware] Unable to update item 'meetvault-demo-store-v1', the given storage is currently unavailable.` per state update. **Root cause undiagnosed.**
|
||||||
|
|
||||||
|
Planned fix: drop the gate — unconditionally `Object.defineProperty(globalThis, 'localStorage', { value: new MemoryStorage(), configurable: true, writable: true })`, then add a functional probe (setItem/getItem roundtrip) so a silent no-op is impossible. Verify success by running `npm test` and confirming BOTH the ExperimentalWarning lines AND all "storage is currently unavailable" messages are gone from output.
|
||||||
|
|
||||||
|
### Step 6 — Clean up stray artifacts
|
||||||
|
|
||||||
|
- **Delete `vite.config.js` and `vite.config.d.ts`.** Verified: `vite.config.js` is a tsc-emitted duplicate of `vite.config.ts` (identical config values, 4-space indent). Vite resolves `.js` before `.ts`, so the stale file can shadow the real one. The real config is `vite.config.ts`; keep it.
|
||||||
|
- **`.gitignore`**: add `*.tsbuildinfo` (3 stray files at repo root: tsconfig.app/node/tests.tsbuildinfo) and `src-tauri/gen/` (tauri-build generated). Keep `src-tauri/Cargo.lock` tracked — this is an application, lockfile should be committed.
|
||||||
|
|
||||||
|
### Step 7 — Full verification (all four must be green)
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
npm run typecheck
|
||||||
|
npm run lint
|
||||||
|
npm test
|
||||||
|
cargo check --manifest-path src-tauri/Cargo.toml
|
||||||
|
```
|
||||||
|
|
||||||
|
Plus the mock-demo smoke path still works: `python speech-engine/server.py --mock` + `npm run desktop` (or at minimum `npm run dev`). AGENTS.md §15: do not claim completion without stating which checks actually ran.
|
||||||
|
|
||||||
|
### Step 8 — Documentation
|
||||||
|
|
||||||
|
- **Create `docs/SMOKE_TEST.md`**: the manual Tauri acceptance walkthrough from AGENTS.md §7 as a checklist (mock mode first; mic permission, levels, transcript segments, speaker rename main↔overlay, pause/resume, finish → Meeting Details playback + transcript, exact placeholder text on backend buttons, restart persistence).
|
||||||
|
- **Update `README.md`**: add the new commands (`typecheck`, `lint`, `test`) and normalize the Python version note to 3.12 (README still says 3.10/3.11 — documented drift per AGENTS.md §6.1; the working venv is 3.12).
|
||||||
|
- **Update `AGENTS.md`** (at workspace root `D:\workspace\MeetVault\AGENTS.md`, tracked in git): §6.6 verification commands now include typecheck/lint/test; repo map gains `tests/`, `src/test/`, `eslint.config.js`, `vitest.config.ts`, `tsconfig.tests.json`; note the placeholder icon situation if relevant.
|
||||||
|
- **Delete this HANDOFF.md** (or replace with a one-line "Phase 0 complete" note).
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 4. Verified ground truth & environment notes
|
||||||
|
|
||||||
|
### 4.1 Mock speech protocol (`speech-engine/server.py`)
|
||||||
|
|
||||||
|
- `/health` → `{ok: true, mode: "mock", sample_rate: 16000}` (real mode reports `"real-capable"`).
|
||||||
|
- WS endpoint `ws://127.0.0.1:<port>/ws/live`; default port 8765; test picks a free port via net server on port 0.
|
||||||
|
- Flow: accept → first text frame is JSON config (`model`, `compute`, `language`, `diarization`, `sample_rate`) → `ready` event. Mock ready payload: `{"type":"ready","engine":"mock-whisper","diarization":"mock-diart"}` (real mode: engine `"Whisper {model}"`).
|
||||||
|
- Mock emits one **final** segment per 38400 samples (2.4 s of 16 kHz int16); speaker cycle `speaker_1 → speaker_2 → speaker_1 → speaker_3`. Test streams 5 s of silence in 3200-byte chunks and expects finals at ~[0.3–2.4] and ~[2.7–4.8].
|
||||||
|
- Client sends `{"type":"stop"}` to end; handler has no explicit close (see Step 2).
|
||||||
|
|
||||||
|
### 4.2 Toolchain / environment
|
||||||
|
|
||||||
|
- Node **v26.10.0**: global `WebSocket` + `fetch` available natively (used by the integration test — no extra deps).
|
||||||
|
- MSVC **14.51.36231** at `C:\Program Files (x86)\Microsoft Visual Studio\18\BuildTools` (VS BuildTools install from this session). cargo check works with it.
|
||||||
|
- `.venv`: Python 3.12, fastapi 0.141.1, starlette 1.7.0, uvicorn 0.54.0, pydantic 2.13.5.
|
||||||
|
- **PowerShell quirks** (this shell is Windows PowerShell 5.x): no `&&` separator — use separate commands or `;`; call operator `& "path"` for paths with spaces; run vitest via `node_modules\.bin\vitest.cmd`; esbuild binary lives at `node_modules\@esbuild\win32-x64\esbuild.exe`, NOT in `.bin`; piping stderr (`*>&1`) produces cosmetic NativeCommandError noise — the command still ran fine.
|
||||||
|
- **ESLint 10**: flat config only; recommended set includes `preserve-caught-error` (rethrows must attach `{ cause }`).
|
||||||
|
|
||||||
|
### 4.3 Known test-output noise (to be fixed in Step 5)
|
||||||
|
|
||||||
|
Unit tests pass but emit, per worker: `(node:...) ExperimentalWarning: localStorage is not available because --localstorage-file was not provided.` and one `[zustand persist middleware] Unable to update item 'meetvault-demo-store-v1', the given storage is currently unavailable.` line per state update. Both should disappear once the setup.ts stub actually takes effect.
|
||||||
|
|
||||||
|
### 4.4 Working-tree changes that predate or accompany this session (keep, don't revert)
|
||||||
|
|
||||||
|
- `src-tauri/Cargo.toml`: tauri feature list changed to `["protocol-asset"]` (from earlier work; needed for asset protocol).
|
||||||
|
- Workspace-root `AGENTS.md` (`D:\workspace\MeetVault\AGENTS.md`): rewritten agent guide (82 → 944 lines); it is the current operating guide.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 5. Constraints (from AGENTS.md / PROJECT_SCOPE.md — do not violate)
|
||||||
|
|
||||||
|
- Keep mock-mode desktop operation working at every checkpoint; no half-finished architecture states.
|
||||||
|
- Placeholder text must remain exactly: `Processing (The job). Please wait a moment`
|
||||||
|
- `tsconfig.app.json` has `strict: true` — never weaken TypeScript strictness to silence an error.
|
||||||
|
- No new **runtime** dependencies (test devDeps are fine and already installed).
|
||||||
|
- Do not invent backend endpoints, credentials, auth behavior, or billing rules; keep `src/lib/backend.ts` a placeholder.
|
||||||
|
- Do not edit generated/dependency dirs: `node_modules/`, `speech-engine/.venv/`, `dist/`, `src-tauri/target/`.
|
||||||
|
- **No commits** — all changes stay uncommitted unless the user asks otherwise.
|
||||||
|
|
||||||
|
## 6. Git snapshot (as of handoff, branch `main`)
|
||||||
|
|
||||||
|
Modified: `AGENTS.md` (workspace root), `package.json`, `package-lock.json`, `src-tauri/Cargo.toml`, `tsconfig.json` (+tests project reference), `tsconfig.node.json` (+noEmit)
|
||||||
|
Untracked: `eslint.config.js`, `vitest.config.ts`, `tsconfig.tests.json`, `src/test/`, `tests/`, `src/store/useAppStore.test.ts`, `HANDOFF.md` (this file), stray `vite.config.js` + `vite.config.d.ts` (delete in Step 6), `*.tsbuildinfo` ×3, `src-tauri/Cargo.lock` (keep tracked), `src-tauri/gen/` (gitignore)
|
||||||
11
Wireframes/Desktop/MeetVault-Desktop-Demo/eslint.config.js
Normal file
11
Wireframes/Desktop/MeetVault-Desktop-Demo/eslint.config.js
Normal file
@@ -0,0 +1,11 @@
|
|||||||
|
import js from '@eslint/js';
|
||||||
|
import tseslint from 'typescript-eslint';
|
||||||
|
|
||||||
|
const tsFiles = ['**/*.{ts,tsx}'];
|
||||||
|
|
||||||
|
export default [
|
||||||
|
{ ignores: ['dist/', 'node_modules/', 'src-tauri/target/', 'speech-engine/.venv/'] },
|
||||||
|
{ files: tsFiles, ...js.configs.recommended },
|
||||||
|
// typescript-eslint entries come last so their overrides of core rules win.
|
||||||
|
...tseslint.configs.recommended.map((entry) => ({ files: tsFiles, ...entry })),
|
||||||
|
];
|
||||||
1759
Wireframes/Desktop/MeetVault-Desktop-Demo/package-lock.json
generated
1759
Wireframes/Desktop/MeetVault-Desktop-Demo/package-lock.json
generated
File diff suppressed because it is too large
Load Diff
@@ -10,7 +10,11 @@
|
|||||||
"tauri": "tauri",
|
"tauri": "tauri",
|
||||||
"desktop": "tauri dev",
|
"desktop": "tauri dev",
|
||||||
"speech": "python speech-engine/server.py",
|
"speech": "python speech-engine/server.py",
|
||||||
"speech:mock": "python speech-engine/server.py --mock"
|
"speech:mock": "python speech-engine/server.py --mock",
|
||||||
|
"typecheck": "tsc -b",
|
||||||
|
"lint": "eslint .",
|
||||||
|
"test": "vitest run",
|
||||||
|
"test:watch": "vitest"
|
||||||
},
|
},
|
||||||
"dependencies": {
|
"dependencies": {
|
||||||
"@tauri-apps/api": "^2.0.0",
|
"@tauri-apps/api": "^2.0.0",
|
||||||
@@ -23,11 +27,17 @@
|
|||||||
"zustand": "^5.0.2"
|
"zustand": "^5.0.2"
|
||||||
},
|
},
|
||||||
"devDependencies": {
|
"devDependencies": {
|
||||||
|
"@eslint/js": "^10.0.1",
|
||||||
"@tauri-apps/cli": "^2.0.0",
|
"@tauri-apps/cli": "^2.0.0",
|
||||||
|
"@types/node": "^26.6.4",
|
||||||
"@types/react": "^18.3.18",
|
"@types/react": "^18.3.18",
|
||||||
"@types/react-dom": "^18.3.5",
|
"@types/react-dom": "^18.3.5",
|
||||||
"@vitejs/plugin-react": "^4.3.4",
|
"@vitejs/plugin-react": "^4.3.4",
|
||||||
|
"eslint": "^10.12.0",
|
||||||
|
"eslint-plugin-react-hooks": "^7.1.1",
|
||||||
"typescript": "~5.6.2",
|
"typescript": "~5.6.2",
|
||||||
"vite": "^6.0.5"
|
"typescript-eslint": "^8.71.0",
|
||||||
|
"vite": "^6.0.5",
|
||||||
|
"vitest": "^3.2.7"
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|||||||
4548
Wireframes/Desktop/MeetVault-Desktop-Demo/src-tauri/Cargo.lock
generated
Normal file
4548
Wireframes/Desktop/MeetVault-Desktop-Demo/src-tauri/Cargo.lock
generated
Normal file
File diff suppressed because it is too large
Load Diff
@@ -13,7 +13,7 @@ crate-type = ["staticlib", "cdylib", "rlib"]
|
|||||||
tauri-build = { version = "2", features = [] }
|
tauri-build = { version = "2", features = [] }
|
||||||
|
|
||||||
[dependencies]
|
[dependencies]
|
||||||
tauri = { version = "2", features = [] }
|
tauri = { version = "2", features = ["protocol-asset"] }
|
||||||
tauri-plugin-fs = "2"
|
tauri-plugin-fs = "2"
|
||||||
tauri-plugin-dialog = "2"
|
tauri-plugin-dialog = "2"
|
||||||
serde = { version = "1", features = ["derive"] }
|
serde = { version = "1", features = ["derive"] }
|
||||||
|
|||||||
@@ -0,0 +1,110 @@
|
|||||||
|
import { describe, expect, it } from 'vitest';
|
||||||
|
import type { TranscriptSegment } from '../types';
|
||||||
|
import { useAppStore } from './useAppStore';
|
||||||
|
|
||||||
|
function makeSegment(overrides: Partial<TranscriptSegment> = {}): TranscriptSegment {
|
||||||
|
return {
|
||||||
|
id: crypto.randomUUID(),
|
||||||
|
speakerId: 'speaker_1',
|
||||||
|
start: 0,
|
||||||
|
end: 2.4,
|
||||||
|
text: 'Hello world.',
|
||||||
|
final: true,
|
||||||
|
...overrides,
|
||||||
|
};
|
||||||
|
}
|
||||||
|
|
||||||
|
function getMeeting(id: string) {
|
||||||
|
const meeting = useAppStore.getState().meetings.find((m) => m.id === id);
|
||||||
|
if (!meeting) throw new Error(`meeting ${id} not found`);
|
||||||
|
return meeting;
|
||||||
|
}
|
||||||
|
|
||||||
|
describe('useAppStore', () => {
|
||||||
|
it('createMeeting prepends a recording meeting and activates it', () => {
|
||||||
|
const before = useAppStore.getState().meetings.length;
|
||||||
|
const id = useAppStore.getState().createMeeting('Test Meeting');
|
||||||
|
|
||||||
|
const state = useAppStore.getState();
|
||||||
|
expect(id).toBeTruthy();
|
||||||
|
expect(state.activeMeetingId).toBe(id);
|
||||||
|
expect(state.meetings.length).toBe(before + 1);
|
||||||
|
expect(state.meetings[0].id).toBe(id);
|
||||||
|
|
||||||
|
const meeting = getMeeting(id);
|
||||||
|
expect(meeting.title).toBe('Test Meeting');
|
||||||
|
expect(meeting.status).toBe('recording');
|
||||||
|
expect(meeting.durationSeconds).toBe(0);
|
||||||
|
expect(meeting.transcript).toEqual([]);
|
||||||
|
expect(meeting.speakerNames).toEqual({});
|
||||||
|
expect(meeting.localOnly).toBe(true);
|
||||||
|
expect(Number.isNaN(Date.parse(meeting.startedAt))).toBe(false);
|
||||||
|
});
|
||||||
|
|
||||||
|
it('appendSegment adds segments to the right meeting only', () => {
|
||||||
|
const id = useAppStore.getState().createMeeting();
|
||||||
|
const otherId = useAppStore.getState().createMeeting();
|
||||||
|
const segment = makeSegment({ text: 'First line.' });
|
||||||
|
|
||||||
|
useAppStore.getState().appendSegment(id, segment);
|
||||||
|
|
||||||
|
expect(getMeeting(id).transcript).toEqual([segment]);
|
||||||
|
expect(getMeeting(otherId).transcript).toEqual([]);
|
||||||
|
});
|
||||||
|
|
||||||
|
it('appendSegment replaces an existing segment with the same id', () => {
|
||||||
|
const id = useAppStore.getState().createMeeting();
|
||||||
|
const original = makeSegment({ text: 'partial text' });
|
||||||
|
useAppStore.getState().appendSegment(id, original);
|
||||||
|
|
||||||
|
const updated = makeSegment({ ...original, text: 'final text', final: true });
|
||||||
|
useAppStore.getState().appendSegment(id, updated);
|
||||||
|
|
||||||
|
expect(getMeeting(id).transcript).toEqual([updated]);
|
||||||
|
});
|
||||||
|
|
||||||
|
it('renameSpeaker stores the trimmed name and falls back to the speaker id for blank names', () => {
|
||||||
|
const id = useAppStore.getState().createMeeting();
|
||||||
|
|
||||||
|
useAppStore.getState().renameSpeaker(id, 'speaker_1', ' Alice ');
|
||||||
|
expect(getMeeting(id).speakerNames.speaker_1).toBe('Alice');
|
||||||
|
|
||||||
|
useAppStore.getState().renameSpeaker(id, 'speaker_2', ' ');
|
||||||
|
expect(getMeeting(id).speakerNames.speaker_2).toBe('speaker_2');
|
||||||
|
});
|
||||||
|
|
||||||
|
it('finishMeeting commits endedAt, duration, processing status and media info', () => {
|
||||||
|
const id = useAppStore.getState().createMeeting();
|
||||||
|
const beforeFinish = Date.now();
|
||||||
|
|
||||||
|
useAppStore.getState().finishMeeting(id, 'meetings/test/recording.webm', 'video/webm');
|
||||||
|
|
||||||
|
const state = useAppStore.getState();
|
||||||
|
const meeting = getMeeting(id);
|
||||||
|
expect(meeting.status).toBe('processing');
|
||||||
|
expect(meeting.endedAt).toBeTruthy();
|
||||||
|
expect(Number.isNaN(Date.parse(meeting.endedAt!))).toBe(false);
|
||||||
|
// A just-created meeting has ~0 elapsed time; the store clamps to at least 1 second.
|
||||||
|
expect(meeting.durationSeconds).toBeGreaterThanOrEqual(1);
|
||||||
|
expect(Date.parse(meeting.endedAt!) - beforeFinish).toBeLessThan(5_000);
|
||||||
|
expect(meeting.mediaPath).toBe('meetings/test/recording.webm');
|
||||||
|
expect(meeting.mediaMime).toBe('video/webm');
|
||||||
|
expect(state.activeMeetingId).toBeUndefined();
|
||||||
|
expect(state.processingMessage).toBe('Processing (The job). Please wait a moment');
|
||||||
|
});
|
||||||
|
|
||||||
|
it('setSettings merges the patch without clobbering other keys', () => {
|
||||||
|
const before = useAppStore.getState().settings;
|
||||||
|
|
||||||
|
useAppStore.getState().setSettings({ theme: 'dark', overlayLines: 6 });
|
||||||
|
|
||||||
|
const settings = useAppStore.getState().settings;
|
||||||
|
expect(settings.theme).toBe('dark');
|
||||||
|
expect(settings.overlayLines).toBe(6);
|
||||||
|
for (const key of Object.keys(before) as Array<keyof typeof before>) {
|
||||||
|
if (key !== 'theme' && key !== 'overlayLines') {
|
||||||
|
expect(settings[key]).toBe(before[key]);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
});
|
||||||
|
});
|
||||||
38
Wireframes/Desktop/MeetVault-Desktop-Demo/src/test/setup.ts
Normal file
38
Wireframes/Desktop/MeetVault-Desktop-Demo/src/test/setup.ts
Normal file
@@ -0,0 +1,38 @@
|
|||||||
|
// Vitest setup: provide a minimal in-memory localStorage so zustand/persist
|
||||||
|
// works under the Node test environment (no DOM available).
|
||||||
|
|
||||||
|
class MemoryStorage implements Storage {
|
||||||
|
private store = new Map<string, string>();
|
||||||
|
|
||||||
|
get length(): number {
|
||||||
|
return this.store.size;
|
||||||
|
}
|
||||||
|
|
||||||
|
clear(): void {
|
||||||
|
this.store.clear();
|
||||||
|
}
|
||||||
|
|
||||||
|
getItem(key: string): string | null {
|
||||||
|
return this.store.has(key) ? (this.store.get(key) as string) : null;
|
||||||
|
}
|
||||||
|
|
||||||
|
key(index: number): string | null {
|
||||||
|
return Array.from(this.store.keys())[index] ?? null;
|
||||||
|
}
|
||||||
|
|
||||||
|
setItem(key: string, value: string): void {
|
||||||
|
this.store.set(key, String(value));
|
||||||
|
}
|
||||||
|
|
||||||
|
removeItem(key: string): void {
|
||||||
|
this.store.delete(key);
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
if (typeof globalThis.localStorage === 'undefined') {
|
||||||
|
Object.defineProperty(globalThis, 'localStorage', {
|
||||||
|
value: new MemoryStorage(),
|
||||||
|
configurable: true,
|
||||||
|
writable: true,
|
||||||
|
});
|
||||||
|
}
|
||||||
@@ -0,0 +1,223 @@
|
|||||||
|
// Integration test for the mock speech-engine WebSocket protocol.
|
||||||
|
// Spawns `speech-engine/server.py --mock` from the local .venv on a free port
|
||||||
|
// and drives it with Node's native WebSocket client. Skipped when the venv is missing.
|
||||||
|
|
||||||
|
import { spawn } from 'node:child_process';
|
||||||
|
import type { ChildProcessWithoutNullStreams } from 'node:child_process';
|
||||||
|
import fs from 'node:fs';
|
||||||
|
import net from 'node:net';
|
||||||
|
import path from 'node:path';
|
||||||
|
import { fileURLToPath } from 'node:url';
|
||||||
|
import { afterAll, beforeAll, describe, expect, it } from 'vitest';
|
||||||
|
|
||||||
|
const repoRoot = path.resolve(path.dirname(fileURLToPath(import.meta.url)), '..');
|
||||||
|
|
||||||
|
interface SpeechEvent {
|
||||||
|
type: string;
|
||||||
|
[key: string]: unknown;
|
||||||
|
}
|
||||||
|
|
||||||
|
interface SegmentPayload {
|
||||||
|
id: string;
|
||||||
|
speakerId: string;
|
||||||
|
start: number;
|
||||||
|
end: number;
|
||||||
|
text: string;
|
||||||
|
final: boolean;
|
||||||
|
}
|
||||||
|
|
||||||
|
function venvPython(): string | null {
|
||||||
|
const candidates =
|
||||||
|
process.platform === 'win32'
|
||||||
|
? [path.join(repoRoot, 'speech-engine', '.venv', 'Scripts', 'python.exe')]
|
||||||
|
: [path.join(repoRoot, 'speech-engine', '.venv', 'bin', 'python')];
|
||||||
|
return candidates.find((candidate) => fs.existsSync(candidate)) ?? null;
|
||||||
|
}
|
||||||
|
|
||||||
|
function findFreePort(): Promise<number> {
|
||||||
|
return new Promise((resolve, reject) => {
|
||||||
|
const server = net.createServer();
|
||||||
|
server.unref();
|
||||||
|
server.on('error', reject);
|
||||||
|
server.listen(0, '127.0.0.1', () => {
|
||||||
|
const address = server.address();
|
||||||
|
if (!address || typeof address === 'string') {
|
||||||
|
reject(new Error('Could not determine a free port'));
|
||||||
|
return;
|
||||||
|
}
|
||||||
|
server.close(() => resolve(address.port));
|
||||||
|
});
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
async function waitForHealth(url: string, timeoutMs: number): Promise<void> {
|
||||||
|
const deadline = Date.now() + timeoutMs;
|
||||||
|
for (;;) {
|
||||||
|
try {
|
||||||
|
const response = await fetch(url);
|
||||||
|
if (response.ok) return;
|
||||||
|
} catch {
|
||||||
|
// Server not up yet.
|
||||||
|
}
|
||||||
|
if (Date.now() > deadline) throw new Error(`Speech engine did not become healthy at ${url}`);
|
||||||
|
await new Promise((resolve) => setTimeout(resolve, 250));
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
function openSocket(url: string): Promise<WebSocket> {
|
||||||
|
return new Promise((resolve, reject) => {
|
||||||
|
const socket = new WebSocket(url);
|
||||||
|
socket.onopen = () => resolve(socket);
|
||||||
|
socket.onerror = (event) => reject(new Error(`WebSocket error: ${String(event.message ?? event.type)}`));
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
function nextEvent(socket: WebSocket, timeoutMs: number): Promise<SpeechEvent> {
|
||||||
|
return new Promise((resolve, reject) => {
|
||||||
|
const timer = setTimeout(() => reject(new Error('Timed out waiting for a speech engine event')), timeoutMs);
|
||||||
|
socket.onmessage = (event) => {
|
||||||
|
clearTimeout(timer);
|
||||||
|
try {
|
||||||
|
resolve(JSON.parse(String(event.data)) as SpeechEvent);
|
||||||
|
} catch (error) {
|
||||||
|
reject(error instanceof Error ? error : new Error('Non-JSON message from speech engine'));
|
||||||
|
}
|
||||||
|
};
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
function waitForClose(socket: WebSocket, timeoutMs: number): Promise<number> {
|
||||||
|
return new Promise((resolve, reject) => {
|
||||||
|
const timer = setTimeout(() => reject(new Error('Timed out waiting for the socket to close')), timeoutMs);
|
||||||
|
socket.onclose = (event) => {
|
||||||
|
clearTimeout(timer);
|
||||||
|
resolve(event.code);
|
||||||
|
};
|
||||||
|
});
|
||||||
|
}
|
||||||
|
|
||||||
|
const python = venvPython();
|
||||||
|
|
||||||
|
describe.skipIf(python === null)(
|
||||||
|
'mock speech protocol (server.py --mock)',
|
||||||
|
() => {
|
||||||
|
let server: ChildProcessWithoutNullStreams | undefined;
|
||||||
|
let port = 0;
|
||||||
|
const stderr: string[] = [];
|
||||||
|
|
||||||
|
beforeAll(async () => {
|
||||||
|
if (!python) return;
|
||||||
|
port = await findFreePort();
|
||||||
|
server = spawn(python, [path.join(repoRoot, 'speech-engine', 'server.py'), '--mock', '--port', String(port)], {
|
||||||
|
cwd: repoRoot,
|
||||||
|
stdio: ['ignore', 'pipe', 'pipe'],
|
||||||
|
});
|
||||||
|
server.stdout.on('data', () => undefined);
|
||||||
|
server.stderr.on('data', (chunk) => stderr.push(String(chunk)));
|
||||||
|
try {
|
||||||
|
await waitForHealth(`http://127.0.0.1:${port}/health`, 30_000);
|
||||||
|
} catch (error) {
|
||||||
|
throw new Error(
|
||||||
|
`Speech engine did not start: ${error instanceof Error ? error.message : String(error)}\n--- stderr ---\n${stderr.join('')}`,
|
||||||
|
);
|
||||||
|
}
|
||||||
|
}, 60_000);
|
||||||
|
|
||||||
|
afterAll(() => {
|
||||||
|
server?.kill();
|
||||||
|
});
|
||||||
|
|
||||||
|
const validConfig = JSON.stringify({
|
||||||
|
model: 'small',
|
||||||
|
compute: 'auto',
|
||||||
|
language: 'auto',
|
||||||
|
diarization: true,
|
||||||
|
sample_rate: 16000,
|
||||||
|
});
|
||||||
|
|
||||||
|
it('reports mock mode on /health', async () => {
|
||||||
|
const response = await fetch(`http://127.0.0.1:${port}/health`);
|
||||||
|
expect(response.ok).toBe(true);
|
||||||
|
const body = (await response.json()) as { ok: boolean; mode: string; sample_rate: number };
|
||||||
|
expect(body.ok).toBe(true);
|
||||||
|
expect(body.mode).toBe('mock');
|
||||||
|
expect(body.sample_rate).toBe(16000);
|
||||||
|
});
|
||||||
|
|
||||||
|
it('completes the ready handshake for a valid config', async () => {
|
||||||
|
const socket = await openSocket(`ws://127.0.0.1:${port}/ws/live`);
|
||||||
|
socket.send(validConfig);
|
||||||
|
|
||||||
|
const ready = await nextEvent(socket, 10_000);
|
||||||
|
expect(ready.type).toBe('ready');
|
||||||
|
expect(String(ready.engine)).toContain('mock');
|
||||||
|
socket.close();
|
||||||
|
});
|
||||||
|
|
||||||
|
it('emits final segments with cycling speakers for streamed PCM', async () => {
|
||||||
|
const socket = await openSocket(`ws://127.0.0.1:${port}/ws/live`);
|
||||||
|
socket.send(validConfig);
|
||||||
|
expect((await nextEvent(socket, 10_000)).type).toBe('ready');
|
||||||
|
|
||||||
|
// Stream 5 seconds of silence (16 kHz mono int16) in 100 ms chunks; the mock
|
||||||
|
// emits one final segment per 2.4 s of samples with cycling speaker labels.
|
||||||
|
const totalBytes = 16000 * 5 * 2;
|
||||||
|
const chunkSize = 3200;
|
||||||
|
for (let offset = 0; offset < totalBytes; offset += chunkSize) {
|
||||||
|
socket.send(new ArrayBuffer(chunkSize));
|
||||||
|
}
|
||||||
|
|
||||||
|
const finals: SpeechEvent[] = [];
|
||||||
|
while (finals.length < 2) {
|
||||||
|
const event = await nextEvent(socket, 15_000);
|
||||||
|
if (event.type === 'error') throw new Error(String(event.message));
|
||||||
|
if (event.type === 'final') finals.push(event);
|
||||||
|
}
|
||||||
|
|
||||||
|
expect(finals).toHaveLength(2);
|
||||||
|
const first = finals[0].segment as SegmentPayload;
|
||||||
|
const second = finals[1].segment as SegmentPayload;
|
||||||
|
|
||||||
|
for (const segment of [first, second]) {
|
||||||
|
expect(segment.id).toBeTruthy();
|
||||||
|
expect(segment.text.length).toBeGreaterThan(0);
|
||||||
|
expect(segment.final).toBe(true);
|
||||||
|
}
|
||||||
|
expect(first.speakerId).toBe('speaker_1');
|
||||||
|
expect(second.speakerId).toBe('speaker_2');
|
||||||
|
expect(first.start).toBeCloseTo(0.3, 1);
|
||||||
|
expect(first.end).toBeCloseTo(2.4, 1);
|
||||||
|
expect(second.start).toBeCloseTo(2.7, 1);
|
||||||
|
expect(second.end).toBeCloseTo(4.8, 1);
|
||||||
|
expect(second.start).toBeGreaterThan(first.end);
|
||||||
|
|
||||||
|
socket.send(JSON.stringify({ type: 'stop' }));
|
||||||
|
const code = await waitForClose(socket, 5_000);
|
||||||
|
expect(code).toBe(1000);
|
||||||
|
});
|
||||||
|
|
||||||
|
it('rejects unsupported sample rates', async () => {
|
||||||
|
const socket = await openSocket(`ws://127.0.0.1:${port}/ws/live`);
|
||||||
|
socket.send(
|
||||||
|
JSON.stringify({ model: 'small', compute: 'auto', language: 'auto', diarization: true, sample_rate: 44100 }),
|
||||||
|
);
|
||||||
|
|
||||||
|
const event = await nextEvent(socket, 10_000);
|
||||||
|
expect(event.type).toBe('error');
|
||||||
|
expect(String(event.message)).toContain('16 kHz');
|
||||||
|
|
||||||
|
await waitForClose(socket, 5_000);
|
||||||
|
});
|
||||||
|
|
||||||
|
it('reports malformed config as an error event', async () => {
|
||||||
|
const socket = await openSocket(`ws://127.0.0.1:${port}/ws/live`);
|
||||||
|
socket.send('this is not json');
|
||||||
|
|
||||||
|
const event = await nextEvent(socket, 10_000);
|
||||||
|
expect(event.type).toBe('error');
|
||||||
|
expect(String(event.message)).not.toBe('');
|
||||||
|
|
||||||
|
await waitForClose(socket, 5_000);
|
||||||
|
});
|
||||||
|
},
|
||||||
|
);
|
||||||
@@ -2,6 +2,7 @@
|
|||||||
"files": [],
|
"files": [],
|
||||||
"references": [
|
"references": [
|
||||||
{ "path": "./tsconfig.app.json" },
|
{ "path": "./tsconfig.app.json" },
|
||||||
{ "path": "./tsconfig.node.json" }
|
{ "path": "./tsconfig.node.json" },
|
||||||
|
{ "path": "./tsconfig.tests.json" }
|
||||||
]
|
]
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -4,7 +4,8 @@
|
|||||||
"skipLibCheck": true,
|
"skipLibCheck": true,
|
||||||
"module": "ESNext",
|
"module": "ESNext",
|
||||||
"moduleResolution": "Bundler",
|
"moduleResolution": "Bundler",
|
||||||
"allowImportingTsExtensions": true
|
"allowImportingTsExtensions": true,
|
||||||
|
"noEmit": true
|
||||||
},
|
},
|
||||||
"include": ["vite.config.ts"]
|
"include": ["vite.config.ts"]
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -0,0 +1,15 @@
|
|||||||
|
{
|
||||||
|
"compilerOptions": {
|
||||||
|
"target": "ES2022",
|
||||||
|
"lib": ["ES2022"],
|
||||||
|
"module": "ESNext",
|
||||||
|
"moduleResolution": "Bundler",
|
||||||
|
"strict": true,
|
||||||
|
"noEmit": true,
|
||||||
|
"skipLibCheck": true,
|
||||||
|
"esModuleInterop": true,
|
||||||
|
"isolatedModules": true,
|
||||||
|
"types": ["node"]
|
||||||
|
},
|
||||||
|
"include": ["tests/**/*.ts", "vitest.config.ts"]
|
||||||
|
}
|
||||||
11
Wireframes/Desktop/MeetVault-Desktop-Demo/vitest.config.ts
Normal file
11
Wireframes/Desktop/MeetVault-Desktop-Demo/vitest.config.ts
Normal file
@@ -0,0 +1,11 @@
|
|||||||
|
import { defineConfig } from 'vitest/config';
|
||||||
|
|
||||||
|
export default defineConfig({
|
||||||
|
test: {
|
||||||
|
environment: 'node',
|
||||||
|
setupFiles: ['./src/test/setup.ts'],
|
||||||
|
include: ['src/**/*.test.{ts,tsx}', 'tests/**/*.test.{ts,tsx}'],
|
||||||
|
testTimeout: 30_000,
|
||||||
|
hookTimeout: 60_000,
|
||||||
|
},
|
||||||
|
});
|
||||||
Reference in New Issue
Block a user