157 lines
8.8 KiB
Markdown
157 lines
8.8 KiB
Markdown
# Nat20 Notes
|
||
|
||
Turn a recorded tabletop RPG session (audio or video) into two documents:
|
||
a full GM/DM session log, and a spoiler-free player recap — using local
|
||
transcription (WhisperX) and either a local LLM (Ollama) or any
|
||
OpenAI-compatible hosted API for summarization.
|
||
|
||
## Requirements
|
||
|
||
- Docker + Docker Compose
|
||
- An NVIDIA GPU with drivers + [NVIDIA Container Toolkit](https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html) installed on the host (for transcription; CPU-only works but is much slower)
|
||
- A free [HuggingFace token](https://huggingface.co/settings/tokens) — needed for speaker diarization. You'll also need to accept the terms on the gated pyannote model page it links you to on first run.
|
||
- Either: [Ollama](https://ollama.com) running somewhere reachable from this app (local or LAN), **or** an API key for a hosted LLM (OpenAI, or any OpenAI-compatible provider)
|
||
|
||
## Quick start
|
||
|
||
```bash
|
||
git clone <this-repo>
|
||
cd nat20-notes
|
||
# Review docker-compose.yml — see Configuration below for env vars
|
||
docker compose up -d --build
|
||
```
|
||
|
||
Open `http://<your-server>:8020` (d20 nod) and **log in** with the default
|
||
admin credentials (`admin` / `admin`, overridable via `NAT20_ADMIN_USERNAME`
|
||
and `NAT20_ADMIN_PASSWORD`). The first user is created automatically on
|
||
first startup — no separate registration step.
|
||
|
||
Then follow the setup wizard:
|
||
|
||
1. Choose a Whisper model size based on your GPU's available VRAM (guidance shown in-app)
|
||
2. Give your campaign a name
|
||
3. Choose local Ollama or a hosted API for summarization, and paste your HF token
|
||
4. Optionally paste campaign/world context (NPC names, places) so summaries recognise them correctly
|
||
|
||
A reference compose file using named volumes only (no host paths) is at
|
||
[`docker-compose.example.yml`](./docker-compose.example.yml).
|
||
|
||
## Using it
|
||
|
||
1. **Log in** with your username and password. Admins can create additional users at `/users`.
|
||
2. **Select a campaign** from the sidebar (or use the default). Each campaign has its own settings, sessions, and player recaps.
|
||
3. **Upload** a recording (audio or video — video is auto-converted, audio-only files skip that step and are much faster). Files over **90 MB** are auto-chunked with 3-way concurrent uploads.
|
||
4. **Transcribe** — runs in the background. On completion you're **automatically taken** to the speaker-naming screen.
|
||
5. **Name your speakers** — a waveform-style "session reel" shows each detected speaker's segments; click any point to jump the audio there and hear who's talking, then type in their name. Use the checkboxes to **merge** speakers (e.g. when diarization over-splits one person into `SPEAKER_00`, `SPEAKER_05`, etc.).
|
||
6. **Generate notes** — click "Done naming" and confirm; notes generation starts automatically and you're taken to the job progress screen. When finished, navigate to the notes viewer.
|
||
7. **Review & tweak** — regenerate notes, rename speakers, or delete jobs/sessions from the session detail page.
|
||
|
||
## Data layout
|
||
|
||
The app stores everything under `/data` (inside the container):
|
||
|
||
| Directory / File | Contents |
|
||
|---|---|
|
||
| `campaigns/{id}/audio/` | Uploaded recordings and extracted audio, organised per campaign |
|
||
| `campaigns/{id}/transcriptions/` | Per-session transcript JSON files |
|
||
| `campaigns/{id}/notes/` | Generated notes (GM log + player recap) |
|
||
| `app.db` | SQLite database (users, campaigns, campaign_shares, sessions, speakers, campaign/global settings, jobs) |
|
||
|
||
## Configuration
|
||
|
||
### Environment variables (`NAT20_*`)
|
||
|
||
Set these under the `backend` service in `docker-compose.yml` to prefill
|
||
the setup wizard and override defaults. All are optional — the wizard
|
||
and Settings page can set them at runtime.
|
||
|
||
#### Transcription
|
||
|
||
| Variable | What it does | Values | Default | Notes |
|
||
|---|---|---|---|---|
|
||
| `NAT20_WHISPER_MODEL` | Whisper model size | `tiny` `base` `small` `medium` `large-v3` | `medium` | Larger = more accurate but uses more VRAM. `medium` fits most 6-8 GB GPUs; `large-v3` needs ~10 GB+ |
|
||
| `NAT20_WHISPER_COMPUTE_TYPE` | Compute precision | `int8` `float16` `float32` | `int8` | `int8` = fastest / least VRAM. `float16` = more accurate, more VRAM. `float32` = full precision, slowest |
|
||
| `NAT20_HF_TOKEN` | HuggingFace token for speaker diarization | `hf_...` | (none) | Required for speaker attribution. Must accept pyannote gated-model terms with the same account first |
|
||
|
||
#### LLM backend
|
||
|
||
| Variable | What it does | Values | Default | Notes |
|
||
|---|---|---|---|---|
|
||
| `NAT20_OLLAMA_HOST` | Ollama server URL | URL | `http://localhost:11434` | Set to your Ollama host |
|
||
| `NAT20_OLLAMA_MODEL` | Ollama model name | any model on your server | `qwen2.5:7b` | |
|
||
| `NAT20_API_BASE_URL` | OpenAI-compatible API base | URL | `https://api.openai.com/v1` | Uncomment and set to switch from Ollama |
|
||
| `NAT20_API_KEY` | API key | string | (none) | |
|
||
| `NAT20_API_MODEL` | API model name | string | `gpt-4o-mini` | |
|
||
|
||
#### Summarization
|
||
|
||
| Variable | What it does | Values | Default | Notes |
|
||
|---|---|---|---|---|
|
||
| `NAT20_CHUNK_WORD_TARGET` | Target words per summarisation chunk | number | `2500` | Long transcripts are split into chunks, each summarised separately, then combined. **Lower** = more LLM calls but finer granularity (good for models with small context windows). **Higher** = more context per chunk but may exceed the model's window. `2500` is safe for most models (8K–128K context) |
|
||
| `NAT20_WORLD_CONTEXT` | Campaign context string | text | (none) | Injected into every summarisation prompt so the LLM recognises NPCs, places, and lore correctly |
|
||
| `NAT20_WORLD_CONTEXT_PATH` | Path to campaign context file inside container | container path | (none) | Alternative to `NAT20_WORLD_CONTEXT` for large campaign bibles you update independently |
|
||
| `NAT20_PLAYER_RECAP_STYLE` | Player recap format | `story` `diary` `bullets` `custom` | `story` | |
|
||
| `NAT20_PLAYER_RECAP_CUSTOM_PROMPT` | Custom prompt (only when style is `custom`) | text | (none) | |
|
||
|
||
#### Auth
|
||
|
||
| Variable | What it does | Values | Default | Notes |
|
||
|---|---|---|---|---|
|
||
| `NAT20_ADMIN_USERNAME` | Admin username | string | `admin` | Created on first startup if no users exist |
|
||
| `NAT20_ADMIN_PASSWORD` | Admin password | string | `admin` | |
|
||
| `NAT20_JWT_SECRET` | JWT signing key | string | (random) | Set to a fixed value so tokens survive restarts |
|
||
|
||
#### Setup wizard
|
||
|
||
| Variable | What it does | Values | Default | Notes |
|
||
|---|---|---|---|---|
|
||
| `NAT20_ONBOARDING_COMPLETED` | Skip the setup wizard | `true` or `false` | `false` | Set to `"true"` after you finish the wizard once |
|
||
|
||
### Campaigns
|
||
|
||
Campaigns organise sessions, settings, and context into separate silos.
|
||
On first run a **default** campaign is created automatically; you can
|
||
name it during the setup wizard and rename it later in Settings.
|
||
|
||
Key features:
|
||
|
||
- **Per-campaign settings** — each campaign has its own Whisper model,
|
||
LLM config, world context, and player recap style. Switch campaigns
|
||
in the sidebar to change which settings apply to new sessions.
|
||
- **Sharing** — the admin can share a campaign with other registered users
|
||
from the campaign's detail page. Shared users can upload sessions and
|
||
generate notes within that campaign.
|
||
- **User management** — admins can create additional users via the Users
|
||
page (`/users`). Every user authenticates with the same login form.
|
||
|
||
### Campaign context
|
||
|
||
You can paste context directly in the Settings page, or point to a file
|
||
inside the container using `NAT20_WORLD_CONTEXT_PATH`. The file version is
|
||
useful for large campaign bibles that you update independently.
|
||
|
||
## Notes on hardware
|
||
|
||
Transcription is GPU-bound and by far the slowest step for long sessions.
|
||
Summarization is comparatively light — a 7-8B parameter local model is
|
||
sufficient for most groups; only step up in size if you have the VRAM
|
||
headroom after Whisper's footprint (they don't run at the same time, so
|
||
you only need enough VRAM for whichever is currently running, plus normal
|
||
system overhead from other GPU-using services).
|
||
|
||
## Architecture
|
||
|
||
- `backend/` — FastAPI, SQLite (no external DB needed), WhisperX as a library (no nested Docker), in-process `ThreadPoolExecutor` background jobs (no Redis/Celery). JWT auth with HMAC-SHA256 tokens, scrypt password hashing.
|
||
- `frontend/` — React 18 + Vite + Tailwind, served via nginx (listens on port **8020**) which proxies `/api` to the backend. Auth state managed via React context; global fetch interceptor adds `Authorization: Bearer` to all API calls.
|
||
|
||
### CI/CD
|
||
|
||
The backend is split into two Docker images for fast rebuilds:
|
||
|
||
| Image | Build time | Contents |
|
||
|---|---|---|
|
||
| `whisperx-base` (`.deps`) | 5-10 min, only when `requirements.txt` changes | PyTorch 2.3.1 + cu121, WhisperX, ctranslate2, ffmpeg |
|
||
| `backend` (thin) | ~10 seconds | App code only (`FROM whisperx-base`) |
|
||
|
||
Both run as standard Docker Compose services — no special orchestration needed beyond GPU passthrough for the backend.
|