Files
Nat20-Notes/README.md
KansaiGaijin f7eeae8469
All checks were successful
Build and Push / build (push) Successful in 1m17s
docs: update README with auth, campaigns, and CI/CD changes
2026-07-30 16:09:54 +12:00

157 lines
8.8 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Nat20 Notes
Turn a recorded tabletop RPG session (audio or video) into two documents:
a full GM/DM session log, and a spoiler-free player recap — using local
transcription (WhisperX) and either a local LLM (Ollama) or any
OpenAI-compatible hosted API for summarization.
## Requirements
- Docker + Docker Compose
- An NVIDIA GPU with drivers + [NVIDIA Container Toolkit](https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html) installed on the host (for transcription; CPU-only works but is much slower)
- A free [HuggingFace token](https://huggingface.co/settings/tokens) — needed for speaker diarization. You'll also need to accept the terms on the gated pyannote model page it links you to on first run.
- Either: [Ollama](https://ollama.com) running somewhere reachable from this app (local or LAN), **or** an API key for a hosted LLM (OpenAI, or any OpenAI-compatible provider)
## Quick start
```bash
git clone <this-repo>
cd nat20-notes
# Review docker-compose.yml — see Configuration below for env vars
docker compose up -d --build
```
Open `http://<your-server>:8020` (d20 nod) and **log in** with the default
admin credentials (`admin` / `admin`, overridable via `NAT20_ADMIN_USERNAME`
and `NAT20_ADMIN_PASSWORD`). The first user is created automatically on
first startup — no separate registration step.
Then follow the setup wizard:
1. Choose a Whisper model size based on your GPU's available VRAM (guidance shown in-app)
2. Give your campaign a name
3. Choose local Ollama or a hosted API for summarization, and paste your HF token
4. Optionally paste campaign/world context (NPC names, places) so summaries recognise them correctly
A reference compose file using named volumes only (no host paths) is at
[`docker-compose.example.yml`](./docker-compose.example.yml).
## Using it
1. **Log in** with your username and password. Admins can create additional users at `/users`.
2. **Select a campaign** from the sidebar (or use the default). Each campaign has its own settings, sessions, and player recaps.
3. **Upload** a recording (audio or video — video is auto-converted, audio-only files skip that step and are much faster). Files over **90 MB** are auto-chunked with 3-way concurrent uploads.
4. **Transcribe** — runs in the background. On completion you're **automatically taken** to the speaker-naming screen.
5. **Name your speakers** — a waveform-style "session reel" shows each detected speaker's segments; click any point to jump the audio there and hear who's talking, then type in their name. Use the checkboxes to **merge** speakers (e.g. when diarization over-splits one person into `SPEAKER_00`, `SPEAKER_05`, etc.).
6. **Generate notes** — click "Done naming" and confirm; notes generation starts automatically and you're taken to the job progress screen. When finished, navigate to the notes viewer.
7. **Review & tweak** — regenerate notes, rename speakers, or delete jobs/sessions from the session detail page.
## Data layout
The app stores everything under `/data` (inside the container):
| Directory / File | Contents |
|---|---|
| `campaigns/{id}/audio/` | Uploaded recordings and extracted audio, organised per campaign |
| `campaigns/{id}/transcriptions/` | Per-session transcript JSON files |
| `campaigns/{id}/notes/` | Generated notes (GM log + player recap) |
| `app.db` | SQLite database (users, campaigns, campaign_shares, sessions, speakers, campaign/global settings, jobs) |
## Configuration
### Environment variables (`NAT20_*`)
Set these under the `backend` service in `docker-compose.yml` to prefill
the setup wizard and override defaults. All are optional — the wizard
and Settings page can set them at runtime.
#### Transcription
| Variable | What it does | Values | Default | Notes |
|---|---|---|---|---|
| `NAT20_WHISPER_MODEL` | Whisper model size | `tiny` `base` `small` `medium` `large-v3` | `medium` | Larger = more accurate but uses more VRAM. `medium` fits most 6-8 GB GPUs; `large-v3` needs ~10 GB+ |
| `NAT20_WHISPER_COMPUTE_TYPE` | Compute precision | `int8` `float16` `float32` | `int8` | `int8` = fastest / least VRAM. `float16` = more accurate, more VRAM. `float32` = full precision, slowest |
| `NAT20_HF_TOKEN` | HuggingFace token for speaker diarization | `hf_...` | (none) | Required for speaker attribution. Must accept pyannote gated-model terms with the same account first |
#### LLM backend
| Variable | What it does | Values | Default | Notes |
|---|---|---|---|---|
| `NAT20_OLLAMA_HOST` | Ollama server URL | URL | `http://localhost:11434` | Set to your Ollama host |
| `NAT20_OLLAMA_MODEL` | Ollama model name | any model on your server | `qwen2.5:7b` | |
| `NAT20_API_BASE_URL` | OpenAI-compatible API base | URL | `https://api.openai.com/v1` | Uncomment and set to switch from Ollama |
| `NAT20_API_KEY` | API key | string | (none) | |
| `NAT20_API_MODEL` | API model name | string | `gpt-4o-mini` | |
#### Summarization
| Variable | What it does | Values | Default | Notes |
|---|---|---|---|---|
| `NAT20_CHUNK_WORD_TARGET` | Target words per summarisation chunk | number | `2500` | Long transcripts are split into chunks, each summarised separately, then combined. **Lower** = more LLM calls but finer granularity (good for models with small context windows). **Higher** = more context per chunk but may exceed the model's window. `2500` is safe for most models (8K128K context) |
| `NAT20_WORLD_CONTEXT` | Campaign context string | text | (none) | Injected into every summarisation prompt so the LLM recognises NPCs, places, and lore correctly |
| `NAT20_WORLD_CONTEXT_PATH` | Path to campaign context file inside container | container path | (none) | Alternative to `NAT20_WORLD_CONTEXT` for large campaign bibles you update independently |
| `NAT20_PLAYER_RECAP_STYLE` | Player recap format | `story` `diary` `bullets` `custom` | `story` | |
| `NAT20_PLAYER_RECAP_CUSTOM_PROMPT` | Custom prompt (only when style is `custom`) | text | (none) | |
#### Auth
| Variable | What it does | Values | Default | Notes |
|---|---|---|---|---|
| `NAT20_ADMIN_USERNAME` | Admin username | string | `admin` | Created on first startup if no users exist |
| `NAT20_ADMIN_PASSWORD` | Admin password | string | `admin` | |
| `NAT20_JWT_SECRET` | JWT signing key | string | (random) | Set to a fixed value so tokens survive restarts |
#### Setup wizard
| Variable | What it does | Values | Default | Notes |
|---|---|---|---|---|
| `NAT20_ONBOARDING_COMPLETED` | Skip the setup wizard | `true` or `false` | `false` | Set to `"true"` after you finish the wizard once |
### Campaigns
Campaigns organise sessions, settings, and context into separate silos.
On first run a **default** campaign is created automatically; you can
name it during the setup wizard and rename it later in Settings.
Key features:
- **Per-campaign settings** — each campaign has its own Whisper model,
LLM config, world context, and player recap style. Switch campaigns
in the sidebar to change which settings apply to new sessions.
- **Sharing** — the admin can share a campaign with other registered users
from the campaign's detail page. Shared users can upload sessions and
generate notes within that campaign.
- **User management** — admins can create additional users via the Users
page (`/users`). Every user authenticates with the same login form.
### Campaign context
You can paste context directly in the Settings page, or point to a file
inside the container using `NAT20_WORLD_CONTEXT_PATH`. The file version is
useful for large campaign bibles that you update independently.
## Notes on hardware
Transcription is GPU-bound and by far the slowest step for long sessions.
Summarization is comparatively light — a 7-8B parameter local model is
sufficient for most groups; only step up in size if you have the VRAM
headroom after Whisper's footprint (they don't run at the same time, so
you only need enough VRAM for whichever is currently running, plus normal
system overhead from other GPU-using services).
## Architecture
- `backend/` — FastAPI, SQLite (no external DB needed), WhisperX as a library (no nested Docker), in-process `ThreadPoolExecutor` background jobs (no Redis/Celery). JWT auth with HMAC-SHA256 tokens, scrypt password hashing.
- `frontend/` — React 18 + Vite + Tailwind, served via nginx (listens on port **8020**) which proxies `/api` to the backend. Auth state managed via React context; global fetch interceptor adds `Authorization: Bearer` to all API calls.
### CI/CD
The backend is split into two Docker images for fast rebuilds:
| Image | Build time | Contents |
|---|---|---|
| `whisperx-base` (`.deps`) | 5-10 min, only when `requirements.txt` changes | PyTorch 2.3.1 + cu121, WhisperX, ctranslate2, ffmpeg |
| `backend` (thin) | ~10 seconds | App code only (`FROM whisperx-base`) |
Both run as standard Docker Compose services — no special orchestration needed beyond GPU passthrough for the backend.