# Production backup, restore, and upgrade runbook Production web and worker processes never migrate on startup. The Compose file also requires `GOCHAT_IMAGE_REF` to be an immutable `image@sha256:digest`; tags such as `latest` are not an acceptable rollback record. The bundled `gochat_storage` volume is shared and durable for replicas on one Docker host. Multi-host replicas must provision that volume with a shared volume driver/filesystem; never use separate node-local volumes. ## Recovery objectives - RPO: 24 hours. Run the encrypted backup at least daily and alert if the latest off-site bundle is older than 24 hours. - RTO: 4 hours. Rehearse a clean-environment restore quarterly and record the script's `rpo_seconds`, `rto_seconds`, image digest, migration version, account count, and attachment count. - Local and off-site backup directories must be different mounts/failure domains. The bundle is AES-256 encrypted; keep the passphrase file in the secret manager, never beside either backup copy. ## Daily encrypted backup Set these operator-owned paths before any Compose command: ```bash export GOCHAT_IMAGE_REF='ghcr.io/gochat/gochat@sha256:' export GOCHAT_BACKUP_DIR='/mnt/backup-local/gochat' export GOCHAT_BACKUP_OFFSITE_DIR='/mnt/gochat-offsite' export GOCHAT_BACKUP_OFFSITE_SOURCE='backup.example.com:/gochat' export GOCHAT_BACKUP_OFFSITE_FSTYPE='nfs4' export GOCHAT_BACKUP_PASSPHRASE_FILE='/run/secrets/gochat-backup-passphrase' export GOCHAT_CONNECTOR_BACKUP_NAME="connector-$(date -u +%Y%m%dT%H%M%SZ).db" ``` Backup/restore are audited one-shot containers and run as root only to read or rebuild Docker volumes; web and worker remain non-root. `GOCHAT_BACKUP_OFFSITE_DIR` must be an existing external mount point provisioned outside this Compose project. Set its approved source and filesystem type to the exact values reported by `findmnt -M "$GOCHAT_BACKUP_OFFSITE_DIR"`; the preflight rejects `/`, in-memory filesystems, unapproved mount metadata, and the local backup device. Create a consistent online Connector backup, then create and verify the single encrypted bundle containing PostgreSQL, attachments, and that Connector copy: ```bash docker compose -f deploy/docker/docker-compose.prod.yml exec shangwutong \ shangwutong backup --output "/backup/$GOCHAT_CONNECTOR_BACKUP_NAME" docker compose -f deploy/docker/docker-compose.prod.yml --profile ops run --rm backup ``` Schedule those two commands daily. Preserve their final `backup=... offsite=...` line as audit evidence. A failed dump, archive checksum, decrypt/list check, or off-site copy exits non-zero and must alert. ## Clean-environment restore rehearsal Use an isolated host/project with empty volumes and an empty PostgreSQL database. Do not point this procedure at the live project. ```bash export COMPOSE_PROJECT_NAME=gochat-restore-$(date +%Y%m%d) export GOCHAT_RESTORE_BUNDLE='gochat-.tar.enc' docker compose -f deploy/docker/docker-compose.prod.yml --profile ops run --rm restore ``` The restore refuses a non-empty database, attachment directory, or Connector DB. It verifies the encrypted bundle and internal checksums before restoring, then prints the RPO/RTO, image version, migration version, accounts, and attachments. Afterward start web/worker with the recorded digest and verify `/health`, one attachment URL, and Connector `readyz`. ## Upgrade preflight 1. Record the running digest as `OLD_IMAGE`; pull and record `NEW_IMAGE` by digest. Never derive rollback state from a mutable tag. 2. Complete the backup above and restore it in the isolated environment. 3. Stop writes or enter the maintenance window. Migration 000079 takes `SHARE ROW EXCLUSIVE` locks; migration 000048 takes `ACCESS EXCLUSIVE` on the rollup table. The migration job uses a 5-second lock timeout and 15-minute statement timeout, so contention fails instead of waiting indefinitely. 4. Capture these pre-migration checks: ```sql SELECT count(*) AS rollups FROM reporting_events_rollups; SELECT count(*) AS swt_source_rows FROM conversations WHERE COALESCE(custom_attributes, '{}'::jsonb) ?| ARRAY['swt_source_url','swt_source_search_term']; SELECT account_id, (config->>'assistant_id')::bigint, count(*) FROM agent_bots WHERE bot_type = 'captain' AND config->>'assistant_id' ~ '^[0-9]+$' GROUP BY 1, 2 HAVING count(*) > 1; ``` While migrating, observe lock waits with: ```sql SELECT pid, wait_event_type, wait_event, clock_timestamp() - query_start AS elapsed, query FROM pg_stat_activity WHERE datname = current_database() AND state <> 'idle'; ``` ## One-shot migration and application rollout Run exactly one named migration service before changing web/worker. Its output is the audit log; `golang-migrate` is idempotent and PostgreSQL serializes competing migration runners, but the deployment pipeline must contain only this one step. ```bash export GOCHAT_IMAGE_REF="$NEW_IMAGE" time docker compose -f deploy/docker/docker-compose.prod.yml --profile ops \ up --abort-on-container-exit --exit-code-from migrate migrate docker compose -f deploy/docker/docker-compose.prod.yml logs migrate docker compose -f deploy/docker/docker-compose.prod.yml up -d --no-deps gochat worker curl -fsS http://127.0.0.1:3000/health ``` Post-migration checks: ```sql SELECT version, dirty FROM schema_migrations; SELECT count(*) AS rollups FROM reporting_events_rollups; SELECT count(*) AS duplicate_bot_bindings FROM ( SELECT account_id, captain_assistant_id FROM agent_bots WHERE captain_assistant_id IS NOT NULL GROUP BY 1, 2 HAVING count(*) > 1 ) duplicates; SELECT count(*) AS orphaned_inbox_bindings FROM agent_bot_inboxes b LEFT JOIN agent_bots a ON a.id = b.agent_bot_id WHERE a.id IS NULL; ``` Migration 000048 preserves every legacy rollup row, 000076 retains legacy Connector attributes for old-image compatibility, and 000079 repairs references before deduplication. Any count mismatch or dirty version stops the rollout. ## Rollback Rollback the application only, using the recorded digest: ```bash export GOCHAT_IMAGE_REF="$OLD_IMAGE" docker compose -f deploy/docker/docker-compose.prod.yml up -d --no-deps gochat worker curl -fsS http://127.0.0.1:3000/health ``` Production `migrate down` and negative `migrate steps` are blocked. Never run a destructive schema down during an application rollback. If the new schema itself is unusable, stop web/worker and restore the verified pre-upgrade bundle into a clean database and volumes.