Files
gochat/backend/configs/prometheus_alerts.yml
T
rogee aeddedf2a3 Reorganize repo: backend/, deploy/, docs/ layout + AGENTS.md
Restructure the monorepo into clear top-level directories:
- backend/: Go module root (cmd, internal, pkg, configs, migrations,
  docs/swagger, scripts, tests, go.mod, Makefile, .air.toml)
- deploy/: Docker (Dockerfile, docker-compose*), quickstart, fluentd
- docs/: project documentation + reports/ (moved from repo root)
- AGENTS.md: new AI coding-agent guide at repo root

Update all references to the new layout:
- Dockerfile: COPY backend/go.mod, COPY backend/ (context = repo root)
- docker-compose files: context ../.., dockerfile deploy/docker/Dockerfile,
  env_file ../../.env, volume mounts ../../backend:/app
- deploy/quickstart/compose.yaml: dockerfile deploy/docker/Dockerfile
- CI: working-directory: backend for go commands, file deploy/docker/Dockerfile,
  coverage path backend/coverage.out, health_check backend/scripts/
- backend/Makefile: docker target uses -f ../deploy/docker/Dockerfile ../
- README: architecture tree, quickstart, config paths updated

Move root stray scripts (rename_models.*, run_m11_tests.sh, verify_build.sh,
gorm_bool_main.go) to backend/scripts/legacy/. All moves via git mv to
preserve history. Build, vet, SQLite tests, and docker compose config verified.
2026-07-07 14:44:12 +08:00

85 lines
2.5 KiB
YAML

# GoChat Prometheus Alert Rules
# Reference: Chatwoot production monitoring with Sidekiq queue alerts
# Adjust thresholds based on your deployment scale
groups:
- name: gochat-app
rules:
# Application down
- alert: GoChatAppDown
expr: up{job="gochat"} == 0
for: 1m
labels:
severity: critical
annotations:
summary: "GoChat application is down"
description: "GoChat instance {{ $labels.instance }} has been down for more than 1 minute."
# High error rate
- alert: GoChatHighErrorRate
expr: rate(http_requests_total{job="gochat", status=~"5.."}[5m]) / rate(http_requests_total{job="gochat"}[5m]) > 0.05
for: 5m
labels:
severity: warning
annotations:
summary: "GoChat error rate above 5%"
description: "Error rate is {{ $value | humanizePercentage }} over the last 5 minutes."
# High memory usage
- alert: GoChatHighMemory
expr: gochat_go_memory_alloc_bytes / (1024 * 1024) > 400
for: 5m
labels:
severity: warning
annotations:
summary: "GoChat memory usage above 400MB"
description: "Memory allocation is {{ $value }}MB."
# Too many goroutines
- alert: GoChatHighGoroutines
expr: gochat_go_goroutines > 1000
for: 5m
labels:
severity: warning
annotations:
summary: "GoChat goroutine count above 1000"
description: "{{ $value }} goroutines running."
- name: gochat-infra
rules:
# PostgreSQL down
- alert: GoChatPostgresDown
expr: up{job="gochat-postgres"} == 0
for: 1m
labels:
severity: critical
annotations:
summary: "PostgreSQL is down"
# Redis down
- alert: GoChatRedisDown
expr: up{job="gochat-redis"} == 0
for: 1m
labels:
severity: critical
annotations:
summary: "Redis is down"
# Redis memory approaching limit
- alert: GoChatRedisMemoryHigh
expr: redis_memory_used_bytes / redis_memory_max_bytes > 0.8
for: 5m
labels:
severity: warning
annotations:
summary: "Redis memory usage above 80%"
# PostgreSQL connections exhausted
- alert: GoChatPostgresConnectionsHigh
expr: pg_stat_activity_count / pg_settings_max_connections > 0.8
for: 5m
labels:
severity: warning
annotations:
summary: "PostgreSQL connection usage above 80%"