# Operations runbook

Day-to-day operation of **C-CARE v1.0.0 Release Candidate 2**. No real
credential, hostname or IP address appears in this document. Paths and service
names follow `PRODUCTION_DEPLOYMENT.md`; adapt them to your host.

## Start, stop and restart

C-CARE is a WSGI application behind a reverse proxy. Manage it with the service
manager, never from a shell inside the container.

```
sudo systemctl status  ccare
sudo systemctl start   ccare
sudo systemctl stop    ccare
sudo systemctl restart ccare
```

After any restart, confirm both probes before returning to business:

```
curl -fsS https://<host>/health/    # {"status":"ok"}
curl -fsS https://<host>/ready/     # {"status":"ready"}
```

`/health/` failing means the process is down or wedged. `/health/` passing while
`/ready/` returns 503 means the process is alive but the database is not usable;
do not send traffic.

Restart is safe at any time. C-CARE holds no in-process state that survives a
restart by design: pending inventory and communication work is either already
committed to PostgreSQL or explicitly recorded as a failed attempt.

## Logs

The application writes only to standard output. There is no log file, no log
database and no error mail, because `ADMINS` is empty by design and a mail
handler could disclose configuration.

| Channel | Level | Content |
| --- | --- | --- |
| `django.request` | ERROR | Unhandled exceptions with a server-side traceback. **Never** shown to a user |
| `django.security` | WARNING | CSRF rejections, disallowed hosts, suspicious operations |
| `django.db.backends` | WARNING | Database problems surfaced to Django |
| `django` | INFO | Lifecycle messages |
| `apps.communications.delivery` | INFO | Each delivery attempt and its definite outcome |

Rotation and retention are the host's responsibility, for example `journald`
with `SystemMaxUse=`, or a container log driver with a size cap. Read recent
errors with:

```
sudo journalctl -u ccare --since '1 hour ago' --no-pager
sudo journalctl -u ccare -p err --since '24 hours ago' --no-pager
```

Never paste raw log lines into a public channel: a traceback can contain
customer identifiers or record references.

## Health interpretation

| Symptom | `/health/` | `/ready/` | Meaning |
| --- | --- | --- | --- |
| Normal | 200 ok | 200 ready | Serving traffic |
| Database down | 200 ok | 503 not_ready | Stop routing; investigate PostgreSQL first |
| Migrating | 200 ok | 200 or 503 | `SELECT 1` still answers; the UI may fail on locked tables |
| Process wedged | timeout | timeout | Restart the service |
| Bad configuration | no response | no response | Check `journalctl`; settings raise at import |

## Migrations

Routine schema change is a deployment step, not a runtime task.

```
cd /opt/ccare/app
sudo -u ccare .venv/bin/python manage.py showmigrations | grep '\[ \]'
```

Any `[ ]` line is unapplied. Apply with `--noinput` **after** confirming a fresh
backup, then restart the service so every worker picks up the new schema:

```
sudo -u ccare .venv/bin/python manage.py migrate --noinput
sudo systemctl restart ccare
```

There is no down-migration path. If a migration must be undone, restore the
pre-deployment backup into a scratch database and decide explicitly whether to
promote it. Never edit an applied migration file.

Stale test or audit databases accumulate when a test run is interrupted. They are
harmless but consume disk; drop them only after confirming no run is active:

```
psql -l | grep -E 'test_|audit_'
```

## Backup

C-CARE stores all business state in one PostgreSQL database. `MEDIA_ROOT` is
empty in RC2 because the release defines no upload-bearing model; re-enable the
media rule in your proxy the day the first upload capability ships.

```
# Full logical backup in PostgreSQL custom format (restorable with pg_restore).
pg_dump -Fc --no-owner --no-privileges -f /var/backups/ccare/ccare-$(date +%F-%H%M).dump \
     -h <host> -p <port> -U ccare_backup <database>
```

Use a dedicated backup role, not the application role. Schedule it daily and
keep hourly dumps only if your recovery objective requires them.

| Concern | Recommendation |
| --- | --- |
| Retention | 24 hourly, 14 daily, 12 monthly, then delete |
| Off-host | **Mandatory.** Copy to object storage or a second host. A backup on the same disk is not a backup |
| Encryption | Encrypt at rest in object storage; for `pg_dump` on disk use `age`, `gpg` or filesystem encryption, and keep the key off-host |
| Verification | After each backup, `pg_restore --list` the file to prove it is readable and non-empty |
| Media files | Nothing to back up in RC2. When uploads exist, back up `MEDIA_ROOT` on the same schedule |

A backup whose restore has never been tested is not a backup. Run the drill
below on every release candidate.

## Restore

Restoring replaces live data. Take a decision, then follow the steps in order.

```
# 0. Stop the application so nothing writes during the restore.
sudo systemctl stop ccare

# 1. Restore into a SCRATCH database first and verify it. Never restore over
#    the live database as a first step.
createdb -h <host> -U ccare_backup ccare_restore_check
pg_restore -d ccare_restore_check --no-owner --no-privileges \
    -h <host> -U ccare_backup /var/backups/ccare/ccare-<timestamp>.dump

# 2. Verify the scratch copy.
DJANGO_SETTINGS_MODULE=config.settings.production DB_NAME=ccare_restore_check \
  .venv/bin/python manage.py check
DJANGO_SETTINGS_MODULE=config.settings.production DB_NAME=ccare_restore_check \
  .venv/bin/python manage.py makemigrations --check
psql -h <host> -U ccare_backup -d ccare_restore_check -tAc \
    'select count(*) from django_migrations'

# 3. Only after the scratch copy is proven: restore into the live database.
#    A database with active connections will refuse pg_restore, so stop the app
#    first and terminate leftover sessions.
dropdb -h <host> -U ccare_backup --if-exists <database>
createdb -h <host> -U ccare_backup -O ccare_app <database>
pg_restore -d <database> --no-owner --no-privileges \
    -h <host> -U ccare_backup /var/backups/ccare/ccare-<timestamp>.dump

# 4. Restart and verify.
sudo systemctl start ccare
curl -fsS https://<host>/ready/
```

Django's connection pool can hold sessions open after a restore and block
`DROP DATABASE`. Use `DROP DATABASE ... WITH (FORCE)` from a maintenance session
when a drop is refused.

### Backup/restore drill (release verification)

```
pg_dump -Fc --no-owner --no-privileges -f /tmp/drill.dump -h <host> -U ccare_backup <scratch_db>
dropdb -h <host> -U ccare_backup --force --if-exists <scratch_db>
createdb -h <host> -U ccare_backup <scratch_db>
pg_restore -d <scratch_db> --no-owner --no-privileges -h <host> -U ccare_backup /tmp/drill.dump
# verify: connect, manage.py check, makemigrations --check, /ready/, row counts
dropdb -h <host> -U ccare_backup --force --if-exists <scratch_db>
rm -f /tmp/drill.dump
```

## Disk monitoring

Alert on:

| Path | Watch for | Cause |
| --- | --- | --- |
| Database data directory | Growth beyond plan | Retention growth, WAL, vacuum lag |
| `STATIC_ROOT` | Unexpected files | A `collectstatic` to the wrong path |
| `MEDIA_ROOT` | Any growth in RC2 | Unexpected: RC2 has no uploads |
| Backup target | Growth | Retention policy not pruning |
| Log volume | Growth | A logging loop or a failing dependency |

Run `VACUUM` and `ANALYZE` on a schedule, or enable `autovacuum` deliberately.
C-CARE's append-only ledger and history tables benefit from regular vacuuming.

## Database monitoring

Watch at minimum:

- Connection count against `max_connections`. Size
  `max_connections ≥ application workers × threads + backup + monitoring + 10`.
- Long-running transactions and idle-in-transaction sessions; these hold locks
  and block the ledger writers.
- Replication lag, if a standby is configured.
- Table bloat and index usage for the history tables.
- Deadlock and lock-timeout errors in the application log.

Two recurring symptoms:

| Symptom | Likely cause | Action |
| --- | --- | --- |
| `too many connections for role` | Worker count exceeds the database limit | Lower workers or raise `max_connections` |
| Hangs during `migrate` or reporting | A long transaction holding a lock | Find the blocking PID, end it deliberately |

## Communication failures

C-CARE records every notification attempt with a **definite** outcome:
accepted, rejected or failed. It never retries silently, so a provider outage
surfaces as failed records rather than invisible delay.

| Outcome | Meaning | Operator action |
| --- | --- | --- |
| Accepted by provider | The provider took responsibility for delivery | None |
| Rejected | The provider definitively refused (bad number, invalid address, policy) | Correct the customer contact value or the template; do not retry unchanged |
| Failed | The attempt could not be completed (timeout, connection error, unconfigured provider) | Check configuration and provider status, then re-send deliberately |

Diagnostics:

```
# Channel configuration and provider selection.
DJANGO_SETTINGS_MODULE=config.settings.production .venv/bin/python -c \
  "import django; django.setup(); from django.conf import settings; \
   print({k: (v or 'unconfigured') for k, v in settings.COMMUNICATION_PROVIDERS.items()}); \
   print('external delivery enabled:', settings.COMMUNICATIONS_ALLOW_EXTERNAL)"

# Delivery history for a case, or the whole history page in the UI.
# Operational workspace -> Case -> Case communications
```

Rules to remember:

- Outbound delivery is disabled unless `COMMUNICATIONS_ALLOW_EXTERNAL=True`.
- A channel with no provider in production raises `ImproperlyConfigured` at send
  time rather than silently succeeding. **RC2 ships no SMS adapter**; leave SMS
  unconfigured.
- `EMAIL_USE_TLS` and `EMAIL_USE_SSL` are mutually exclusive, and SMTP
  credentials must be supplied as a pair or not at all.
- An in-memory email backend is rejected in production; a production SMTP
  backend with external delivery disabled is also rejected.
- Delivery failures never block the business transaction that triggered them.
  A failed notification must not roll back a repair or an invoice.

## Scheduled jobs

RC2 ships three read-only or idempotent commands. Run them from cron or a timer;
none requires a task queue.

| Command | Purpose | Safety |
| --- | --- | --- |
| `monitor_service_sla` | Derive SLA state and escalate overdue cases | Read-only except for recorded escalations; idempotent |
| `queue_appointment_reminders` | Queue appointment reminder notifications | Idempotent per event; respects `COMMUNICATIONS_ALLOW_EXTERNAL` |
| `audit_inventory` | Reconcile the inventory ledger | Read-only |

Run `monitor_service_sla` frequently (every few minutes) and the reminder queue
on a short interval. Keep a single instance of each to avoid duplicated work.

## User and administrator recovery

C-CARE has no self-service password reset: passwords are never emailed. All
recovery is performed by an administrator.

```
# Create a replacement administrator interactively (prompts, never echoes).
sudo -u ccare .venv/bin/python manage.py createsuperuser

# Reset an existing user's password by activating the account in the shell.
sudo -u ccare .venv/bin/python manage.py shell -c "
from django.contrib.auth import get_user_model
from django.contrib.auth.password_validation import validate_password
user = get_user_model().objects.get(username='<username>')
password = input('new password: ')
validate_password(password, user)
user.set_password(password)
user.is_active = True
user.save()
print('password updated')"
```

Recovery rules:

- Never reuse a password. Record who reset what, when and why.
- If an administrator account is compromised, deactivate it first, then create a
  replacement. Deactivation cascades to its role assignments and organizational
  scope.
- Deactivating a company, region, service center or department deactivates its
  descendants. Reactivate the top of the affected branch, never individual
  leaves, or the activation path stays inconsistent.
- Authorization is role-permission plus organizational scope. A user with the
  right permission but the wrong company, region, center or department sees
  nothing and is not an error to work around.

## Common operational failures

| Symptom | Cause | Resolution |
| --- | --- | --- |
| `ImproperlyConfigured: Set DJANGO_SECRET_KEY in your local .env or environment.` | Environment not loaded by the service manager | Confirm `EnvironmentFile=` and that the service, not your shell, sees it |
| `ImproperlyConfigured: Production requires explicit DJANGO_ALLOWED_HOSTS` | Missing or wildcard host list | Set explicit hostnames |
| Redirect loop on every request | Proxy TLS and `DJANGO_TRUSTED_PROXY_HEADER` disagree | Set the variable to the header the proxy actually sends, and strip client-supplied variants at the proxy |
| `400 Bad Request` on every POST | CSRF origin mismatch | Add the exact `https://host` to `DJANGO_CSRF_TRUSTED_ORIGINS` |
| 404 on the dashboard after adding a user | The user has no organizational assignment | Assign the user to a company, region, center or department |
| 403 everywhere for one persona | Permission without scope, or scope without permission | Check both the role's permissions and the organizational assignment |
| `Server Unavailable` in the UI | Unapplied migration after a deploy | `migrate --noinput`, then restart |
| Static assets 404 | `collectstatic` not run, or the proxy alias is wrong | Re-run `collectstatic`; verify the proxy alias matches `STATIC_ROOT` |
| Readiness 503, application works | Transient database latency, or the readiness statement timeout is too low | Check the database first; only then raise `READY_STATEMENT_TIMEOUT_MS` |
| `pg_restore` refuses to drop the database | An open session from the application pool | Stop the service, or use `DROP DATABASE ... WITH (FORCE)` |

## Incident response basics

1. **Record first.** Note the time, the observed symptom and whether `/health/`
   and `/ready/` agree. Do not change anything before you can describe it.
2. **Classify.** Availability (probes failing), integrity (wrong or missing
   business data), security (unexpected access or disclosure), or delivery
   (notifications only).
3. **Contain.** For an availability incident, stop routing traffic. For a
   suspected security incident, deactivate the affected account or organization
   branch through the application, not by editing rows.
4. **Preserve evidence.** Copy the relevant log window and the current database
   state with a fresh `pg_dump` before any repair.
5. **Repair.** Prefer a configuration correction. Deploy, migrate or restore only
   with a verified backup in hand.
6. **Verify.** Re-run `/health/`, `/ready/` and the smoke tests from
   `PRODUCTION_DEPLOYMENT.md` §12.
7. **Record.** Note the cause, the action taken and any follow-up. For a data
   integrity incident, reconcile the affected domain before declaring it closed:
   inventory, commercial, SLA and communications records are append-only and must
   never be edited to "fix" a discrepancy. Raise a correction through the
   application's own workflow or record the discrepancy for review.