# Service SLA monitoring — Phase 4C

SLA is operational evidence. **SLA status never changes ServiceCase state.**
The frozen service, authorization, taxonomy, inventory, commercial and
communications domains retain their existing behavior.

## Policy and enrollment

`apps.sla` defines `SlaPolicy`, `ServiceSla`, `SlaEscalation` and
`EscalationCommunication`. Migration `sla.0001_initial` adds only new tables and
constraints. No historical migration is changed.

A policy belongs to one company and optionally one service center and/or one
Product Category. Product Category is the serviced device's ProductModel category;
it is not a complaint classification or ServiceCategory. No priority, calendar,
holiday or shift is invented.

Targets and warning windows are integer elapsed minutes. Target must be positive;
warning must be between zero and the target. Effective dates are inclusive and
use the service acceptance date in the configured project timezone (Asia/Dhaka).
Active policies match in this order:

1. Center and Product Category.
2. Center, all product categories.
3. Company and Product Category.
4. Company default.

Active policies with identical company/center/category scope may not overlap
effective dates. Company locking serializes validation and enrollment, including
concurrent inserts. Matching fails on a tie rather than picking an arbitrary row.
Inactive policies do not match. Policy company and code cannot change; deactivate
instead of deleting. Other edits apply only to future enrollments.

Enrollment is explicit (`enroll`, case monitor action, or monitoring command).
There is no new hook in frozen intake. The first enrollment snapshots the current
matching policy configuration applicable to the case's acceptance date, including
target, warning, dimensions, effective dates, policy version timestamp, clock
definition and optional notification-template references. Later policy changes or
deactivation do not reselect or change this snapshot.

For delayed enrollment, this is **not** a reconstruction of what a mutable policy
looked like at intake. Configure policies before operations and run monitoring
regularly. Cases already cancelled, ready for delivery, delivered or closed are
not newly enrolled; no synthetic historical SLA is backfilled. Existing snapshots
remain available after completion. Unmatched active cases are counted by the
command and are absent from the enrolled-case dashboard.

## Clock and observations

Start: authoritative `ServiceCase.received_at`, formal service acceptance.
Stop: immutable `ServiceCaseDeliveryRelease.readied_at`, created by the existing
QC-passed-to-ready-for-delivery transition. This measures **acceptance to ready**.
Repair completion and QC completion are earlier, separate events. Customer
handover and case closure occur later and never extend this SLA clock.

`due_at = received_at + target_minutes`; `warning_at = due_at - warning_minutes`.
All timestamps are aware. Time is elapsed continuously: no automatic pause for
parts, approval, quotation, payment, weekends or pickup. Pause/resume is future
work. Missing required release/cancellation evidence raises a validation error;
`updated_at` is never substituted for authoritative evidence.

State is derived from the snapshot and authoritative timestamp at observation:

| State | Rule |
| --- | --- |
| ON_TRACK | Not stopped, observation before warning |
| DUE_SOON | Not stopped, warning <= observation <= due |
| OVERDUE | Not stopped, observation > due |
| COMPLETED_ON_TIME | Ready timestamp <= due |
| COMPLETED_LATE | Ready timestamp > due |
| CANCELLED | Authoritative cancellation observed; not a successful completion |

At exactly due, an unfinished case is DUE_SOON, not overdue; completion is on time.
With zero warning, only the exact due instant is DUE_SOON. Durations stop at ready
or cancellation; remaining/overdue durations are clamped at zero. Cancellation
stops new escalation and is not counted as on-time completion. No duplicate
ServiceCase status or completion timestamp is stored in SLA.

## Escalation and communications

Each snapshot can have one DUE_SOON and one OVERDUE event. A first observation
already overdue records only OVERDUE; it does not manufacture an earlier warning.
Events retain the observation time, actor, MANUAL/MONITOR origin, case reference,
due time and policy snapshot. They are append-only, including after completion.

Policies may select existing Phase 4B customer templates independently for each
level. There is no invented internal recipient model. Templates should describe
the **recorded SLA alert**, not promise that a case is still overdue at delivery.
The supported generic `request_notification` API receives `SLA_DUE_SOON` or
`SLA_OVERDUE`, the service case reference and a stable escalation UUID event key.
Only existing placeholders are supplied: `customer_name`,
`service_case_reference`, `device_name`. Current validated destination selection,
template rendering and dispatch remain Phase 4B responsibilities. A template
reference is frozen on enrollment; message content is snapshotted by Phase 4B
when queued, not when SLA is enrolled.

Monitoring commits SLA evidence before attempting communications. Queuing occurs
in a separate transaction/savepoint; errors create sanitized append-only
`NOTIFICATION_REQUEST_FAILED` results without changing the case or SLA. Subsequent
runs retry unqueued historical events. A successful link is unique per escalation;
the communications idempotency key is also stable. Pending/failed/cancelled
notifications are managed by Phase 4B and are never replaced by SLA retries.
This subsystem never sends messages or calls providers.

## Command and UI

```powershell
python manage.py monitor_service_sla --actor <authorized-user-uuid>
python manage.py monitor_service_sla --actor <authorized-user-uuid> --company <company-uuid> --center <center-uuid>
```

An explicit active actor is required; no hidden superuser/system bypass is used.
The command reports enrolled/unmatched cases, observed state counts, new
escalations, queued messages and communication failures. Repeated execution is
safe. It must run outside an enclosing transaction to preserve the commit boundary.
No scheduler or Celery is installed; a future scheduler can invoke this command.

- `/sla/`: scoped, paginated operational dashboard; center/state/ready-since
  filters; target, start, due, state, engineer and durations. Counts respect center
  and ready-since filters before the selected state. “Completed late today” uses
  today's project-local date, an explicit recent window.
- `/sla/cases/<case UUID>/`: snapshot and escalation/communication history.
- `/sla/cases/<case UUID>/monitor/`: CSRF-protected confirmation and POST operation
  for one case, using exactly the same services as the command.
- `/sla/policies/`, `new/`, `<policy UUID>/`: scoped policy management.
- Existing ServiceCase Admin detail gains authorized SLA/monitor links through a
  narrow template override retaining existing object tools. No Admin lifecycle
  logic is changed; SLA history is not editable in generic Admin.

This is operational monitoring, not a replacement for Phase 3D analytics. Lists
and counts are SQL-scoped before pagination; arbitrary IDs cannot expand scope.

## Authorization and concurrency

Existing authorization APIs evaluate three new capabilities: `sla.view_sla`,
`sla.monitor_sla`, and `sla.manage_slapolicy`. View and monitor are evaluated
against the case ServiceCenter; company/region/center/department containment
remains unchanged. Policy management requires company-level scope. Staff status,
Groups and direct Django permissions alone provide no business scope. Notification
queuing additionally requires existing `communications.send_notification`.

Actor SHARE, company locks and the existing role/assignment dependency locks
protect write authorization. Policy saves and enrollment take Company UPDATE;
escalation observes locked ServiceCase and ServiceSla rows, compatible with
authoritative service transitions. Communication requests take Company UPDATE
before case/escalation locks, then reuse the frozen queue abstraction.

Database uniqueness protects one SLA per case, one event per level and one
successful communication link per event. Supported APIs reject record mutation,
bulk writes and deletion of evidence. Policy overlap prevention is transactional
through the supported service; direct SQL is not a supported administration API.

## Limitations and verification

No retroactive policy revision reconstruction, SLA revision/resnapshot workflow,
pause calendar, periodic escalation ladder, automatic dispatcher or scheduler.
Company locks favor correctness over maximum policy/enrollment throughput.
Monitoring scans authorized cases; no full-project analytics materialization is
introduced. SLA setup and monitoring permissions must be explicitly assigned.

Focused tests cover policy validity and precedence, effective dates, timezone and
boundary calculations, real ready-for-delivery evidence, cancellation, immutable
snapshots, scope/IDOR/CSRF, command execution, notification isolation and PostgreSQL
lock contention. Verification results are reported with implementation delivery;
the full project audit is intentionally not part of Phase 4C.

Verified on 2026-10-01 (Asia/Dhaka): 47 focused SLA tests passed in 88.601 seconds;
the single `audit_project --quick` run passed all checks and 19 smoke tests
(29.021 seconds for tests, 69.260 seconds total). Migration `sla.0001_initial` is
applied. No full audit, frozen-test edits, historical-migration edits, staging,
commit or push was performed.
