An alert count is not an incident count. A service may produce dozens of notifications during one problem, while a single failure can affect several tills. A useful report first defines what counts as an incident and how durations are measured.
Automation collects tickets and events, correlates operational episodes and calculates metrics. AI can summarise the results while preserving uncertainty about causes.
A fictional incident register
The example observation period runs from 5 October 2026 at 00:00 to 12 October at 00:00, Europe/Rome, with the end excluded. Each row represents a distinct incident for this exercise; operational correlation needs checking.
| Incident | Store | Start | Recovery | Available evidence |
|---|---|---|---|---|
| INC101 | ST017 | 5 October 09:00 | 5 October 09:30 | Recovered after restart; cause unconfirmed |
| INC102 | ST017 | 6 October 14:00 | 6 October 14:45 | Network interruption confirmed in the ticket |
| INC103 | ST018 | 11 October 23:50 | Open at period end | Delayed data upload; investigation ongoing |
INC101 and INC102 last 30 and 45 minutes respectively. INC103 has ten observed minutes within the period and is still open. Those ten minutes are not its final duration.
Choose metrics that answer clear questions
Three incidents started in the period, two recovered before its end and one remained open. The mean duration for the two recovered cases is 37.5 minutes, from recorded onset to recovery. An open-incident count must also include earlier incidents still active, although none appear in this example.
The sum of observed durations is 85 minutes. It is not automatically total business downtime: incidents can overlap or affect the same service. Downtime requires an affected scope and a rule for merging overlapping intervals. These durations do not establish lost sales.
Preserve the source evidence
Normalise timestamps, correlate events with stable identifiers and keep links to source tickets. Administrative ticket closure may happen later than recovery. Store both timestamps separately when verified.
A restart before recovery is an action, not necessarily the cause. Keep symptoms, actions, confirmed causes and open investigations in separate fields. A missing cause should remain missing in the report.
A prompt for the weekly summary
Write an operational commentary using the attached rows and metrics.
3 incidents started in the period; 2 recovered; 1 open at period end.
Mean duration of the 2 recovered cases: 37.5 minutes from onset to recovery.
INC103 remains open with 10 observed minutes in the period.
Separate facts, confirmed causes and checks still needed.
Do not describe summed durations as total downtime.
Do not infer economic impact or missing causes.
State the number of cases behind every mean and cite incident IDs.
An authored reference commentary is:
Three incidents started during the period. Two recovered, with a mean duration of 37.5 minutes across two cases. INC103 remains open at the observation cutoff. INC102 confirms a network interruption; recovery after restart in INC101 does not establish its cause. Complete the investigation of INC101 and update INC103.
Test the reporting boundaries
Include incidents spanning weeks, duplicate tickets, missing timestamps and overlapping intervals. Open cases must not enter the mean of final durations. An earlier incident can be excluded from new-incident counts while remaining part of open counts and observed duration.
Every report should show its observation period, extraction time and source coverage. Retain ticket links with appropriate permissions so readers can verify conclusions and assign follow-up work.
Practise providing facts and verification criteria in the free prompt engineering mini-course, open without registration.