An incident is a confirmed period of degraded or unavailable service that requires human attention. Incidents are usually opened automatically when monitoring detects sustained failures and closed when service is verified restored.
Mature incident management distinguishes between an alert (a single failed check), an incident (a confirmed sustained failure), and a major incident (one that triggers customer-facing communication and post-mortem). Each has its own lifecycle, severity, and routing.
Incident records typically include: start and end times, affected components, severity, root cause, customer impact, and a post-mortem document for major incidents.
Treating every alert as an incident drowns the on-call rotation. Treating no alerts as incidents leaves outages uninvestigated. The right balance — confirm-then-escalate logic, severity-based routing, and clear ownership — is the spine of operational reliability.
See it in the product: Incident management.