Skip to main content

Understand alerts and delivery

Status: Upcoming — not yet available.

Section: DOC-IS-alerts#overview.

Catch rising failures or missing activity before they go unnoticed. Define the condition that needs attention, send it to the people who can act, and follow the incident through recovery. Alerts report what is happening; spending limits and safety controls enforce their own rules.

Configurable Insights alert rules, incident history and acknowledge/mute controls are not yet available. Existing notifications provide sending, inbox and preference operations; a notification destination alone does not create an alert rule.

1. Configure an actionable error-rate alert​

Status: Upcoming — not yet available.

Section: DOC-IS-alerts#alert-conditions.

Build an alert for a sustained rise in request errors, with a notice to the people who can fix it. Before starting, have a supported error signal, an authorized project/audience, a destination you control and an agreed action for the recipient.

  1. Choose the signal and population, then select a threshold and observation window suitable for that workload.
  2. Save the rule inactive and inspect its scope, missing-data handling and destination.
  3. Enable it after those choices are reviewed.
  4. Follow the resulting incident and its delivery attempt separately.
  5. During an outage, acknowledge or mute optional notices without hiding the incident; close the investigation only after its recovery evidence is current.

A useful alert identifies the condition, the affected audience and the next action. It does not grant additional spending or prevent work by itself.

Choose a condition your team can act on, such as a rise in errors or a missing heartbeat. The rule catalog identifies the event sources and conditions supported by your deployment, including event-count thresholds, error rates and missing-heartbeat checks. Statistical conditions require enough baseline data before they can be evaluated.

Before enabling a supported rule, review these choices:

ChoiceWhat to specify
ScopeProject, test or live environment, and exact customer audience; unassigned is a specific audience
ConditionSupported event/metric, unit, threshold and observation window
Data qualityHow recent and complete observations must be before a result is trustworthy
ConfirmationHow long a breach or recovery must persist before changing incident state
DestinationAuthorized recipients and channel, with an owner who can fix delivery failures
RemindersWhen another notice is useful while the same incident remains active

For this example, scope the rule to one project’s live, unassigned customer audience and send an email to the operations inbox after the selected request-error threshold holds over a complete five-minute window. The threshold and window are policy choices for this recipe, not defaults, a latency guarantee or a claim that this signal is available today.

Suggested rules start inactive. After explicitly creating your first project, inspect its scope and destinations before enabling suggestions. A cross-project rule requires an explicit authorized project set. A project rule does not override a broader rule: independently configured matching rules can both notify their permitted audiences.

2. Follow the condition through recovery​

Status: Upcoming — not yet available.

Section: DOC-IS-alerts#incident-lifecycle.

Incident history separates the condition being observed from notification delivery. One incident can have an initial notice and several reminders.

State or eventMeaning
New incidentA supported condition met its breach confirmation rule
ReminderThe same incident remains active and its reminder period elapsed
RecoveryFresh, sufficiently complete observations met the recovery condition
Unknown or staleThe source or evaluation cannot establish the current condition
CancelledThe rule was disabled/deleted or authority ended; this does not prove recovery

Missing data must not appear as a healthy zero. For a heartbeat rule, distinguish “the monitored activity stopped” from “the observation pipeline stopped.” Check source freshness before interpreting silence. Total notification delay includes source delay, evaluation, confirmation and delivery; it is not an instantaneous enforcement guarantee.

A later breach after recovery creates a new incident. Keep its first occurrence and recovery separate from earlier incidents when measuring duration or frequency.

Example: After an error-rate breach, a missing observation window shows Unknown, not Recovered. Fresh observations that meet the recovery condition close that incident; a later breach opens a new one.

3. Recover a notice that did not arrive​

Status: Upcoming — not yet available.

Section: DOC-IS-alerts#alert-delivery.

Detection, queued delivery, provider acceptance and recipient visibility are different outcomes. When a notice is missing, check the destination, recipient permissions, delivery failure and any declared suppression before concluding that the condition never fired. A test or dry run must say whether it sends a real message or contacts a paid provider.

Send notices through SMTP or webhook destinations in a self-hosted deployment, or choose an optional hosted destination. Delivery history shows pending notices and missing intervals after an interruption, so you can distinguish an undelivered notice from a condition that never fired.

Example: An incident can show Detected, an email attempt Failed, and another destination Delivered. Retrying delivery remains attached to the same incident.

Variant: reduce reminders while the team investigates​

Status: Upcoming — not yet available.

Section: DOC-IS-alerts#acknowledge-and-mute.

Acknowledge an incident to record that it has your attention, or mute optional notices for a limited time. Neither action changes the observed facts. A personal mute applies to your delivery only; shared suppression requires authority over the entire rule audience. Required notice classes remain identified and follow their own delivery policy.

Before applying suppression, check the incident, audience and expiry. Muting leaves evaluation running and the incident visible to other authorized operators. Existing notification preferences are separate controls; do not assume they implement these future alert semantics.

Example: Mute optional notices for yourself until a chosen time while a colleague continues to receive them and both can see the incident. Shared suppression requires permission for the whole audience.

Status: Upcoming — not yet available.

Section: DOC-IS-alerts#related-workflows.

  • Use notification sending and the inbox for their documented, deployment-supported tasks.
  • Read usage and cost before choosing a threshold based on sampled, estimated or delayed reporting.
  • Use evaluation to investigate permitted quality evidence behind an incident.

Document ID: DOC-IS-alerts. Section identities and revisions.