Skip to content

Alerts

The Alerts module is where you decide what should notify you, and where those notifications go.

Who can access this: admins and operators. Both can create, edit and delete alert rules and notification targets.

An alert has two halves, and you configure them separately:

  • an alert rule — the condition that fires, and which devices it watches;
  • one or more notification targets — an email address or a webhook endpoint the alert is delivered to.

A rule with no targets is evaluated but notifies nobody, so set up a target first.

Open the Notification targets tab and click Add target.

Give the target a name and an email address. Nothing else is required — the platform uses the mail relay your deployment is already configured with.

Give the target a name and an HTTPS URL. The platform will POST a JSON body to that URL whenever an alert fires or resolves. See Webhook payload below for the exact shape.

When you save a webhook target, the platform generates a signing secret and shows it to you once, in a dialog.

The URL must be publicly reachable. Addresses that point back into private infrastructure — localhost, loopback and link-local addresses, private ranges such as 10.x or 192.168.x, cloud metadata endpoints, and internal cluster hostnames — are rejected when you save. Fleet Manager fetches this URL from inside its own network, so allowing those would let a webhook target reach services that are not meant to be reachable from outside.

The channel cannot be changed after a target is created. Create a new target instead — switching an email target to a webhook would leave every alert already recorded against it claiming to have been delivered over a transport that never ran.

Open the Rules tab and click Create rule. A rule has four parts.

Type Fires when
Device offline A device stops reporting for longer than the threshold you set.
Metric threshold A reported metric crosses the threshold in your expression.
Workload crash A workload on the device reports failed or repeatedly restarting containers.

The rule type is fixed once the rule is created.

For Device offline, set Minutes offline before firing. The default is five minutes.

For the other two types, write an expression. The field checks your syntax as you type and marks the exact character it could not read, so you find out before saving rather than after. See Expression syntax.

Scope Applies to
All devices Every device in your tenant.
By label Devices carrying every label key and value you list.
Single device One device, by its ID.

By label uses the same label matching as Groups and deployment targeting: a device matches only if it carries all the pairs you list.

Tick the targets this rule should notify. A rule can notify several, and each one receives its own copy.

Click Send test on any rule to dispatch a test notification through every target wired to it. The result tells you, per target, whether delivery succeeded — so a single misconfigured endpoint is identified by name rather than hidden behind a generic failure.

A test never counts as a real alert: it creates no history entry and cannot suppress a genuine alert that fires afterwards.

The History tab lists every alert occurrence: which rule fired, on which device, when it fired and when it resolved, and whether the notification was actually delivered.

Alert state and delivery outcome are shown separately on purpose. “The alert fired” and “somebody received it” are different facts, and the case worth catching — a correct alert whose endpoint was unreachable — is only visible when the two are not merged.

An alert stays Open while its condition holds and becomes Resolved when the condition clears. You receive one notification per occurrence, not one per check — a device offline for an hour notifies you once, and a second, separate outage notifies you again.

Expressions compare a field against a value, optionally combined with AND and optionally required to hold for a duration:

field operator value [AND field operator value ...] [for duration]

Examples:

cpu > 80
memory >= 90 AND disk > 75
cpu > 80 AND device.label.tier = "edge" for 5m
containersFailed > 0
device.status = OFFLINE

Operators: >, >=, <, <=, =, !=. The ordering operators need a numeric value.

Durations: 30s, 5m, 2h. With a for clause the condition must hold across the whole window — a single momentary spike does not fire the rule.

Fields:

Field Meaning
cpu CPU usage, percent
memory Memory used, percent of total
disk Disk used, percent of total
restarts Container restart count
containersFailed Number of failed containers
containersRunning Number of running containers
errors Reported error count
uptime Device uptime, seconds
device.status ONLINE, OFFLINE, PENDING or ERROR
device.name Device name
device.label.<key> The value of a device label

A webhook target receives a POST with Content-Type: application/json and this body:

{
"id": "e9a1c2f0-5b3d-4c8a-9f21-7d6e4b8a1c33",
"type": "alert.triggered",
"timestamp": "2026-09-29T10:00:00.000Z",
"data": {
"ruleId": "b2c3d4e5-6f70-4a81-9b2c-3d4e5f607182",
"ruleName": "CPU above 80% on edge nodes",
"ruleType": "METRIC_THRESHOLD",
"condition": "cpu > 80 for 5m",
"organizationId": "1a2b3c4d-5e6f-4071-8293-a4b5c6d7e8f9",
"deviceId": "7f8e9d0c-1b2a-4394-8576-1c2d3e4f5061",
"deviceName": "edge-berlin-01",
"metricValue": 93.5,
"firedAt": "2026-09-29T10:00:00.000Z",
"deliveryId": "3c4d5e6f-7081-4192-a3b4-c5d6e7f80912"
}
}

type is alert.triggered when the alert fires and alert.resolved when it clears; a resolution also carries resolvedAt. metricValue is omitted entirely for rule types that have no metric, so absent and zero remain distinguishable. deliveryId is stable across the fire and resolve messages of the same occurrence, so you can correlate the pair.

Header Value
Content-Type application/json
X-Webhook-Id Unique per delivery — use it to discard duplicates
X-Webhook-Signature sha256=:<hex>: — HMAC-SHA256 of the raw body, keyed by the secret

Verify the signature by computing HMAC-SHA256 over the raw request body (before parsing it) with your target’s signing secret, and comparing it to the hex digest between the colons. Compare with a constant-time function.

A delivery is attempted up to three times with exponential backoff starting at one second, and each attempt times out after ten seconds. Fleet Manager retries on connection failures, 5xx responses and 429. It does not retry other 4xx responses — those say the request itself was unacceptable, and repeating it unchanged cannot help.

Respond 2xx as soon as you have accepted the payload, and do slow work afterwards. A receiver that holds the connection open while processing will hit the timeout and be retried.

  • View history — jump to the History tab filtered to that rule.
  • Send test — dispatch a test notification to the rule’s targets.
  • Edit — change the name, condition, scope, targets or enabled state.
  • Delete — removes the rule and its delivery history, after confirmation.