Skip to content

Device Telemetry (OTLP Ingest)

Fleet Manager accepts device metrics and logs over the industry-standard OpenTelemetry Protocol (OTLP). A device (or a local collector/agent running alongside it) exports directly to Fleet Manager using its existing device identity — no separate telemetry account or credential is needed.

Who this is for: integrators configuring an edge device’s OpenTelemetry exporter, or anyone evaluating what device observability data Fleet Manager can ingest.

This is a Fleet Manager–native API, not part of Margo. The Margo section of this site covers device onboarding, capability reporting and workload orchestration — telemetry is a separate, independent surface with its own host and its own routes. The one thing the two share is the device certificate: if your device has already completed Margo onboarding, it already holds everything it needs to send telemetry too.

https://telemetry.<your-environment-domain>/<tenantSlug>/api/v1/telemetry

The telemetry. host is a sibling of your Fleet Manager application URL, on the same base domain — not a separate domain to look up. If you access Fleet Manager at https://fleet-manager.example.com, telemetry ingest is at https://telemetry.example.com. Ask your administrator if you’re unsure of your organization’s exact domain.

Configure this as the base endpoint in a standard OTLP/HTTP exporter — it appends /v1/metrics and /v1/logs itself:

OTEL_EXPORTER_OTLP_ENDPOINT=https://telemetry.example.com/acme-corp/api/v1/telemetry

<tenantSlug> is your organization’s slug — the same one used for device onboarding.

On-premises/OpenShift deployments: telemetry ingest is off by default and must be explicitly enabled by the platform administrator. If your Fleet Manager instance runs on-premises and telemetry data isn’t arriving, confirm with your administrator that ingest has been enabled and that a TLS certificate has been provisioned for the telemetry. host — this step is manual on OpenShift and doesn’t happen automatically.

Telemetry ingest uses mutual TLS with the device’s existing certificate — the same certificate and private key issued during device onboarding. There is no separate telemetry credential to provision or rotate.

  • The server presents a publicly trusted certificate (validate against your normal public trust store — this is not the Fleet Manager Root CA used for onboarding).
  • TLS 1.3 minimum; SNI is required.
  • A certificate that has been revoked is rejected immediately on the next request — there is no delay for cache expiry beyond a few minutes.

Both standard OTLP encodings are accepted — send whichever your exporter uses by default:

  • Content-Type: application/json — OTLP/JSON
  • Content-Type: application/x-protobuf — OTLP/protobuf (the default for most OpenTelemetry SDKs and collector distributions)

Metrics: only gauge and sum data points are stored today; histogram, exponentialHistogram and summary are accepted but silently dropped. Prefer gauges and sums for anything you need visible in the dashboard.

Logs: the log body must be a plain string (stringValue); structured bodies are not read. Attach service.name as a log-record or data-point attribute (not a resource attribute) if you want it to appear as the log source in the Devices view.

Fleet Manager reads a specific, small set of OTel metric names out of whatever your collector sends — anything else is accepted (not an error) but has no effect on the dashboard. This table is the exact interop contract; every name below was verified against a real OTel Collector Contrib payload or its receiver documentation. An earlier revision of this integration expected two metric names (process.runtime.restart_count, container.running/container.failed) that no OTel receiver has ever emitted — if you’re configuring a new integration, use this table, not older guidance you may have seen.

Several of these metrics are reported as multiple data points per collection cycle, one per state attribute value (e.g. system.cpu. utilization reports user/system/idle/nice/interrupt/softirq/ steal/wait together, summing to ~1.0). Fleet Manager reads the specific state named below and ignores the others for that metric — sending the metric without a state attribute at all means it is dropped, not guessed-at.

Fleet Manager field OTel metric Required state Enabled by default?
CPU usage % system.cpu.utilization (hostmetrics cpu scraper) none read directly — derived as 100 - state="idle" Yes
Memory used (bytes) system.memory.usage (hostmetrics memory scraper) used Yes
Memory total (bytes) system.memory.limit, or derived from system.memory.utilization{state="used"} if system.memory.limit isn’t sent n/a (limit) / used (utilization) No (system.memory.limit is opt-in — see below)
Disk used (bytes) system.filesystem.usage (hostmetrics filesystem scraper) used Yes
Disk total (bytes) Derived from system.filesystem.utilization n/a — this metric carries no state attribute No — there is no system.filesystem.limit metric; enable system.filesystem.utilization if you want disk %
Uptime (seconds) system.uptime (hostmetrics system scraper) n/a Yes
Container restarts container.restarts (docker_stats receiver) — summed across every container reporting n/a No — optional metric, enable it explicitly
Containers running container.state.status{state="running"} (docker_stats receiver) — summed running No — optional metric, enable it explicitly
Containers failed container.state.status{state="exited"} or {state="dead"} (docker_stats receiver) — summed exited / dead No — optional metric, enable it explicitly
Error count (not a metric) — counted server-side from log records at ERROR severity in the same window — n/a — send logs if you want this populated

Not read at all: system.cpu.usage (a cumulative CPU-seconds counter, not a ratio — converting it to a percentage requires a rate-over-time calculation this integration does not perform; send system.cpu.utilization instead, which the cpu scraper already computes for you).

Minimal collector config to populate every field above, assuming the hostmetrics and docker_stats receivers:

receivers:
hostmetrics:
collection_interval: 30s
scrapers:
cpu:
metrics:
system.cpu.utilization:
enabled: true
memory:
metrics:
system.memory.utilization:
enabled: true # populates memory total without needing system.memory.limit
filesystem:
metrics:
system.filesystem.utilization:
enabled: true # the only way to get a disk total — there is no "limit" metric
system: {} # system.uptime is on by default
docker_stats:
metrics:
container.restarts:
enabled: true
container.state.status:
enabled: true

(Exact YAML keys depend on your collector distribution/version — check your receiver’s own documentation.md if a key above doesn’t match.)

To protect the shared platform, ingest is rate-limited per device and per organization, and each request is capped in size and record count:

Limit Default
Request body size 4 MB
Requests per device 300 per 60 seconds
Requests per organization 3000 per 60 seconds
Metric data points per request 10000
Log records per request 500

A rate-limit breach returns 429 with a Retry-After header — back off and retry. Exceeding a per-request record cap does not fail the whole batch: the records up to the cap are accepted and stored, and the response reports how many were dropped so you can reduce your batch size.

If your fleet’s expected export volume is close to either rate limit, contact your Fleet Manager administrator — these limits are configurable per deployment.

Recommended: batch on a 10–60 second interval and apply exponential backoff with jitter on 429 and 5xx responses, regardless of the defaults above.

A fully accepted request returns 200 with an empty JSON body. A request where some records exceeded the per-request cap still returns 200, with an OTLP partial-success body reporting what was dropped.

Status Meaning What to do
200 Accepted (fully, or partially — check the response) Continue; reduce batch size if partial
400 Malformed request body Fix the payload — retrying unchanged will not help
401 No certificate presented, or it is unknown/revoked Check device onboarding/certificate status
403 Tenant slug does not match the certificate’s org Fix the configured slug
404 Unknown or inactive tenant slug Fix the configured slug
413 Request body too large Reduce batch size
415 Unsupported Content-Type Use application/json or application/x-protobuf
429 Rate limited Back off (honor Retry-After) and retry

Once telemetry is flowing, view it per device under Devices → Health and in the Dashboard’s fleet health metrics.