Device Telemetry (OTLP Ingest)
Fleet Manager accepts device metrics and logs over the industry-standard OpenTelemetry Protocol (OTLP). A device (or a local collector/agent running alongside it) exports directly to Fleet Manager using its existing device identity — no separate telemetry account or credential is needed.
Who this is for: integrators configuring an edge device’s OpenTelemetry exporter, or anyone evaluating what device observability data Fleet Manager can ingest.
This is a Fleet Manager–native API, not part of Margo. The Margo section of this site covers device onboarding, capability reporting and workload orchestration — telemetry is a separate, independent surface with its own host and its own routes. The one thing the two share is the device certificate: if your device has already completed Margo onboarding, it already holds everything it needs to send telemetry too.
Endpoint
Section titled “Endpoint”https://telemetry.<your-environment-domain>/<tenantSlug>/api/v1/telemetryThe telemetry. host is a sibling of your Fleet Manager application URL,
on the same base domain — not a separate domain to look up. If you access
Fleet Manager at https://fleet-manager.example.com, telemetry ingest is at
https://telemetry.example.com. Ask your administrator if you’re unsure of
your organization’s exact domain.
Configure this as the base endpoint in a standard OTLP/HTTP exporter — it
appends /v1/metrics and /v1/logs itself:
OTEL_EXPORTER_OTLP_ENDPOINT=https://telemetry.example.com/acme-corp/api/v1/telemetry<tenantSlug> is your organization’s slug — the same one used for device
onboarding.
On-premises/OpenShift deployments: telemetry ingest is off by default
and must be explicitly enabled by the platform administrator. If your
Fleet Manager instance runs on-premises and telemetry data isn’t arriving,
confirm with your administrator that ingest has been enabled and that a TLS
certificate has been provisioned for the telemetry. host — this step is
manual on OpenShift and doesn’t happen automatically.
Authentication
Section titled “Authentication”Telemetry ingest uses mutual TLS with the device’s existing certificate — the same certificate and private key issued during device onboarding. There is no separate telemetry credential to provision or rotate.
- The server presents a publicly trusted certificate (validate against your normal public trust store — this is not the Fleet Manager Root CA used for onboarding).
- TLS 1.3 minimum; SNI is required.
- A certificate that has been revoked is rejected immediately on the next request — there is no delay for cache expiry beyond a few minutes.
Payload format
Section titled “Payload format”Both standard OTLP encodings are accepted — send whichever your exporter uses by default:
Content-Type: application/json— OTLP/JSONContent-Type: application/x-protobuf— OTLP/protobuf (the default for most OpenTelemetry SDKs and collector distributions)
Metrics: only gauge and sum data points are stored today; histogram,
exponentialHistogram and summary are accepted but silently dropped. Prefer
gauges and sums for anything you need visible in the dashboard.
Logs: the log body must be a plain string (stringValue); structured
bodies are not read. Attach service.name as a log-record or data-point
attribute (not a resource attribute) if you want it to appear as the log
source in the Devices view.
Supported metrics
Section titled “Supported metrics”Fleet Manager reads a specific, small set of OTel metric names out of
whatever your collector sends — anything else is accepted (not an error) but
has no effect on the dashboard. This table is the exact interop contract;
every name below was verified against a real OTel Collector Contrib payload
or its receiver documentation. An earlier revision of this integration
expected two metric names (process.runtime.restart_count,
container.running/container.failed) that no OTel receiver has ever
emitted — if you’re configuring a new integration, use this table, not
older guidance you may have seen.
Several of these metrics are reported as multiple data points per
collection cycle, one per state attribute value (e.g. system.cpu. utilization reports user/system/idle/nice/interrupt/softirq/
steal/wait together, summing to ~1.0). Fleet Manager reads the specific
state named below and ignores the others for that metric — sending the
metric without a state attribute at all means it is dropped, not
guessed-at.
| Fleet Manager field | OTel metric | Required state |
Enabled by default? |
|---|---|---|---|
| CPU usage % | system.cpu.utilization (hostmetrics cpu scraper) |
none read directly — derived as 100 - state="idle" |
Yes |
| Memory used (bytes) | system.memory.usage (hostmetrics memory scraper) |
used |
Yes |
| Memory total (bytes) | system.memory.limit, or derived from system.memory.utilization{state="used"} if system.memory.limit isn’t sent |
n/a (limit) / used (utilization) |
No (system.memory.limit is opt-in — see below) |
| Disk used (bytes) | system.filesystem.usage (hostmetrics filesystem scraper) |
used |
Yes |
| Disk total (bytes) | Derived from system.filesystem.utilization |
n/a — this metric carries no state attribute |
No — there is no system.filesystem.limit metric; enable system.filesystem.utilization if you want disk % |
| Uptime (seconds) | system.uptime (hostmetrics system scraper) |
n/a | Yes |
| Container restarts | container.restarts (docker_stats receiver) — summed across every container reporting |
n/a | No — optional metric, enable it explicitly |
| Containers running | container.state.status{state="running"} (docker_stats receiver) — summed |
running |
No — optional metric, enable it explicitly |
| Containers failed | container.state.status{state="exited"} or {state="dead"} (docker_stats receiver) — summed |
exited / dead |
No — optional metric, enable it explicitly |
| Error count | (not a metric) — counted server-side from log records at ERROR severity in the same window |
— | n/a — send logs if you want this populated |
Not read at all: system.cpu.usage (a cumulative CPU-seconds counter, not
a ratio — converting it to a percentage requires a rate-over-time
calculation this integration does not perform; send system.cpu.utilization
instead, which the cpu scraper already computes for you).
Minimal collector config to populate every field above, assuming the
hostmetrics and docker_stats receivers:
receivers: hostmetrics: collection_interval: 30s scrapers: cpu: metrics: system.cpu.utilization: enabled: true memory: metrics: system.memory.utilization: enabled: true # populates memory total without needing system.memory.limit filesystem: metrics: system.filesystem.utilization: enabled: true # the only way to get a disk total — there is no "limit" metric system: {} # system.uptime is on by default docker_stats: metrics: container.restarts: enabled: true container.state.status: enabled: true(Exact YAML keys depend on your collector distribution/version — check your
receiver’s own documentation.md if a key above doesn’t match.)
Rate limits and payload caps
Section titled “Rate limits and payload caps”To protect the shared platform, ingest is rate-limited per device and per organization, and each request is capped in size and record count:
| Limit | Default |
|---|---|
| Request body size | 4 MB |
| Requests per device | 300 per 60 seconds |
| Requests per organization | 3000 per 60 seconds |
| Metric data points per request | 10000 |
| Log records per request | 500 |
A rate-limit breach returns 429 with a Retry-After header — back off and
retry. Exceeding a per-request record cap does not fail the whole batch: the
records up to the cap are accepted and stored, and the response reports how
many were dropped so you can reduce your batch size.
If your fleet’s expected export volume is close to either rate limit, contact your Fleet Manager administrator — these limits are configurable per deployment.
Recommended: batch on a 10–60 second interval and apply exponential
backoff with jitter on 429 and 5xx responses, regardless of the defaults
above.
Responses and errors
Section titled “Responses and errors”A fully accepted request returns 200 with an empty JSON body. A request
where some records exceeded the per-request cap still returns 200, with an
OTLP partial-success body reporting what was dropped.
| Status | Meaning | What to do |
|---|---|---|
| 200 | Accepted (fully, or partially — check the response) | Continue; reduce batch size if partial |
| 400 | Malformed request body | Fix the payload — retrying unchanged will not help |
| 401 | No certificate presented, or it is unknown/revoked | Check device onboarding/certificate status |
| 403 | Tenant slug does not match the certificate’s org | Fix the configured slug |
| 404 | Unknown or inactive tenant slug | Fix the configured slug |
| 413 | Request body too large | Reduce batch size |
| 415 | Unsupported Content-Type |
Use application/json or application/x-protobuf |
| 429 | Rate limited | Back off (honor Retry-After) and retry |
Once telemetry is flowing, view it per device under Devices → Health and in the Dashboard’s fleet health metrics.