Skip to main content

Monitor NVIDIA GPUs with OpenTelemetry

Use this integration to collect NVIDIA GPU hardware metrics exposed by NVIDIA DCGM Exporter. The Coralogix OpenTelemetry Collector scrapes the exporter over HTTP and sends the metrics through the existing metrics pipeline.

This guide covers hardware telemetry, including GPU utilization, framebuffer memory, temperature, power, clocks, and hardware errors.

Prerequisites​

  • An NVIDIA GPU with a supported NVIDIA driver.
  • NVIDIA DCGM Exporter running on the GPU host or Kubernetes cluster. The NVIDIA GPU Operator can deploy it.
  • The Coralogix OpenTelemetry Integration Helm chart installed in the cluster.
  • Network access from the OpenTelemetry Collector to DCGM Exporter's /metrics endpoint, normally port 9400.

The Collector only scrapes HTTP metrics. Do not mount GPU devices or grant NVIDIA capabilities to the Collector.

Schedule DCGM Exporter only on GPU nodes. Configure the node selector, affinity, and tolerations in the NVIDIA DCGM Exporter chart values; do not add them to the Coralogix OpenTelemetry Integration values.

AWS EKS​

For EKS, use a supported GPU node type and an EKS-optimized NVIDIA accelerated AMI. When the AMI already provides the NVIDIA driver and container toolkit, configure GPU Operator not to install either component again. See EKS-optimized accelerated AMIs.

Warning

Choose one GPU metric collector. Do not enable this scrape together with AWS CloudWatch Container Insights GPU monitoring, which deploys and manages DCGM collection itself. Running both creates duplicate metrics and ambiguous troubleshooting.

Verify DCGM Exporter​

First verify the exporter before changing the Collector configuration. Replace the namespace and Service name with values from your installation.

kubectl -n <gpu-namespace> port-forward service/<dcgm-exporter-service> 9400:9400
curl -s http://localhost:9400/metrics | grep '^DCGM_FI_' | head

The command must return one or more DCGM_FI_* metrics. If it does not, fix the GPU driver or DCGM Exporter before configuring the Collector.

Configure Target Allocator discovery​

Use the integration's Target Allocator with a Prometheus Operator ServiceMonitor.

Target Allocator requires the Prometheus Operator ServiceMonitor CRD. Enable it in the values file used to install the Coralogix OpenTelemetry Integration:

opentelemetry-agent:
targetAllocator:
enabled: true
allocationStrategy: per-node
prometheusCR:
enabled: true

The NVIDIA DCGM Exporter chart creates its own ServiceMonitor by default. Do not create a second ServiceMonitor for the same exporter. Configure its interval in the NVIDIA DCGM Exporter chart values file:

serviceMonitor:
enabled: true

Target Allocator discovers this ServiceMonitor, resolves its endpoints, and assigns each target to one agent. With the per-node strategy, an exporter endpoint is assigned to the agent on the same node.

Validate collection​

Confirm the Target Allocator was enabled in the rendered agent configuration:

helm template <release> coralogix-charts-virtual/otel-integration \
--namespace <otel-namespace> \
--values values.yaml | grep -n 'targetAllocator'

Check the Target Allocator's assigned targets and then search Coralogix Metrics Explorer for DCGM_FI_DEV_GPU_UTIL or another metric returned by the exporter:

kubectl -n <otel-namespace> port-forward service/<target-allocator-service> 8080:8080
curl -s http://localhost:8080/jobs
curl -s http://localhost:8080/scrape_configs

Each DCGM exporter endpoint must be assigned once. Inspect agent logs for scrape failures.

Metrics and labels​

The integration preserves raw DCGM metric names and labels for compatibility. Start with the metrics DCGM Exporter exposes by default. High-value metrics include:

  • DCGM_FI_DEV_GPU_UTIL and DCGM_FI_DEV_MEM_COPY_UTIL
  • DCGM_FI_DEV_FB_USED and DCGM_FI_DEV_FB_FREE
  • DCGM_FI_DEV_GPU_TEMP and DCGM_FI_DEV_POWER_USAGE
  • DCGM_FI_DEV_TOTAL_ENERGY_CONSUMPTION
  • DCGM_FI_DEV_XID_ERRORS

Start with the exporter defaults. Kubernetes pod labels and pod UIDs are disabled by default; keep them disabled unless you need workload correlation. The following is an NVIDIA DCGM Exporter chart value, not a Coralogix OpenTelemetry Integration value. Add it at the root of the values file used to install or upgrade gpu-helm-charts/dcgm-exporter:

kubernetes:
enablePodLabels: true
enablePodUID: false
podLabelAllowlistRegex:
- '^app.kubernetes.io/name$'
- '^environment$'

Keep enablePodUID: false: a pod UID creates a new series when a workload restarts. This guide does not enable process-level metrics or add a custom DCGM counter file. Review PID, process, and extra profiling dimensions before enabling them because they can significantly increase metric cardinality and ingestion cost.

Last updated on