Monitor NVIDIA GPUs with OpenTelemetry
Use this integration to collect NVIDIA GPU hardware metrics exposed by NVIDIA DCGM Exporter. The Coralogix OpenTelemetry Collector scrapes the exporter over HTTP and sends the metrics through the existing metrics pipeline.
This guide covers hardware telemetry, including GPU utilization, framebuffer memory, temperature, power, clocks, and hardware errors.
Prerequisites
- An NVIDIA GPU with a supported NVIDIA driver.
- NVIDIA DCGM Exporter running on the GPU host or Kubernetes cluster. The NVIDIA GPU Operator can deploy it.
- The Coralogix OpenTelemetry Integration Helm chart installed in the cluster.
- Network access from the OpenTelemetry Collector to DCGM Exporter's
/metricsendpoint, normally port9400.
The Collector only scrapes HTTP metrics. Do not mount GPU devices or grant NVIDIA capabilities to the Collector.
Schedule DCGM Exporter only on GPU nodes. Configure the node selector, affinity, and tolerations in the NVIDIA DCGM Exporter chart values; do not add them to the Coralogix OpenTelemetry Integration values.
AWS EKS
For EKS, use a supported GPU node type and an EKS-optimized NVIDIA accelerated AMI. When the AMI already provides the NVIDIA driver and container toolkit, configure GPU Operator not to install either component again. See EKS-optimized accelerated AMIs.
Choose one GPU metric collector. Do not enable this scrape together with AWS CloudWatch Container Insights GPU monitoring, which deploys and manages DCGM collection itself. Running both creates duplicate metrics and ambiguous troubleshooting.
Verify DCGM Exporter
First verify the exporter before changing the Collector configuration. Replace the namespace and Service name with values from your installation.
kubectl -n <gpu-namespace> port-forward service/<dcgm-exporter-service> 9400:9400
curl -s http://localhost:9400/metrics | grep '^DCGM_FI_' | head
The command must return one or more DCGM_FI_* metrics. If it does not, fix the GPU driver or DCGM Exporter before configuring the Collector.
Configure Target Allocator discovery
Use the integration's Target Allocator with a Prometheus Operator ServiceMonitor.
Target Allocator requires the Prometheus Operator ServiceMonitor CRD. Enable it in the values file used to install the Coralogix OpenTelemetry Integration:
opentelemetry-agent:
targetAllocator:
enabled: true
allocationStrategy: per-node
prometheusCR:
enabled: true
The NVIDIA DCGM Exporter chart creates its own ServiceMonitor by default. Do not create a second ServiceMonitor for the same exporter. Configure its interval in the NVIDIA DCGM Exporter chart values file:
serviceMonitor:
enabled: true
Target Allocator discovers this ServiceMonitor, resolves its endpoints, and assigns each target to one agent. With the per-node strategy, an exporter endpoint is assigned to the agent on the same node.
Validate collection
Confirm the Target Allocator was enabled in the rendered agent configuration:
helm template <release> coralogix-charts-virtual/otel-integration \
--namespace <otel-namespace> \
--values values.yaml | grep -n 'targetAllocator'
Check the Target Allocator's assigned targets and then search Coralogix Metrics Explorer for DCGM_FI_DEV_GPU_UTIL or another metric returned by the exporter:
kubectl -n <otel-namespace> port-forward service/<target-allocator-service> 8080:8080
curl -s http://localhost:8080/jobs
curl -s http://localhost:8080/scrape_configs
Each DCGM exporter endpoint must be assigned once. Inspect agent logs for scrape failures.
Metrics and labels
The integration preserves raw DCGM metric names and labels for compatibility. Start with the metrics DCGM Exporter exposes by default. High-value metrics include:
DCGM_FI_DEV_GPU_UTILandDCGM_FI_DEV_MEM_COPY_UTILDCGM_FI_DEV_FB_USEDandDCGM_FI_DEV_FB_FREEDCGM_FI_DEV_GPU_TEMPandDCGM_FI_DEV_POWER_USAGEDCGM_FI_DEV_TOTAL_ENERGY_CONSUMPTIONDCGM_FI_DEV_XID_ERRORS
Start with the exporter defaults. Kubernetes pod labels and pod UIDs are disabled by default; keep them disabled unless you need workload correlation. The following is an NVIDIA DCGM Exporter chart value, not a Coralogix OpenTelemetry Integration value. Add it at the root of the values file used to install or upgrade gpu-helm-charts/dcgm-exporter:
kubernetes:
enablePodLabels: true
enablePodUID: false
podLabelAllowlistRegex:
- '^app.kubernetes.io/name$'
- '^environment$'
Keep enablePodUID: false: a pod UID creates a new series when a workload restarts. This guide does not enable process-level metrics or add a custom DCGM counter file. Review PID, process, and extra profiling dimensions before enabling them because they can significantly increase metric cardinality and ingestion cost.