Copy as Markdown[Open in ChatGPT](https://chatgpt.com/?q=Read%20https%3A%2F%2Fcoralogix.com%2Fdocs%2Fopentelemetry%2Fintegrations%2Fnvidia-gpu-monitoring.md%20and%20help%20me%20with%20my%20question%20about%20this%20Coralogix%20documentation%20page.)[Open in Claude](https://claude.ai/new?q=Read%20https%3A%2F%2Fcoralogix.com%2Fdocs%2Fopentelemetry%2Fintegrations%2Fnvidia-gpu-monitoring.md%20and%20help%20me%20with%20my%20question%20about%20this%20Coralogix%20documentation%20page.)

# Monitor NVIDIA GPUs with OpenTelemetry

Use this integration to collect NVIDIA GPU hardware metrics exposed by [NVIDIA DCGM Exporter](https://docs.nvidia.com/datacenter/dcgm/latest/installation/install-dcgm-exporter.html). The Coralogix OpenTelemetry Collector scrapes the exporter over HTTP and sends the metrics through the existing metrics pipeline.

This guide covers hardware telemetry, including GPU utilization, framebuffer memory, temperature, power, clocks, and hardware errors.

## Prerequisites[​](#prerequisites "Direct link to Prerequisites")

* An NVIDIA GPU with a supported NVIDIA driver.
* NVIDIA DCGM Exporter running on the GPU host or Kubernetes cluster. The [NVIDIA GPU Operator](https://docs.nvidia.com/datacenter/cloud-native/gpu-operator/latest/index.html) can deploy it.
* The [Coralogix OpenTelemetry Integration Helm chart](https://coralogix.com/docs/opentelemetry/kubernetes-observability/kubernetes-complete-observability-basic-configuration.md) installed in the cluster.
* Network access from the OpenTelemetry Collector to DCGM Exporter's `/metrics` endpoint, normally port `9400`.

The Collector only scrapes HTTP metrics. Do not mount GPU devices or grant NVIDIA capabilities to the Collector.

Schedule DCGM Exporter only on GPU nodes. Configure the node selector, affinity, and tolerations in the NVIDIA DCGM Exporter chart values; do not add them to the Coralogix OpenTelemetry Integration values.

### AWS EKS[​](#aws-eks "Direct link to AWS EKS")

For EKS, use a supported GPU node type and an EKS-optimized NVIDIA accelerated AMI. When the AMI already provides the NVIDIA driver and container toolkit, configure GPU Operator not to install either component again. See [EKS-optimized accelerated AMIs](https://docs.aws.amazon.com/eks/latest/userguide/ml-eks-optimized-ami.html).

Warning

Choose one GPU metric collector. Do not enable this scrape together with AWS CloudWatch Container Insights GPU monitoring, which deploys and manages DCGM collection itself. Running both creates duplicate metrics and ambiguous troubleshooting.

## Verify DCGM Exporter[​](#verify-dcgm-exporter "Direct link to Verify DCGM Exporter")

First verify the exporter before changing the Collector configuration. Replace the namespace and Service name with values from your installation.

```
kubectl -n <gpu-namespace> port-forward service/<dcgm-exporter-service> 9400:9400

curl -s http://localhost:9400/metrics | grep '^DCGM_FI_' | head
```

The command must return one or more `DCGM_FI_*` metrics. If it does not, fix the GPU driver or DCGM Exporter before configuring the Collector.

## Configure Target Allocator discovery[​](#configure-target-allocator-discovery "Direct link to Configure Target Allocator discovery")

Use the integration's Target Allocator with a Prometheus Operator `ServiceMonitor`.

Target Allocator requires the Prometheus Operator `ServiceMonitor` CRD. Enable it in the values file used to install the Coralogix OpenTelemetry Integration:

```
opentelemetry-agent:

  targetAllocator:

    enabled: true

    allocationStrategy: per-node

    prometheusCR:

      enabled: true
```

The NVIDIA DCGM Exporter chart creates its own `ServiceMonitor` by default. Do not create a second `ServiceMonitor` for the same exporter. Configure its interval in the NVIDIA DCGM Exporter chart values file:

```
serviceMonitor:

  enabled: true
```

Target Allocator discovers this `ServiceMonitor`, resolves its endpoints, and assigns each target to one agent. With the `per-node` strategy, an exporter endpoint is assigned to the agent on the same node.

## Validate collection[​](#validate-collection "Direct link to Validate collection")

Confirm the Target Allocator was enabled in the rendered agent configuration:

```
helm template <release> coralogix-charts-virtual/otel-integration \

  --namespace <otel-namespace> \

  --values values.yaml | grep -n 'targetAllocator'
```

Check the Target Allocator's assigned targets and then search Coralogix Metrics Explorer for `DCGM_FI_DEV_GPU_UTIL` or another metric returned by the exporter:

```
kubectl -n <otel-namespace> port-forward service/<target-allocator-service> 8080:8080

curl -s http://localhost:8080/jobs

curl -s http://localhost:8080/scrape_configs
```

Each DCGM exporter endpoint must be assigned once. Inspect agent logs for scrape failures.

## Metrics and labels[​](#metrics-and-labels "Direct link to Metrics and labels")

The integration preserves raw DCGM metric names and labels for compatibility. Start with the metrics DCGM Exporter exposes by default. High-value metrics include:

* `DCGM_FI_DEV_GPU_UTIL` and `DCGM_FI_DEV_MEM_COPY_UTIL`
* `DCGM_FI_DEV_FB_USED` and `DCGM_FI_DEV_FB_FREE`
* `DCGM_FI_DEV_GPU_TEMP` and `DCGM_FI_DEV_POWER_USAGE`
* `DCGM_FI_DEV_TOTAL_ENERGY_CONSUMPTION`
* `DCGM_FI_DEV_XID_ERRORS`

Start with the exporter defaults. Kubernetes pod labels and pod UIDs are disabled by default; keep them disabled unless you need workload correlation. The following is an NVIDIA DCGM Exporter chart value, not a Coralogix OpenTelemetry Integration value. Add it at the root of the values file used to install or upgrade `gpu-helm-charts/dcgm-exporter`:

```
kubernetes:

  enablePodLabels: true

  enablePodUID: false

  podLabelAllowlistRegex:

    - '^app.kubernetes.io/name$'

    - '^environment$'
```

Keep `enablePodUID: false`: a pod UID creates a new series when a workload restarts. This guide does not enable process-level metrics or add a custom DCGM counter file. Review PID, process, and extra profiling dimensions before enabling them because they can significantly increase metric cardinality and ingestion cost.
