Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

Connect to the GPU server and run its vendor’s monitoring tool there: use nvidia-smi for a quick NVIDIA check, or AMD SMI’s monitor command on a system where AMD SMI is installed. For ongoing NVIDIA monitoring across one or more hosts, run DCGM Exporter on each GPU node and have Prometheus scrape its metrics endpoint.

Check GPU telemetry interactively over a remote shell

The monitoring command runs on the GPU host; your remote shell is simply how you reach that host. Connect using your organization’s approved remote-access method, then run the command for the GPU vendor.

NVIDIA: start with nvidia-smi

  1. Connect to the server that has the GPU.
  2. Run nvidia-smi.
  3. Check the fields available on that machine for utilization and temperature. The exact output and fields can vary by GPU and software version.

NVIDIA identifies nvidia-smi as an initial diagnostic check for confirming that the driver discovers the GPUs. If it does not show the expected device, investigate GPU and driver discovery before troubleshooting an exporter. See NVIDIA’s DCGM Exporter installation and troubleshooting guide.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

AMD: use AMD SMI’s monitor options

On an AMD system with AMD SMI installed, its CLI monitor options include temperature in Celsius, graphics and memory utilization, and VRAM use. The --watch INTERVAL option repeats output at the interval you specify in seconds. For example, use --watch 1 for a one-second interval if that cadence suits your check. Consult the installed version’s CLI help for the exact command form and available options; AMD’s AMD SMI CLI documentation describes the monitor options.

#1 Best Overall
Thermal Grizzly WireView Pro II 12V-2x6 GPU Power Meter Normal
  • CHECK COMPATIBILITY BEFORE PURCHASE: This product is only compatible with specific models. Please review the Compatibility List in the A+ Content below before ordering to ensure your device/model is supported.
  • GPU POWER METER FOR 12V-2X6 CONNECTIONS – WireView Pro II monitors graphics-card power delivery directly at the GPU cable path.
  • HARDWARE-BASED MONITORING WITHOUT REQUIRED SOFTWARE – Shows key values directly on the display, with optional software use.
  • EXTENDED 2-YEAR WARRANTY - For qualifying damage to the 12VHPWR or 12V-2x6 connector, Thermal Grizzly provides repair or, if repair is not possible, an equivalent replacement
  • DESIGNED FOR ADDITIONAL PC SAFETY – Supports early detection of abnormal power behavior on compatible 12V-2x6 GPU setups.

Choose between a live check and retained metrics

Approach Useful for What it provides
Vendor CLI on the GPU host A quick check or short-lived observation Current telemetry in terminal output; it does not, by itself, create centrally retained history.
DCGM Exporter scraped by Prometheus Central collection for dashboards, historical views, or alerting Selected NVIDIA DCGM fields exposed in Prometheus format for a Prometheus server to collect.

AMD SMI’s documented CLI is a direct monitoring route. The cited AMD documentation does not establish an equivalent standardized Prometheus-exporter workflow, so do not assume the NVIDIA setup below applies to AMD GPUs.

Set up persistent NVIDIA monitoring with DCGM Exporter

DCGM Exporter converts selected DCGM telemetry fields to Prometheus exposition format. Its documented default listener is :9400, with metrics available at /metrics. The exporter runs on the GPU node; a Prometheus server on another host must be able to reach that node’s endpoint. See the DCGM Exporter command reference and installation guide.

Rank #2
Thermalright Trofeo Vision LCD AIO Display 9.16” PC Monitor
  • 9.16” Wide LCD Screen – Features a crisp 1920×480 resolution display, perfect for showcasing system stats, hardware performance, or personalized visuals inside your gaming PC.
  • Real-Time Hardware Monitoring – Easily track CPU/GPU temps, fan speed, memory usage, and more, giving you complete control of your system health at a glance.
  • TRCC Software with DIY Options – Includes Thermalright TRCC app with multiple preset themes and DIY customization, so you can design your own unique interface
  • Plug & Play USB-C Connection – Simple Type-C interface ensures quick setup and compatibility with most Windows systems, no complicated drivers required.
  • Compact & Stylish Build – At only L251 x W68 x H17 mm, this slim display fits seamlessly inside or outside your PC case, adding both function and aesthetic appeal for modders and enthusiasts.

Pick a deployment that matches how you manage the host

  • Host-managed server: NVIDIA documents a package-managed service or a standalone OCI container. These approaches fit a server managed directly rather than through Kubernetes.
  • Kubernetes: NVIDIA documents deployment as a Helm-managed DaemonSet on selected GPU nodes. If the NVIDIA GPU Operator already manages the GPU software stack, its lifecycle is another documented route.
  • More than one GPU node: Run an exporter on each node you want to monitor, and configure collection so Prometheus can reach each exporter.

Before deployment, check NVIDIA’s support guidance for the GPU, driver, DCGM, and exporter combination. NVIDIA pairs exporter releases with DCGM versions; a mismatched pairing may work, but is not tested or supported. For a separately managed host engine, the DCGM client library must be at least as new as that engine. That condition alone does not make an otherwise mismatched release pair supported.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Verify the endpoint locally, then from Prometheus

  1. On the GPU host, check that the exporter process or pod is running and that it listens on the expected port.
  2. On a host-managed installation, run curl --fail http://localhost:9400/metrics on the GPU node. A successful response should include metric names beginning with DCGM_.
  3. For Kubernetes, follow NVIDIA’s guide to forward the exporter service before checking the endpoint locally.
  4. Make sure the Prometheus server can reach the endpoint over the network, and configure it to scrape the GPU node’s address and port. A local curl check does not prove remote network access.

Choose the metrics and collection cadence deliberately

The exporter exposes selected fields, so confirm that the fields you need are supported for the deployed GPU and software combination. NVIDIA’s guide includes examples for GPU utilization and framebuffer memory in MiB, as well as a five-second watch-group example for GPU temperature and board power. Treat that five-second value as a documented example, not a universal setting.

The guide documents a default collection interval of 30,000 ms, which can be configured. Each HTTP scrape returns the latest cached sample; it is not a new, independently stored measurement. Prometheus retention and scrape configuration determine how collected time series are recorded. See NVIDIA’s DCGM Exporter guide for field selection and configuration details.

Troubleshoot missing or incomplete metrics

Work from the GPU host outward so you can distinguish a device-discovery problem from an exporter, field, or network problem.

Rank #4
Thermalright Trofeo Vision LCD AIO Display 9.16” PC Monitor
  • 9.16” Wide LCD Screen – Features a crisp 1920×480 resolution display, perfect for showcasing system stats, hardware performance, or personalized visuals inside your gaming PC.
  • Real-Time Hardware Monitoring – Easily track CPU/GPU temps, fan speed, memory usage, and more, giving you complete control of your system health at a glance.
  • TRCC Software with DIY Options – Includes Thermalright TRCC app with multiple preset themes and DIY customization, so you can design your own unique interface
  • Plug & Play USB-C Connection – Simple Type-C interface ensures quick setup and compatibility with most Windows systems, no complicated drivers required.
  • Compact & Stylish Build – At only L251 x W68 x H17 mm, this slim display fits seamlessly inside or outside your PC case, adding both function and aesthetic appeal for modders and enthusiasts.
  1. Confirm GPU discovery: Run nvidia-smi on the node. If the expected GPU is not visible there, address that before debugging DCGM Exporter.
  2. Check the exporter: Confirm its process or Kubernetes pod is running, then inspect its logs for startup or collection errors.
  3. Check the endpoint: Verify the listener and port, then test /metrics locally. If local access works but Prometheus cannot scrape it, investigate routing, firewall rules, and endpoint reachability.
  4. Check field support: If the endpoint responds but a metric is absent, verify that the selected field is supported for the GPU/entity and software version in use.
  5. Check versions and host-engine access: Confirm the documented exporter/DCGM pairing. If using a separately managed host engine, also check client-library compatibility and the required remote transport.
  6. Check profiling requirements: Some profiling metrics have GPU-support and SYS_ADMIN requirements. Do not assume they are available just because basic utilization or temperature metrics work.

NVIDIA’s troubleshooting sequence and compatibility notes are in its installation guide.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Protect the metrics endpoint

Do not expose port 9400 broadly just to make scraping convenient. Restrict access to the Prometheus server and other required operators. NVIDIA documents TLS or basic authentication through the exporter’s --web-config-file option; consult the command reference for configuration details.

Best Value
WOWNOVA 5" Computer Temp Monitor, Dynamic Theme Supported, ARGB PC Case Sensor Panel, IPS Type-C USB Mini Secondary Screen, CPU RAM HDD Data Monitor (Black)
  • 【Upgraded 5" with Self-developed Software】In response to some customers' needs for a larger computer temp monitor, we have developed this upgraded 5-inch pannel. The PC Temperature Display works great with our English version software. You can use this with our software as a "second monitor" to view computer's Temperature and usage of CPU, GPU ,RAM, FPS and HDD Data etc. More professional and occupy less resoures.
  • 【Dynamic Vedio Theme & Cool!!】There are a lot of cool and cute dynamic videos preset in it, and the temporary computer monitor supports customizing your own dynamic video theme. Attached 16G flash card allows you DIY more and a lots dynamic videos.
  • 【Just One USB & Great Viewing Angles】Our Computer Temp Monitor only needs the single USB-C cable so it can be mounted completely internally off a usb header without the need of a port on the GPU which is a huge plus to you. No HDMI required, no power required. Just One USB Type-C cable. IPS full view. 5inch panel screen. Display area: 1.93*2.91". Overall size: 2.17*3.35". Resolution: 800*480. Thickness: 0.39". Shell material: Aluminum Housing
  • 【Simple & Feature-rich】Image&video UI support. Customizable screen layout. Horizontal and vertial screen switching. Visual theme editor: drag the mouse arbitarily to realize your creativity. Energy saving & environmental protection. One-click operation, Auto-Start, turn off the screen automatically and Comfortable eye protection Brightness adjustment.
  • 【Continuously Updated Theme & Great Customer Service】We have professional artists and techie who continuously updated the images and videos theme. We respect and value each customer's product and service satisfaction. We want to offer you premium products for a Long-Lasting Experience. If any issue, please kindly contact us for a solution.

Review container privileges before production deployment, particularly when enabling profiling fields. Protect runtime sockets, debug dumps, and profiling endpoints: diagnostic data may reveal host, process, or workload information. The appropriate network and runtime controls depend on how the exporter is deployed.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.