Skip to content

Monitoring

GitHub

Monitoring¶

April 3, 2025
in Monitoring, AMD, NVIDIA
2 min read

Built-in UI for monitoring essential GPU metrics

AI workloads generate vast amounts of metrics, making it essential to have efficient monitoring tools. While our recent update introduced the ability to export available metrics to Prometheus for maximum flexibility, there are times when users need to quickly access essential metrics without the need to switch to an external tool.

Previously, we introduced a CLI command that allows users to view essential GPU metrics for both NVIDIA and AMD hardware. Now, with this latest update, we’re excited to announce the addition of a built-in dashboard within the dstack control plane.

April 1, 2025
in Monitoring, NVIDIA
2 min read

Exporting GPU, cost, and other metrics to Prometheus

Effective AI infrastructure management requires full visibility into compute performance and costs. AI researchers need detailed insights into container- and GPU-level performance, while managers rely on cost metrics to track resource usage across projects.

While dstack provides key metrics through its UI and dstack metrics CLI, teams often need more granular data and prefer using their own monitoring tools. To support this, we’ve introduced a new endpoint that allows real-time exporting all collected metrics—covering fleets and runs—directly to Prometheus.

October 22, 2024
in AMD, NVIDIA, Monitoring
2 min read

Monitoring essential GPU metrics via CLI

While it's possible to use third-party monitoring tools with dstack, it is often more convenient to debug your run and track metrics out of the box. That's why, with the latest release, dstack introduced dstack stats, a new CLI (and API) for monitoring container metrics, including GPU usage for NVIDIA, AMD, and other accelerators.