Keyboard shortcuts

Press or to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Measuring the energy consumption of an AI model during training & inference

This page explains how to monitor an AI model deployed locally during training and inference.

This provides plugin configuration examples. Depending on your hardware, the required plugins may slightly vary. Refer to the page of each plugin to find the corresponding configuration to put in alumet-config.toml. The framework used to deploy the model (vLLM, Ollama …) does not impact the configuration.

The 3 key components in an AI workflow that need monitoring are the CPU, the RAM and the GPU (if you have one).

Energy consumption

For the CPU and RAM, you should enable the rapl plugin if it is supported.
If not supported, you should use energy-estimation-tdp, which estimates the energy consumption of your processor based on its TDP (Thermal Design Power).

For the GPU, you should enable either the nvml plugin for NVIDIA GPUs or amd-gpu for AMD GPUs.

Component usage

If you want to monitor each component’s usage, it is also possible!

CPU and RAM

For the CPU and RAM, the required plugin will vary depending on the environment:

  • If the code runs inside a Kubernetes pod, a OAR or a Slurm job, you should enable respecively the k8s, oar or slurm plugin.
  • For bare metal training & inference (as in : the code runs directly on your machine), you should enable the cgroups plugin.
  • For Grid’5000 clusters, you should use Kwollect-input plugin

The relevant metrics are cpu_time_delta, cpu_percent for CPU usage and memory_usage for RAM usage.

GPU

For the GPU, amd-gpu and nvml already provide different utilization metrics.
There are different usage metrics corresponding to different parts of the GPU:

MetricName (nvml)Name (amd-gpu)
GPU usagenvml_gpu_utilizationamd_gpu_activity_usage (attribute “graphic_core”)
SM (= compute units) usagenvml_sm_utilizationamd_gpu_process_compute_unit_occupancy
VRAM usagenvml_memory_utilizationamd_gpu_activity_usage (attribute “unified_memory_controller”)
VRAM allocationnvml_gpu_memory_infoamd_gpu_memory_usage (total) or amd_gpu_process_memory_usage (per process)

Configuration examples

Here are a few example of scenarios, and the corresponding Alumet plugins configuration:

ComponentGPU pluginCPU pluginInfra plugin
Scenario #1nvmlraplk8s
Scenario #2amd-gpuraplcgroups
Scenario #3amd-gpuenergy-estimation-tdpslurm

Scenario 1: x86 CPU + NVIDIA GPU + K8S

  • GPU: any NVIDIA GPU
  • CPU : any RAPL compatible CPU
  • Infra : Kubernetes
Alumet config file
[plugins.k8s]
k8s_api_url = "http://127.0.0.1:8080"
token_retrieval = "auto"
poll_interval = "5s"
annotate_foreign_measurements = false
annotate_containers = false

[plugins.rapl]
poll_interval = "1s"
flush_interval = "5s"
no_perf_events = false

[plugins.nvml]
poll_interval = "1s"
flush_interval = "5s"
skip_failed_devices = true
mode = "full"

Scenario 2: x86 CPU + AMD GPU

  • GPU: any AMD GPU
  • CPU : any RAPL compatible CPU
  • Bare metal (runs directly on your machine)
Alumet config file
[plugins.cgroups]
poll_interval = "5s"

[plugins.rapl]
poll_interval = "1s"
flush_interval = "5s"
no_perf_events = false

[plugins.amd-gpu]
poll_interval = "1s"
flush_interval = "5s"
skip_failed_devices = true

Scenario 3: other CPU + AMD GPU + Slurm

  • GPU: any AMD GPU
  • CPU : not RAPL compatible CPU
  • Infra : Slurm
Alumet config file
[plugins.slurm]
poll_interval = "1s"
ignore_non_jobs = true
jobs_monitoring_level = "job"
add_source_in_pause_state = false
annotate_foreign_measurements = false

[plugins.energy-estimation-tdp]
poll_interval = "30s"
tdp = 100.0
nb_vcpu = 1.0
nb_cpu = 1.0

[plugins.amd-gpu]
poll_interval = "1s"
flush_interval = "5s"
skip_failed_devices = true