SciTech Pulse
AI

Constraint-aware GPU scheduler beats FIFO by up to 33 points on utilization, benchmark shows

A benchmark from Dharma-AI shows its constraint-aware GPU allocator lifts utilization by up to 33 percentage points over a FIFO scheduler on identical hardware, and more than doubles priority-weighted output in the…

Step by step

  1. 1

    Real-time demand read as a curve

  2. 2

    GPUs allocated per timestep, not peak

  3. 3

    Batch jobs fill the demand troughs

  4. 4

    Jobs ranked by priority, not arrival

  5. 5

    Utilization and priority output rise

Dharma-AI has published a benchmark comparing its constraint-aware GPU allocator against a FIFO — first in, first out — scheduler across seven workload scenarios. On identical hardware running the same jobs, the allocator raised by as much as 33 percentage points, and priority-weighted output rose in every scenario, by as much as 105 percent, the company reports.

A GPU cluster running enterprise AI has to serve four kinds of jobs at once — training, batch , model , and real-time inference. The first three are batch-like and need a block of GPUs held without interruption until they finish, while real-time inference is elastic and rises and falls with user traffic. Two shapes competing for the same hardware in the same window is the core of the problem.

According to the writeup, FIFO pays for the mismatch in two ways. To guarantee real-time capacity, it reserves each application's peak demand for the entire day, leaving GPUs idle through the off-peak hours. And because it places jobs in arrival order without weighing what a job is worth, high-priority work waits behind whichever request landed first.

The Dharma-AI allocator treats real-time demand as a curve rather than a ceiling, allocating against that demand timestep by timestep and letting batch-like jobs occupy the troughs. Batch jobs are ordered by priority across the full scheduling horizon rather than by arrival. In a training-heavy 8-GPU scenario, utilization rose from 53.6 percent to 87 percent and priority-weighted value more than doubled.

Across five contended scenarios, utilization moved from a 52–88 percent band under FIFO to a 72–88 percent band under the allocator, with priority-weighted gains averaging 52 percent, the company reports.

Terms explained

The story so far

  1. New MIT Method Predicts Extreme Storms Without Past Extreme Data
  2. Ordinary Wi-Fi Can Identify People With Near-Perfect Accuracy, Study Finds
  3. New AI Model Learns to Design Proteins Unlike Any in Nature
  4. DeepMind Runs First 'Double-Blind' Test of an AI Model to Stop Cheating on Benchmarks
  5. Machine-Learning Model Uses Routine Blood Tests to Flag a Common Type of Heart Failure
  6. New S-DEIM Method Estimates Sea Surface Temperatures 40% More Accurately
  7. Constraint-aware GPU scheduler beats FIFO by up to 33 points on utilization, benchmark shows
#ai infrastructure#gpu scheduling#dharma ai#benchmark#machine learning
Rate this story

Related stories