NVIDIA NCA-AIIO: GPU Scheduling in Kubernetes
Kubernetes can schedule GPUs as extended resources, but a production AI platform needs more than a request for nvidia.com/gpu. Different workloads require different accelerator models, memory capacities, isolation levels, topology, runtime stacks, and sharing behavior. The scheduler can make a good placement decision only when those requirements are represented accurately.
This makes GPU scheduling a natural bridge between cloud-native infrastructure and AI infrastructure. Kubernetes provides the control-plane model, while NVIDIA operators and device plugins expose accelerator capabilities to that model. The current NCA-AIIO scope emphasizes infrastructure operations because keeping accelerators usable is an orchestration problem as much as a hardware problem.
The main design goal is to allocate scarce GPUs to workloads in a way that preserves service objectives and cluster efficiency. That requires admission rules, resource classes, sharing choices, queueing, topology awareness, and telemetry rather than a first-come, first-served collection of pod requests.
Expose GPUs as governed resources
The NVIDIA device plugin advertises GPU resources to Kubernetes so pods can request them. The GPU Operator can manage drivers, the container toolkit, device plugin, monitoring components, and related lifecycle tasks. Centralizing these components reduces configuration drift compared with hand-building each GPU node.
Node labels and feature discovery should describe capabilities that matter to placement: GPU family, memory class, MIG capability, network features, and possibly workload-specific pools. Labels become part of the scheduling contract, so they should be created by controlled automation rather than by users trying to force their own jobs onto preferred nodes.
Kubernetes events are often the fastest place to see why a pod is pending. Insufficient GPU resources, node selectors, taints, affinity, quotas, and admission policy can all prevent placement. Reading the scheduler story before inspecting application logs avoids troubleshooting a container that has never started.
Reserve whole GPUs when isolation matters
The simplest model gives a pod exclusive access to one or more GPUs. This is predictable and appropriate for jobs that need full memory, stable performance, or low interference. The drawback is fragmentation: a light workload can reserve a device while using only a fraction of its compute and memory.
Exclusive allocation works best when workload size is understood. If teams habitually request the maximum number of GPUs because queue time is unpredictable, the cluster can show low utilization despite being fully allocated. Request guidance and queue policy should make right-sizing easier than over-reserving.
Infrastructure bottlenecks should therefore include allocation efficiency. A platform may buy more accelerators when the real constraint is that existing GPUs are stranded behind oversized requests or incompatible placement rules.
Use MIG when hardware isolation fits the workload
Multi-Instance GPU partitions supported NVIDIA GPUs into isolated GPU instances with defined compute and memory resources. Kubernetes can advertise these instances as schedulable resources through the GPU Operator and device plugin. MIG is useful when several workloads need stronger isolation and predictable slices of a large accelerator.
Partitioning introduces capacity-planning decisions. A MIG profile that matches one workload may leave unusable fragments for another, and changing the partition layout can require operational coordination. Standardize a small number of profiles based on real demand instead of creating a unique shape for every team.
GPU topology still matters with MIG. Instances share the physical placement and network environment of the parent GPU, so multi-node jobs need the same attention to NIC locality, rack placement, and communication paths as full-GPU workloads.
Use time-slicing for shareable workloads with clear expectations
NVIDIA GPU Operator can expose multiple time-sliced replicas of one physical GPU. The workloads interleave on the device, which can raise utilization for lightweight or bursty jobs. Unlike MIG, time-slicing does not provide the same memory or fault isolation between replicas, so it should be presented as shared capacity rather than as a smaller dedicated GPU.
Time-slicing is most useful when the workload can tolerate variable performance and shared memory pressure. Interactive notebooks, development, some inference workloads, and low-duty-cycle jobs may fit. Large training jobs that assume stable memory and compute access usually need a different allocation model.
Observability also changes. Some per-container metric associations are limited when time-slicing is used, so platform teams should test whether their chargeback and troubleshooting model still provides enough visibility before expanding shared pools.
Combine queues, priorities, and quotas with GPU requests
A scheduler can place a pod only after the cluster decides the pod is allowed to consume capacity. Namespace quotas, priority classes, admission controllers, and higher-level queueing systems help prevent one team or workload class from monopolizing accelerators. These policies should reflect business priority and service commitments rather than informal knowledge of which namespace belongs to whom.
Preemption can improve responsiveness for urgent work but can waste expensive computation if a long training job is killed without a recent checkpoint. Workload policy should therefore include checkpoint frequency and restart cost. Scheduling is not only about fitting resources; it is about controlling the cost of interrupting them.
Kubernetes security also shapes scheduling operations. Users who can alter quotas, priority, node labels, or device-plugin configuration can influence access to scarce infrastructure. Limit those permissions and audit changes so capacity policy cannot be bypassed through control-plane objects.
Make topology a scheduling input for distributed jobs
Multi-GPU jobs often need a particular topology more than they need any arbitrary set of devices. Placing ranks across slow paths can increase communication time and make scaling inefficient. Node affinity, topology managers, scheduler extensions, or workload-specific operators can help preserve locality when the application has strong communication requirements.
The platform should decide which topology properties are stable enough to expose. Overly detailed placement rules can make jobs impossible to schedule, while overly generic resources waste high-bandwidth groups on workloads that do not need them. Start with the distinctions that measurably affect performance and refine from telemetry.
Cluster topology should be part of capacity reviews. Queue time for eight tightly connected GPUs can grow even while many individual GPUs remain idle. The shape of free capacity matters as much as the total count.
Use utilization and queue time together
GPU monitoring shows whether placed workloads are using the devices they received, while scheduler metrics show whether unplaced workloads are waiting for capacity. Those two views answer different questions. Low utilization with a long queue suggests fragmentation or workload inefficiency; high utilization with a long queue suggests genuine capacity pressure.
Measure start latency by workload class, failed placements, preemptions, GPU-hours allocated, GPU-hours active, and the distribution of requested GPU counts. These metrics reveal whether a policy change improves the platform or simply shifts delay from one team to another.
The wider NVIDIA certification path provides useful context because operating accelerators at scale requires both infrastructure and workload awareness. Scheduling is the mechanism that converts hardware inventory into usable service capacity, so it deserves the same engineering rigor as network and storage design.
Separate interactive, batch, and serving policies
Different workload classes benefit from different scheduling behavior. Interactive development values short start time and can often tolerate shared accelerators or smaller partitions. Batch training may tolerate queueing but needs large, stable allocations for long periods. Online serving needs predictable latency, replica availability, and controlled rollout behavior. Treating these classes identically creates either poor utilization or poor service quality.
Create resource pools and queue policies that make the distinction explicit. Development pools can emphasize fairness and sharing, training pools can preserve topology for gang-scheduled jobs, and serving pools can reserve headroom for failover and scaling. The policy should be simple enough that users know where to submit work and what service characteristics to expect.
Measure each class with its own success metrics. Notebook users care about time to an available GPU, training teams care about throughput and completion time, and serving teams care about latency and availability. A single fleet-wide utilization target can push the platform toward decisions that improve the dashboard while making one of these workload classes worse.
The scheduling model should also define what happens when demand exceeds capacity. Waiting, preemption, bursting, quota borrowing, and priority overrides all have business consequences. Making those rules explicit prevents incident-time negotiations over whose job should be stopped when the cluster is full.
Admission policy can also prevent impossible requests from occupying the queue. Validate requested GPU counts, supported device classes, sharing modes, and topology requirements before a job waits for hours. Clear rejection with an actionable message is better than indefinite pending. Platform documentation should show valid examples so users can correct requests without relying on an operator to interpret every scheduling failure.
GPU scheduling in Kubernetes is a resource-modeling problem before it is a placement problem. The cluster needs accurate device descriptions, clear sharing semantics, controlled policy, topology awareness, and workload telemetry so the scheduler can make decisions that match real accelerator behavior.
When those pieces are aligned, utilization improves without turning the platform into a noisy neighbor experiment. Teams know what class of GPU service they are receiving, operators can explain queue behavior, and capacity planning reflects the shape of demand rather than only the number of devices installed.