Hugging Face Spaces ZeroGPU
Hugging Face Spaces ZeroGPU is a hosted shared-GPU service for Gradio applications on the Hugging Face Hub. A developer marks GPU work with the Python `@spaces.GPU` decorator. When that function runs, the platform allocates GPU capacity for the task and releases it afterwards instead of reserving an accelerator continuously for one Space. `ZeroGPU` is a Hugging Face product name, not a generic synonym for serverless inference.
Origin and context
The public launch was announced in May 2024 as shared infrastructure for community AI demos. Contemporary coverage described A100 accelerators and mostly inference workloads. The implementation has since changed: a 2025 Hugging Face engineering article documented H200 slices, while the current reviewed documentation lists RTX Pro 6000 Blackwell sizes. Those changes are part of the product history, not interchangeable current specifications.
Why it matters
Interactive model demos often receive sparse, bursty traffic, so a permanently attached GPU can sit idle. ZeroGPU gives Spaces a provider-managed allocation path and lets one application request more than one GPU concurrently when capacity permits. It also makes every Gradio Space callable through generated API endpoints. A peer-reviewed NAACL demonstration used ZeroGPU to host a GPU-backed reviewing system, providing independent evidence of practical adoption while explicitly noting quota and responsiveness costs.
Example
A team publishes a Gradio image-generation demo and decorates its inference function with `@spaces.GPU(duration=...)`. Model setup remains at module level, while the decorated call receives the actual GPU. The team chooses a realistic maximum duration because shorter requests receive better queue treatment, tests the supported runtime versions and exposes the resulting Space through its generated Gradio API. For repeated short-lived processes, ahead-of-time compilation may avoid rebuilding an optimized graph on every task.
How it differs
Workload–Router–Pool Architecture (WRP)
Workload–Router–Pool is a broader architectural framing. ZeroGPU is one vendor service with a specific decorator, scheduler, supported runtime and quota model; the public documentation does not establish it as the implementation of that proposed taxonomy.
Router models / Cascade routing
Model or cascade routing selects which model should answer a request. ZeroGPU allocates compute to a function after the application has already selected its code and model.
GPU-rich and GPU-poor
GPU-poor/GPU-rich describes unequal access to compute. ZeroGPU can lower the entry barrier for demos, but quotas and queues mean it does not eliminate compute scarcity or prove equal access.
Maturity and evidence
Maturity is 3: ZeroGPU has operated since 2024, has a documented developer interface and version matrix, appears in current pricing, supports API-accessible Spaces and has independent deployment evidence. It is not rated higher because compatibility is narrower than ordinary GPU Spaces and core parameters—hardware, quotas, queue priority and supported versions—remain mutable service policy rather than a portable standard.
Limits and open questions
The service is currently Gradio-only, supports a bounded set of Python and PyTorch versions and may behave differently from a dedicated GPU Space. Users share quotas and queue capacity, so cold starts and waiting can reduce responsiveness; the independent OpenReviewer deployment reports this directly. `torch.compile` is not supported in the ordinary path, although Hugging Face documents ahead-of-time alternatives. Current hardware, included usage and prices must be checked at decision time. Free access is quota-limited and should not be described as unlimited or as an SLA-backed production endpoint.
Related terms
References
- Spaces ZeroGPU: Dynamic GPU Allocation for SpacesHugging Face · 2026 · class A
- Make your ZeroGPU Spaces go brrr with ahead-of-time compilationHugging Face · 2025-09-02 · class A
- Spaces as API endpointsHugging Face · 2026 · class A
- Hugging Face pricingHugging Face · 2026 · class A
- Hugging Face to make $10M worth of old Nvidia GPUs freely available to AI devsThe Register · 2024-05-17 · class B
- OpenReviewer: A Specialized Large Language Model for Generating Critical Scientific Paper ReviewsAssociation for Computational Linguistics · 2025-04 · class A
Last updated: 2026-09-07
This term is also covered in the Skills Atlas as hugging face skill.
This term is also covered in the Skills Atlas as gradio skill.
This term is also covered in the Skills Atlas as pytorch skill.
This term is also covered in the Skills Atlas as gpu acceleration skill.
This term is also covered in the Skills Atlas as model deployment skill.
This term is also covered in the Skills Atlas as llm inference serving skill.