Glossary · term

Hugging Face Spaces ZeroGPU

Hugging Face Spaces ZeroGPU is a hosted shared-GPU service for Gradio applications on the Hugging Face Hub. A developer marks GPU work with the Python `@spaces.GPU` decorator. When that function runs, the platform allocates GPU capacity for the task and releases it afterwards instead of reserving an accelerator continuously for one Space. `ZeroGPU` is a Hugging Face product name, not a generic synonym for serverless inference.

Products2024-05-16Wave 2 · 2024Maturity: 3/5

Origin and context

The public launch was announced in May 2024 as shared infrastructure for community AI demos. Contemporary coverage described A100 accelerators and mostly inference workloads. The implementation has since changed: a 2025 Hugging Face engineering article documented H200 slices, while the current reviewed documentation lists RTX Pro 6000 Blackwell sizes. Those changes are part of the product history, not interchangeable current specifications.

Sources: s1, s2, s5

Why it matters

Interactive model demos often receive sparse, bursty traffic, so a permanently attached GPU can sit idle. ZeroGPU gives Spaces a provider-managed allocation path and lets one application request more than one GPU concurrently when capacity permits. It also makes every Gradio Space callable through generated API endpoints. A peer-reviewed NAACL demonstration used ZeroGPU to host a GPU-backed reviewing system, providing independent evidence of practical adoption while explicitly noting quota and responsiveness costs.

Sources: s1, s3, s6

Example

A team publishes a Gradio image-generation demo and decorates its inference function with `@spaces.GPU(duration=...)`. Model setup remains at module level, while the decorated call receives the actual GPU. The team chooses a realistic maximum duration because shorter requests receive better queue treatment, tests the supported runtime versions and exposes the resulting Space through its generated Gradio API. For repeated short-lived processes, ahead-of-time compilation may avoid rebuilding an optimized graph on every task.

Sources: s1, s2, s3

How it differs

Workload–Router–Pool Architecture (WRP)

Workload–Router–Pool is a broader architectural framing. ZeroGPU is one vendor service with a specific decorator, scheduler, supported runtime and quota model; the public documentation does not establish it as the implementation of that proposed taxonomy.

Router models / Cascade routing

Model or cascade routing selects which model should answer a request. ZeroGPU allocates compute to a function after the application has already selected its code and model.

GPU-rich and GPU-poor

GPU-poor/GPU-rich describes unequal access to compute. ZeroGPU can lower the entry barrier for demos, but quotas and queues mean it does not eliminate compute scarcity or prove equal access.

Maturity and evidence

Maturity is 3: ZeroGPU has operated since 2024, has a documented developer interface and version matrix, appears in current pricing, supports API-accessible Spaces and has independent deployment evidence. It is not rated higher because compatibility is narrower than ordinary GPU Spaces and core parameters—hardware, quotas, queue priority and supported versions—remain mutable service policy rather than a portable standard.

Sources: s1, s3, s4, s5, s6

Limits and open questions

The service is currently Gradio-only, supports a bounded set of Python and PyTorch versions and may behave differently from a dedicated GPU Space. Users share quotas and queue capacity, so cold starts and waiting can reduce responsiveness; the independent OpenReviewer deployment reports this directly. `torch.compile` is not supported in the ordinary path, although Hugging Face documents ahead-of-time alternatives. Current hardware, included usage and prices must be checked at decision time. Free access is quota-limited and should not be described as unlimited or as an SLA-backed production endpoint.

Sources: s1, s2, s4, s6

Related terms

References

Last updated: 2026-09-07

In the Skills Atlas

This term is also covered in the Skills Atlas as hugging face skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as gradio skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as pytorch skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as gpu acceleration skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as model deployment skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as llm inference serving skill.