Glossary · term

LLMOps

LLMOps is the engineering and governance discipline for developing, deploying, monitoring, and improving large-language-model systems in production. It adapts MLOps and DevOps practices to artifacts and failure modes such as prompts, model and provider versions, retrieval data, open-ended evaluations, safety controls, traces, token cost, latency, and human feedback. The operational unit may be an application assembled around an external model, not only a model trained in-house.

LLMOps2024-06-25Wave 2 · 2024Maturity: 3/5

Origin and context

By 2024, LLMOps had become common enough for researchers to synthesize practitioner definitions while noting that scientific literature had not converged on one boundary. The SpliTech paper described it as an MLOps adaptation for LLM-specific business, infrastructure, and lifecycle challenges. A 2025 review independently examined the transition from MLOps to LLMOps, including prompt work, generative evaluation, deployment, monitoring, security, and ethical auditing.

Sources: s1, s2

Why it matters

An LLM feature can change when a prompt, retrieval corpus, tool, policy, model snapshot, or provider behavior changes. Its outputs are probabilistic and often cannot be covered by exact-match tests. LLMOps makes those dependencies versioned and observable, links release decisions to evaluations, and gives teams a way to monitor quality, cost, latency, safety, and compliance throughout the lifecycle rather than only at model deployment.

Sources: s1, s2

Example

Before changing the model behind a support assistant, a team records the candidate model and prompt versions, runs a representative evaluation suite, checks retrieval and safety regressions, and compares cost and latency. It deploys to a small traffic segment, retains traces under an approved data policy, monitors failure indicators, and keeps a rollback path. The same release record links code, prompts, data snapshots, evaluations, and approval evidence.

Sources: s1, s2

How it differs

LLM evaluations (evals)

Evals are tests and measurement procedures. LLMOps is the broader lifecycle discipline that versions evals, decides when they gate a release, monitors production signals, and connects findings to rollback or improvement work. Running one benchmark is not a complete operations practice.

Agent observability

Agent observability focuses on traces, state, tool calls, and behavior of agentic workflows. It can be part of LLMOps, but LLMOps also covers development, evaluation, deployment, cost, governance, and non-agent LLM applications.

Compound AI Systems

Compound AI systems describe an architecture composed of interacting components. LLMOps describes how such a system is versioned, tested, released, observed, governed, and improved. One is system structure; the other is lifecycle practice.

Maturity and evidence

Maturity is rated 3. Independent peer-reviewed work agrees on a recognizable lifecycle discipline and its relationship to MLOps, but terminology, stages, metrics, and platform boundaries still vary. Evidence is stronger than a vendor buzzword and weaker than a settled standard.

Sources: s1, s2

Limits and open questions

LLMOps has no universal control framework, and vendor platforms often bundle different capabilities under the label. More telemetry does not guarantee useful diagnosis, while retaining prompts and outputs can create privacy and access risks. Automated judges can introduce their own bias, and a passing offline suite may not predict production behavior. Teams should define scoped service objectives, data-retention rules, ownership, escalation paths, and release gates instead of treating purchase of an LLMOps tool as operational maturity.

Sources: s1, s2

Related terms

References

Last updated: 2026-09-03

In the Skills Atlas

This term is also covered in the Skills Atlas as model deployment skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as prompt management skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as llm testing skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as experiment tracking skill.