LLMOps
LLMOps is the engineering and governance discipline for developing, deploying, monitoring, and improving large-language-model systems in production. It adapts MLOps and DevOps practices to artifacts and failure modes such as prompts, model and provider versions, retrieval data, open-ended evaluations, safety controls, traces, token cost, latency, and human feedback. The operational unit may be an application assembled around an external model, not only a model trained in-house.
Origin and context
By 2024, LLMOps had become common enough for researchers to synthesize practitioner definitions while noting that scientific literature had not converged on one boundary. The SpliTech paper described it as an MLOps adaptation for LLM-specific business, infrastructure, and lifecycle challenges. A 2025 review independently examined the transition from MLOps to LLMOps, including prompt work, generative evaluation, deployment, monitoring, security, and ethical auditing.
Why it matters
An LLM feature can change when a prompt, retrieval corpus, tool, policy, model snapshot, or provider behavior changes. Its outputs are probabilistic and often cannot be covered by exact-match tests. LLMOps makes those dependencies versioned and observable, links release decisions to evaluations, and gives teams a way to monitor quality, cost, latency, safety, and compliance throughout the lifecycle rather than only at model deployment.
Example
Before changing the model behind a support assistant, a team records the candidate model and prompt versions, runs a representative evaluation suite, checks retrieval and safety regressions, and compares cost and latency. It deploys to a small traffic segment, retains traces under an approved data policy, monitors failure indicators, and keeps a rollback path. The same release record links code, prompts, data snapshots, evaluations, and approval evidence.
How it differs
LLM evaluations (evals)
Evals are tests and measurement procedures. LLMOps is the broader lifecycle discipline that versions evals, decides when they gate a release, monitors production signals, and connects findings to rollback or improvement work. Running one benchmark is not a complete operations practice.
Agent observability
Agent observability focuses on traces, state, tool calls, and behavior of agentic workflows. It can be part of LLMOps, but LLMOps also covers development, evaluation, deployment, cost, governance, and non-agent LLM applications.
Compound AI Systems
Compound AI systems describe an architecture composed of interacting components. LLMOps describes how such a system is versioned, tested, released, observed, governed, and improved. One is system structure; the other is lifecycle practice.
Maturity and evidence
Limits and open questions
LLMOps has no universal control framework, and vendor platforms often bundle different capabilities under the label. More telemetry does not guarantee useful diagnosis, while retaining prompts and outputs can create privacy and access risks. Automated judges can introduce their own bias, and a passing offline suite may not predict production behavior. Teams should define scoped service objectives, data-retention rules, ownership, escalation paths, and release gates instead of treating purchase of an LLMOps tool as operational maturity.
Related terms
References
- Large Language Model Operations (LLMOps): Definition, Challenges, and Lifecycle ManagementTECNALIA Publications / IEEE · 2024-06-25 · class A
- Transitioning from MLOps to LLMOps: Navigating the Unique Challenges of Large Language ModelsInformation (MDPI) · 2025-01 · class A
Last updated: 2026-09-03
This term is also covered in the Skills Atlas as model deployment skill.
This term is also covered in the Skills Atlas as prompt management skill.
This term is also covered in the Skills Atlas as llm testing skill.
This term is also covered in the Skills Atlas as experiment tracking skill.