LLMOps

AI Operations Intermediate

LLMOps is the set of practices for running LLM-based systems reliably in production, prompt versioning, evaluation pipelines, deployment, monitoring, and rollback, the AI-specific evolution of what DevOps and MLOps do for traditional software and machine learning systems.

In simple terms

Shipping an LLM feature isn't a one-time deploy; prompts change, models get upgraded, and outputs vary in ways a normal software rollback strategy doesn't fully account for. LLMOps is the discipline of managing that: testing changes before they ship, watching quality after they do, and being able to roll back fast when something regresses.

Why it matters

An LLM's behavior can shift from a prompt tweak, a model version upgrade, or even a change in the underlying provider's serving infrastructure, none of which look like a code change in a traditional sense, so traditional CI/CD alone doesn't catch these regressions.

How it works

Prompts and model configurations are version-controlled like code. Changes run through an evaluation pipeline (automated tests against a golden dataset, sometimes human review) before deployment. Production traffic is monitored for quality and cost drift, often with canary or shadow deployments that test a change against a small slice of real traffic before a full rollout.

Where it fits

Prompt / model change → Evaluation pipeline → Canary or shadow deployment → Production monitoring → Full rollout or rollback

Production impact

Teams without an evaluation gate in front of prompt changes routinely ship silent quality regressions, a prompt tweak that fixes one case while quietly breaking three others, that nobody notices until a user complains.

Learn this first

Related terms

Learn next

← All terms · Knowledge map →