Blog

Kubernetes drift detection is a different problem to Terraform drift

A Deployment can drift from its Helm release without anyone running kubectl, because controllers, admission webhooks and autoscalers all write to live objects by design. Here is what that means for detection, and why ownership is the harder half.

Terraform drift has a clean definition: the state file says one thing, the cloud says another. Kubernetes does not give you that, and teams who carry the Terraform mental model across get confused quickly.

The difference is that in Kubernetes, things are supposed to change the live object. That is not a failure mode, it is the architecture.

Four writers, all legitimate

Only the last of those is what anyone means by drift. A detector that cannot separate the four produces a wall of findings on its first run, and the team turns it off in a week.

The two questions worth asking

Does this object still match what its release declared?

Not “does it match the chart on disk”, which is a different and less useful question: the rendered manifest at release time is the intention, and Helm records it. Comparing against that is what makes the answer meaningful, because it survives values files, overrides and templating.

Is this object declared by anything at all?

This is the Kubernetes version of the unmanaged-resource problem, and it is usually where the interesting findings are. A Deployment with no Helm ownership metadata, no owner reference to anything, and no annotation tying it to a pipeline is a workload nobody’s process created. It might be an experiment somebody forgot. It might be how production actually runs. Both are worth knowing, and neither shows up in a helm diff.

Why ownership is the harder half

Finding a drifted object takes an afternoon. Working out who should care is the part that decides whether the finding gets fixed.

A namespace tells you less than you would hope: shared namespaces are the norm, and default is where things land when nobody decided. Labels help when they are applied consistently, which is to say they help on the clusters that least need help. The Helm release name is often the best available signal, and it still only tells you which chart, not which team.

There is no clean automatic answer here. Any tool that claims to route findings to the right team without you telling it your ownership model is guessing, and a confidently wrong owner is worse than an unrouted finding: the first gets ignored by someone who concluded it was not theirs, the second at least stays in a queue.

The workable pattern is to assign ownership deliberately, once, and let the tool carry it. Manual, and honest about being manual.

What point-in-time means here

Kubernetes changes faster than a scan interval. A detector that scans on a schedule is telling you what was true at that moment, not what is true now, and on a cluster with an active HPA those are genuinely different statements.

That is a real limitation and it is worth saying plainly rather than describing scheduled scans as continuous. What scheduled comparison is good at is the persistent difference: the replica count that has been wrong for a fortnight, the image tag that never matched, the Deployment nobody’s chart declares. Those do not disappear between scans, and they are the ones that matter.

Where this leaves you

Kubernetes drift detection is worth having when it answers the two questions above, separates the four writers, and gives a finding somewhere to end. It is not worth having as a diff of live objects against your charts, because the API server’s own defaulting will bury you.


Cloudkeel-DD reads Helm release metadata to decide what a workload was declared to be, reports objects that no release owns, and names which cluster a finding came from so two identically-named workloads are never confused. Ownership is assigned by you, not inferred.

It is self-hosted and read-only: it runs in your cluster with a read-only service account, and it never writes to your objects. Scans are point-in-time on an interval you choose.

Point it at a cluster: one Helm command, no form and no account.