Terraform drift has a clean definition: the state file says one thing, the cloud says another. Kubernetes does not give you that, and teams who carry the Terraform mental model across get confused quickly.
The difference is that in Kubernetes, things are supposed to change the live object. That is not a failure mode, it is the architecture.
Four writers, all legitimate
- Controllers. A Deployment’s controller writes status, and the HPA writes
spec.replicas. Your chart said 3; the cluster says 7; nothing is wrong. - Admission webhooks. A sidecar injector adds a container that appears in no chart you wrote. Service meshes and policy engines do this constantly.
- Defaulting. The API server fills in dozens of fields you never specified:
imagePullPolicy,terminationGracePeriodSeconds, a wholesecurityContext. A naive diff against your chart reports every one of them. - People.
kubectl edit,kubectl scale, or a patch from a pipeline that is not the pipeline that owns the object.
Only the last of those is what anyone means by drift. A detector that cannot separate the four produces a wall of findings on its first run, and the team turns it off in a week.
The two questions worth asking
Does this object still match what its release declared?
Not “does it match the chart on disk”, which is a different and less useful question: the rendered manifest at release time is the intention, and Helm records it. Comparing against that is what makes the answer meaningful, because it survives values files, overrides and templating.
Is this object declared by anything at all?
This is the Kubernetes version of the unmanaged-resource problem, and it is
usually where the interesting findings are. A Deployment with no Helm ownership
metadata, no owner reference to anything, and no annotation tying it to a
pipeline is a workload nobody’s process created. It might be an experiment
somebody forgot. It might be how production actually runs. Both are worth
knowing, and neither shows up in a helm diff.
Why ownership is the harder half
Finding a drifted object takes an afternoon. Working out who should care is the part that decides whether the finding gets fixed.
A namespace tells you less than you would hope: shared namespaces are the norm,
and default is where things land when nobody decided. Labels help when they
are applied consistently, which is to say they help on the clusters that least
need help. The Helm release name is often the best available signal, and it
still only tells you which chart, not which team.
There is no clean automatic answer here. Any tool that claims to route findings to the right team without you telling it your ownership model is guessing, and a confidently wrong owner is worse than an unrouted finding: the first gets ignored by someone who concluded it was not theirs, the second at least stays in a queue.
The workable pattern is to assign ownership deliberately, once, and let the tool carry it. Manual, and honest about being manual.
What point-in-time means here
Kubernetes changes faster than a scan interval. A detector that scans on a schedule is telling you what was true at that moment, not what is true now, and on a cluster with an active HPA those are genuinely different statements.
That is a real limitation and it is worth saying plainly rather than describing scheduled scans as continuous. What scheduled comparison is good at is the persistent difference: the replica count that has been wrong for a fortnight, the image tag that never matched, the Deployment nobody’s chart declares. Those do not disappear between scans, and they are the ones that matter.
Where this leaves you
Kubernetes drift detection is worth having when it answers the two questions above, separates the four writers, and gives a finding somewhere to end. It is not worth having as a diff of live objects against your charts, because the API server’s own defaulting will bury you.
Cloudkeel-DD reads Helm release metadata to decide what a workload was declared to be, reports objects that no release owns, and names which cluster a finding came from so two identically-named workloads are never confused. Ownership is assigned by you, not inferred.
It is self-hosted and read-only: it runs in your cluster with a read-only service account, and it never writes to your objects. Scans are point-in-time on an interval you choose.
Point it at a cluster: one Helm command, no form and no account.