Blog

Terraform drift detection when your state is not all in one place

Drift checking usually assumes one workspace system holds everything. Real estates keep state in S3, a storage account, a GCS bucket, a workspace and someone's laptop. Here is what that does to detection, and how to get one answer.

Most Terraform drift tooling assumes a shape: your state lives in one system, that system runs your plans, and drift checking is a feature of it.

That shape is real, and plenty of teams have it. Plenty do not. And the ones who do not are usually not in the middle of a migration to it. They have arrived at something more ordinary.

The shape most estates actually have

None of that is a failure of discipline. It is what happens when infrastructure outlives the decision that shaped it, and when teams merge.

What that does to drift detection

Drift checking attached to a workspace system covers the workspaces it manages. That is the correct behaviour and it is also the limit. State in a bucket it does not run plans for is outside its world: not reported as unhealthy, simply not considered.

So a scattered estate ends up with:

The gap is not that any individual tool is bad. It is that the question spans systems and each tool sees one.

What “beyond” needs to mean

To get one answer over a scattered estate, a drift check has to:

Read state where it lives. Raw .tfstate from S3, from an Azure storage account, from GCS, and workspace state through a hosted API: as sources, not as a migration target. If the answer to “where is your state?” has to be “in one place” before the tool works, it is the same shape you already have.

Treat the union as the definition of managed. Whether a resource is undeclared is only answerable against every state source you have connected. Answer it against one and you will call a resource unmanaged because a different team’s state file, which you did not connect, declares it perfectly well.

Say what it did not look at. A tool that reports on the sources it has and stays silent about the rest is the one that produces false confidence. Coverage you did not connect is unexamined, not clean.

Compare against the live cloud, not just against itself. Two state files agreeing with each other tells you nothing about whether either matches reality.

The part that is genuinely hard

Reading several state sources is engineering. Making their answers comparable is the hard bit.

The same resource can appear in two state files during a migration. Two subscriptions can hold identically-named resources, so a finding that says nsg-prod without saying which scope is unactionable. And the resources that appear in no state file are the whole point of doing this, which means the tool has to be equally careful about the boundary of what it enumerated: one AWS account, one GCP project, or the subscriptions a credential can see.

Get those wrong and you produce a confident, wrong inventory, which is worse than no inventory because people act on it.

If you are already consolidated

Then most of this does not apply, and the drift checking built into your workspace system is the obvious answer. We have written separately about where that leaves teams who are not consolidated, including what those tools do better.

The case for something else is narrower than vendors like to admit: it is scattered state, or the unmanaged-resource question, or wanting the comparison to happen somewhere you control. If none of those is your situation, use what you have.


Cloudkeel-DD reads Terraform state from S3, Azure blob storage, GCS, local files and hosted workspaces, treats the union as the definition of managed, and cross-checks it against live AWS, Azure, GCP and Kubernetes.

It runs in your own cluster, read-only, with credentials that stay in your environment. Scans are point-in-time on an interval you set, and remediation arrives as a pull request you review.

Connect a state source and see: one Helm command, no form and no account.