Blog

What multi-cloud drift detection can honestly promise

Most tools that claim drift detection across AWS, Azure and GCP are describing three different depths at once. Here are the three bars worth separating, why coverage is path-dependent rather than cloud-dependent, and what to ask a vendor.

“Drift detection across AWS, Azure and GCP” is a sentence that can mean three very different things, and most of the time nobody says which one.

It matters, because the three differ by more than an order of magnitude in what they have actually been shown to catch. If you are evaluating tools, the useful skill is telling them apart, including in our own material.

The three bars

Bar one: inventory. The tool can list what exists. It enumerates resources in an account and tells you they are there. This is genuinely useful for finding things nothing declares, and it says nothing about whether a resource’s settings match what you declared.

Bar two: field-level comparison, fixture-verified. The tool has a specification for a resource type (which fields to compare, how to normalise them) and that specification passes tests against recorded payloads. This is the bar most coverage tables report, and it is a real engineering artifact. It is not the same as having watched it work on a live account.

Bar three: proven against a live cloud. Somebody changed a real resource in a real account and the tool detected that specific field change. This is the only bar that has survived contact with the messiness of a real provider API: the defaulting, the eventual consistency, the fields that come back in a different shape than the documentation implies.

A vendor quoting one number for all three is not necessarily lying. They are usually quoting bar two and letting you hear bar three.

Ours, stated separately

Field-level comparison specs for 200 resource types, each fixture-verified: 79 Azure, 62 AWS, 59 GCP. Drift has been injected into a live account and detected field-level for three: Azure NSG rules, AWS security-group rules, GCP firewall rules.

Unmanaged detection enumerates 79 Azure, 64 AWS and 59 GCP types.

Those are three different sentences on purpose, and the gap between the second number and the third is the honest state of this category, not a quirk of one product. The coverage page generates its tables from the code, so it cannot drift from what actually ships.

Coverage is path-dependent, not cloud-dependent

This is the part that surprises people, and it is more useful than any per-cloud total.

What a drift tool can compare depends on how it gets your declared state, not only which cloud you are on. Reading a Terraform plan gives you a different set of comparable types than reading raw state from a bucket, because a plan is a summary and state is the full record.

For us, concretely: from a Terraform plan the AWS cross-check reaches exactly one type. AWS field-level diffing effectively wants raw .tfstate in S3. So “62 AWS types” is a true number attached to a specific path, and quoting it without the path is the wrong shape even though the figure is right.

Ask a vendor which path a coverage number applies to. If the answer is a single number for every ingestion method, it is a marketing number.

What multi-cloud actually costs you

Three consoles is the visible cost. The invisible ones matter more:

Questions worth asking any vendor

  1. Which of the three bars is that coverage number? Ask for the live-cloud one specifically.
  2. Which ingestion path does it apply to: plan, raw state, or an API?
  3. Is severity comparable across clouds, or only detection?
  4. Where does the depth end, and is that written down anywhere a customer can read?

A vendor who answers the fourth without flinching is telling you something the first three cannot.


Cloudkeel-DD cross-checks your declared Terraform state against live AWS, Azure and GCP, and against Kubernetes. A cloud credential is a cross-check target, not a scan target: you connect the state, and the cloud is what it gets compared against.

It is self-hosted and read-only: it runs in your own cluster, your credentials stay in your environment, and scans are point-in-time on an interval you set.

Read the coverage table: it is generated from the code, including where the depth ends.