§ 01Cloud and Infrastructure Engineering

Admission control policies: the ones we ship and the ones we withdrew

Admission control is the cheapest control in a Kubernetes estate and the easiest to configure into uselessness. This is bench work on our own estate, four clusters we run for ourselves, together with what I learned running policy at a logistics operator earlier in my career. Nothing below is drawn from a client Kubernetes estate.

· 4 min read · Diego Sotomayor, Principal Engineer, Cloud and Infrastructure Engineering

What this ran on

The estate is ours. Four clusters: two on Amazon EKS carrying our build and artefact pipeline, one on Azure Kubernetes Service running internal identity and directory tooling, and one kubeadm cluster on a pair of machines we use as a detection laboratory. Between them they run 96 workloads. That is the whole population, and every figure below can be checked by hand.

In four years as a platform engineer at a logistics operator, I inherited the same failure in every cluster I was handed. The policy lived in the cluster rather than in the repository that built the cluster. A rebuild lost it. Nobody noticed until someone asked when a given rule had last changed and who approved it, and the answer had to be reconstructed from memory and a Slack thread. We built this library so that question has an answer in Git.

It is one set of Gatekeeper constraint templates, versioned and reconciled by Argo CD, with exceptions expressed as data rather than as forked policy. We wrote 12 constraints. Nine are enforced today.

  • CPU and memory requests and limits set on every container
  • Images pulled only from our own registry, and only with a valid cosign signature
  • Privileged containers, host networking and host path mounts refused outside a named exception
  • An owner label on every namespace, matching an entry in the service catalogue
  • A named TLS secret on every Ingress object

Audit first, and 3 weeks is not long enough

Gatekeeper will run a constraint in dryrun, which records violations without refusing anything. We left all 12 in dryrun for 3 weeks and read the audit output twice a week. The first pass returned just over 1,850 violations, and about 90 per cent came from four namespaces: a monitoring stack, a service mesh, a backup agent and an ingress controller, all vendor Helm charts shipping defaults that no policy set anywhere would pass. Most were closed by chart values rather than by exceptions. Three weeks does not cover a monthly batch job or a quarterly certificate renewal, so there are workloads in this estate the audit has never seen, and we know which ones they are.

Dryrun is not a stage of the rollout. It is most of the rollout, and enforcement is the last five minutes of it.

The three we withdrew

The resource limits constraint refused batch executor pods created by a scheduler that sets its own limits through its own mutating webhook, which runs after ours. Webhook ordering would fix it, and we have not finished testing the change, so that namespace carries an exemption with a review date on it. The namespace owner label constraint refused namespaces created by a provisioning tool that creates first and labels second, which is the tool's design and not a defect we can argue with. The seccomp profile constraint refused a vendor agent that requires an unconfined profile and offers no supported alternative. In each case the constraint was correct and the environment was not going to change, and a constraint you cannot enforce is documentation pretending to be a control.

What we got wrong

We enforced image provenance on the build cluster before the mirroring pipeline was signing third-party images. Nightly builds failed for 4 days before anyone connected the two events, because the failure surfaced as a stuck Argo CD sync rather than as an admission error. We had read the registry inventory, which lists the images, and not the pull logs, which would have shown which of them were signed. The inventory was ours and it was three weeks old.

The other limit is the size of the estate. Four clusters and 96 workloads test whether a constraint is correct; they do not test whether it survives in a place with hundreds of namespaces owned by teams who did not agree to it, where the argument is political before it is technical. We will publish what happens the first time we run this at that size. Until then the set we ship is 9 constraints, smaller than most published libraries, and it is the set we can defend line by line. The next change is ValidatingAdmissionPolicy, which moves the same checks into the API server using CEL and takes the external webhook out of the failure path. It is in dryrun on one cluster. The constraints will not change; the way they fail will.

Scoping is done by the director who will sign the report

The first scoping call is free. Where an assessment follows, it is a fixed fee agreed in writing before it starts.