Problem
The high-availability and scaling pages say extra admission-controller replicas are for availability and scale. That is true for AdmissionReview QPS. It is easy to read as if the policy set is partitioned across replicas.
It is not. Each admission replica watches every Policy and ClusterPolicy and keeps a full in-process cache (pkg/policycache in kyverno/kyverno). Horizontal scale does not reduce per-pod memory. Per-pod RSS is O(policy count).
That matters when a cluster stores thousands of namespaced Policy resources (for example one Policy per tenant or workload). Adding replicas then multiplies the same cache instead of splitting it. If the webhook uses failurePolicy: Fail and every replica OOMs under a burst of admission traffic, matching admission is blocked cluster-wide.
Related scale report (3.5k Policy CRs): kyverno/kyverno#10458. Maintainers there pointed at collapsing many namespaced Policies into a ClusterPolicy plus Global Context, not at sharding the cache.
What already exists
- Policy count is already a gauge:
kyverno_policy_rule_info_total (metrics).
- There is no cache-size-in-bytes metric. Counting policies plus pod RSS is enough to size vertically.
- Sharding the cache would be a controller redesign (any replica must still evaluate any matching policy). This issue is documentation only.
Ask
State on the scaling and high-availability pages that:
- Extra admission replicas distribute AdmissionReview load, not the policy cache.
- Size admission memory for policy count; add replicas for webhook throughput and availability.
- Prefer fewer, broader policies over thousands of near-identical namespaced
Policy resources.
I can send a docs PR.
Problem
The high-availability and scaling pages say extra admission-controller replicas are for availability and scale. That is true for AdmissionReview QPS. It is easy to read as if the policy set is partitioned across replicas.
It is not. Each admission replica watches every
PolicyandClusterPolicyand keeps a full in-process cache (pkg/policycachein kyverno/kyverno). Horizontal scale does not reduce per-pod memory. Per-pod RSS is O(policy count).That matters when a cluster stores thousands of namespaced
Policyresources (for example one Policy per tenant or workload). Adding replicas then multiplies the same cache instead of splitting it. If the webhook usesfailurePolicy: Failand every replica OOMs under a burst of admission traffic, matching admission is blocked cluster-wide.Related scale report (3.5k Policy CRs): kyverno/kyverno#10458. Maintainers there pointed at collapsing many namespaced Policies into a
ClusterPolicyplus Global Context, not at sharding the cache.What already exists
kyverno_policy_rule_info_total(metrics).Ask
State on the scaling and high-availability pages that:
Policyresources.I can send a docs PR.