Operator unavailable see history edit this page

Talks about: , , , , and

Linked from the jaas_operator_available == 0 signal and from GET /operator returning 503. The pod is up and the Jsonnet renderer answers; the controller manager is what cannot start.

Symptom

jaas_operator_available 0
kubectl --namespace <jaas-ns> exec deploy/<release> -- \
  wget --quiet --output-document=- http://127.0.0.1:8081/operator
# {"status":"unavailable","reason":"create manager: setup SnippetReconciler: register watch indexes: failed to get server groups: Get \"https://10.24.64.1:443/api\": dial tcp 10.24.64.1:443: i/o timeout","since":"2026-09-24T09:14:02Z","attempts":7}

Cause

Everything the manager needs from the apiserver happens while it is being built — discovery for the field indexes, the RESTMapper lookups behind the Flux source watches — so anything that blocks those calls keeps the operator down:

  1. Egress to the apiserver is denied. A cluster-wide default-deny NetworkPolicy with no matching allow rule is the common case. The reason string names the apiserver address and ends in i/o timeout or connection refused. See Network policy for the rule each policy engine needs.
  2. The operator’s own ClusterRole is incomplete. The reason carries forbidden, naming the verb and resource. The watch-layer variant of this is in operator-watch-silent .
  3. The JaaS CRDs are not installed. The reason names the missing kind. A chart install with crds.create=false and no out-of-band apply lands here.
  4. The apiserver is genuinely unreachable — control-plane outage, DNS failure inside the cluster, a mesh sidecar that has not started yet.
  5. The leader-election lease was lost after a period of healthy operation. The reason is leader election lost; the next manager blocks on acquiring the lease again, which is the behaviour the lease exists for. Repeated losses point at apiserver latency or a clock problem, not at JaaS.

Diagnosis

# The reason, verbatim, with how long it has been true and how many starts.
kubectl --namespace <jaas-ns> port-forward deploy/<release> 8081:8081 &
curl --silent http://127.0.0.1:8081/operator

# The same reason in the logs, once per transition plus one per retry.
kubectl --namespace <jaas-ns> logs deploy/<release> | grep -i "operator unavailable\|still unavailable"

# Is it reachability? Resolve and dial the apiserver from the pod's namespace.
kubectl --namespace <jaas-ns> get networkpolicy
kubectl --namespace default get endpoints kubernetes

An i/o timeout against the address that get endpoints kubernetes reports, with a default-deny policy in the namespace, is cause 1. A forbidden naming a resource is cause 2. A no matches for kind is cause 3.

Remediation

Fix the cause; no restart is needed. The manager is retried on a backoff that caps at five minutes, so a repaired cluster recovers within that at worst, and jaas_operator_available returns to 1 on its own.

For denied egress, admit the apiserver. On the Calico, Cilium, and ClusterNetworkPolicy engines the chart renders the rule for you once networkPolicy.egress.enabled is set; on the upstream kubernetes engine it needs the endpoint CIDRs:

networkPolicy:
  enabled: true
  egress:
    enabled: true
    kubernetesAPI:
      ipBlocks:
        - 10.24.64.1/32
      port: 6443

For an incomplete ClusterRole, reinstall or upgrade the chart so the generated roles match the running binary, then confirm:

kubectl auth can-i --as=system:serviceaccount:<jaas-ns>:<release> \
  list jsonnetsnippets.jaas.metio.wtf --all-namespaces

For missing CRDs, apply them and wait — the operator picks them up without a restart:

kubectl apply --server-side --filename config/crd/bases/