Skip to content

Add a post-deploy health check

This content is for the v0.1 version. Switch to the latest version for up-to-date documentation.

Use kix.healthCheck when Kubernetes readiness is necessary but not enough. For example, ready Pods do not prove that an application can reach its database or that a monitoring agent is successfully writing samples.

The helper adds a Job that runs after the package root and every endpoint used by the probe are ready. A non-zero exit fails the deployment, and the deploy output includes the last lines printed by the probe.

This guide assumes your package already has a workload, a Service exposed as self.service, and a list-valued build field.

Expose the standard options in the package:

how-to/application/web-package.nix (L47–L47)
healthCheck = kix.options.healthCheck;

View source on GitHub ↗

This gives cluster authors two settings:

OptionDefaultPurpose
healthCheck.enablefalseRender and run the probe
healthCheck.retryFor90Seconds the script may retry before failing

You can extend the option set with package-specific fields by merging another attribute set into kix.options.healthCheck.

Keep the script beside the package so it can be reviewed and tested on its own. This shell probe retries an HTTP endpoint for the period Kix supplies in RETRY_FOR:

how-to/application/health-check.sh (L3–L20)
set -u
deadline=$(( $(date +%s) + RETRY_FOR ))
while true; do
if curl --fail --silent --show-error --max-time 10 "$APP_URL" >/dev/null; then
echo "ok GET $APP_URL"
exit 0
fi
echo "FAIL GET $APP_URL"
if [ "$(date +%s)" -ge "$deadline" ]; then
echo "gave up after ${RETRY_FOR}s"
exit 1
fi
sleep 5
done

View source on GitHub ↗

Keep one pass through the checks below 90 seconds. Kix reserves that much time after the retry window so the Job can report its final failure before its deadline expires.

Append kix.healthCheck to the package’s build list:

how-to/application/web-package.nix (L130–L138)
(kix.healthCheck {
runtime = kix.healthCheck.runtimes.shell;
env =
{ self, ... }:
{
APP_URL = self.service.out.url { };
};
script = builtins.readFile ./health-check.sh;
})

View source on GitHub ↗

The callback value self.service.out.url { } matters. It gives the Job both the address to call and the dependency information attached to that address. Kix uses it to order the Job after the Service and to derive network-policy egress when network policy is enabled.

For a probe with several targets, add each target through env. The helper also provides deps, containing every dependency injected into any build entry in the package.

Use requires only for resources the probe must wait for but does not address. The package root is always included automatically.

Enable it on an instance:

how-to/application/cluster.nix (L43–L46)
healthCheck = {
enable = true;
retryFor = 60;
};

View source on GitHub ↗

Check the rendered cluster before deploying:

Run in kix-examples/
❱ kix check how-to-application
 TOOL         RESULT  DETAILS                                                       
 eval         pass    16 manifests evaluated                                        
 kubeconform  pass    skipped (this validation tool is not yet integrated with Kix) 
 pluto        pass    skipped (this validation tool is not yet integrated with Kix) 
 kyverno      pass    skipped (this validation tool is not yet integrated with Kix) 
 scorecard    pass    0 errors, 11 warnings, 2 info

The output should include no health-check validation errors. Render the Job to confirm the helper is enabled:

Run in kix-examples/ Output excerpt
❱ kix build how-to-application -o json Show output
{
  "apiVersion": "batch/v1",
  "kind": "Job",
  "metadata": {
    "annotations": {
      "kix.run/depends-on": "8767b7nzgc1x5bpfa9gv71cpk9z8h0iw,cd3l507fvaf0wg5azfza49x6wjsiiynz,dlzbv1nahmar8f7b9jm6rmyvmqczmzdb,requires:jh345cnnffy1bnj4cz5x0hwq6ibxvkkr",
      "kix.run/identity-hash": "67bgr7ggiw82bhlc2c3ddxma573vpibj",
      "kix.run/package": "production",
      "kix.run/package-namespace": "how-to-app",
      "kix.run/rerun": "on-change"
    },
    "labels": {
      "app.kubernetes.io/component": "health-check",
      "app.kubernetes.io/instance": "production",
      "app.kubernetes.io/managed-by": "kix",
      "app.kubernetes.io/name": "production-health"
    },
    "name": "production-health",
    "namespace": "how-to-app"
  },
  "spec": {
    "activeDeadlineSeconds": 150,
    "backoffLimit": 0,
    "template": {
      "metadata": {
        "labels": {
          "app.kubernetes.io/component": "health-check",
          "app.kubernetes.io/name": "production-health"
        }
      },
      "spec": {
        "automountServiceAccountToken": false,
        "containers": [
          {
            "command": [
              "/bin/sh",
              "-eu",
              "/probe/health-check.sh"
            ],
            "env": [
              {
                "name": "APP_URL",
                "value": "http://production.how-to-app.svc.cluster.local:80"
              },
              {
                "name": "RETRY_FOR",
                "value": "60"
              }
            ],
            "image": "docker.io/curlimages/curl:8.14.1@sha256:9a1ed35addb45476afa911696297f8e115993df459278ed036182dd2cd22b67b",
            "name": "probe",
            "resources": {
              "limits": {
                "memory": "128Mi"
              },
              "requests": {
                "cpu": "50m",
                "memory": "64Mi"
              }
            },
            "securityContext": {
              "allowPrivilegeEscalation": false,
              "capabilities": {
                "drop": [
                  "ALL"
                ]
              },
              "readOnlyRootFilesystem": true,
              "runAsNonRoot": true
            },
            "terminationMessagePolicy": "FallbackToLogsOnError",
            "volumeMounts": [
              {
                "mountPath": "/probe",
                "name": "production-health-script",
                "readOnly": true
              }
            ]
          }
        ],
        "restartPolicy": "Never",
        "securityContext": {
          "runAsGroup": 65534,
          "runAsNonRoot": true,
          "runAsUser": 65534,
          "seccompProfile": {
            "type": "RuntimeDefault"
          }
        },
        "volumes": [
          {
            "configMap": {
              "name": "production-health-script"
            },
            "name": "production-health-script"
          }
        ]
      }
    },
    "ttlSecondsAfterFinished": 3600
  }
}

Deploy the cluster:

Run in kix-examples/
❱ kix deploy how-to-application -y Show output
Building cluster 'how-to-application'...
Cluster how-to-application: 16 manifests
Connecting to cluster...
No previous activation on cluster — first deploy.

  _cluster
    ~ cluster-level resources (4 added)
  how-to-app
    + preview 1.0.0 (3 resources)
    + production 1.0.0 (5 resources)
  kube-system
    + platform-dns (0 resources)

  Plan: cluster-level changes, 3 added
  Resources: 4 real content, 0 dep-affected
  ↻ 1 under the rerun rule (deleted first when live): Job/production-health@how-to-app
plan: 16 nodes
  ~ Namespace/kube-system configured
  + Namespace/how-to-app created
  ✔ Namespace/kube-system ready
  ✔ Namespace/how-to-app ready
  + CustomResourceDefinition/packageinstances.kix.run created
  + ConfigMap/production-health-script@how-to-app created
  ✔ ConfigMap/production-health-script@how-to-app ready
  + CustomResourceDefinition/activations.kix.run created
  + ConfigMap/preview@how-to-app created
  ✔ ConfigMap/preview@how-to-app ready
  + ConfigMap/production@how-to-app created
  ✔ ConfigMap/production@how-to-app ready
  ✔ CustomResourceDefinition/activations.kix.run ready
  + Deployment/preview@how-to-app created
  + Deployment/production@how-to-app created
   0.759471667s  WARN kix::cluster::client: apiserver request failed, retrying method=GET path="/apis/kix.run/v1alpha1/activations" attempt=1 max_attempts=7 delay_ms=1000 reason="429 Too Many Requests"
  ✔ CustomResourceDefinition/packageinstances.kix.run ready
  + PackageInstance/platform-dns@kube-system created
  ✔ PackageInstance/platform-dns@kube-system ready
  ✔ Deployment/preview@how-to-app ready
  + Service/preview@how-to-app created
  ✔ Service/preview@how-to-app ready
  + PackageInstance/preview@how-to-app created
  ✔ PackageInstance/preview@how-to-app ready
  ✔ Deployment/production@how-to-app ready
  + Service/production@how-to-app created
  ✔ Service/production@how-to-app ready
  + recreate Job/production-health@how-to-app created
  ✔ Job/production-health@how-to-app ready
  + PackageInstance/production@how-to-app created
  ✔ PackageInstance/production@how-to-app ready
  ~ Activation/how-to-application-aq1zdgf9gsnb configured
  ✔ Activation/how-to-application-aq1zdgf9gsnb ready
  • activation 'how-to-application-aq1zdgf9gsnb' → Active

Deploy complete: 14 created, 2 configured, 0 unchanged, 0 failed

The deployment does not become active until the Job succeeds. If the script exits non-zero, Kix reports the Job as failed and prints its final output below the failure.

For example, a probe that reaches the application but rejects its response is reported with the final lines from the script:

Run in kix-examples/
❱ kix deploy how-to-application -y
Building cluster 'how-to-application'...
Cluster how-to-application: 16 manifests
Connecting to cluster...
Active activation: how-to-application-aq1zdgf9gsnb (aq1zdgf9...)

  how-to-app
    ~ production 1.0.0 (1 changed, 2 dep-affected)

  Plan: 1 updated, 2 unchanged
  Resources: 1 real content, 2 dep-affected
  ↻ 1 under the rerun rule (deleted first when live): Job/production-health@how-to-app
plan: 16 nodes
  ~ ConfigMap/production-health-script@how-to-app configured
  ✔ ConfigMap/production-health-script@how-to-app ready
  + recreate Job/production-health@how-to-app created
  ⚠ Job/production-health@how-to-app JobBackoffExceeded: Job failed 1 times (backoffLimit=0)
      ok    GET http://production.how-to-app.svc.cluster.local:80
      the application returned an unexpected response

✗ halt-on-first-failure: JobBackoffExceeded: Job failed 1 times (backoffLimit=0)
Deploy complete: 0 created, 1 configured, 12 unchanged, 1 failed, 2 cancelled
(exit code: 1)

The Job carries kix.run/rerun: on-change. A later deploy recreates it when its identity changes or its previous run failed. A deploy with no relevant change skips a successful probe.