Skip to content

Install the VictoriaMetrics HA monitoring stack

Use the VictoriaMetrics packages to install clustered metrics storage, an HA scrape agent, HA alerting, and Grafana.

This guide assumes Kubernetes 1.25 or later and a default StorageClass. The example requests three 20 GiB vmstorage volumes. Check that the target cluster has enough storage, CPU, and memory before deploying it.

Declare the five package instances in one namespace:

how-to/package-stacks/victoria-metrics-cluster.nix (L21–L46)
instances.monitoring-system = {
vm-operator = {
package = packages.victoria-metrics-operator;
};
vm-cluster = {
package = packages.victoria-metrics-cluster;
config = {
retentionPeriod = "30d";
vmstorage.storage.resources.requests.storage = "20Gi";
};
};
vm-agent = {
package = packages.victoria-metrics-agent;
config.externalLabels.cluster = "example";
};
vm-alert = {
package = packages.victoria-metrics-alert;
};
grafana = {
package = packages.grafana;
};
};

View source on GitHub ↗

The package dependencies connect the components:

  • vm-agent sends samples to vm-cluster and uses its scrape interval.
  • vm-alert reads from and writes recording rules to vm-cluster. It also creates a two-replica VMAlertmanager.
  • Grafana discovers vm-cluster through the prometheusServer alias and adds its Prometheus-compatible query endpoint as a datasource.
  • The operator supplies the custom resource definitions and reconciles the VictoriaMetrics workloads.

The catalogue defaults create three vmstorage replicas, two vmselect replicas, two vminsert replicas, two VMAgent replicas, and two replicas for each alerting component. Pod disruption budgets allow one unavailable replica per component, and preferred pod anti-affinity spreads replicas across nodes when possible.

Set retentionPeriod to the amount of metrics history you need. Each vmstorage replica gets the volume request under vmstorage.storage.resources.requests.storage, so account for all three volumes when estimating capacity.

The example leaves Alertmanager state on ephemeral storage. For a long-lived cluster, set vm-alert.config.alertmanager.storage to a persistent volume claim spec. This preserves silences and the notification log across a simultaneous restart of both replicas.

Before making Grafana reachable outside the cluster, configure its admin credentials with an existing Secret. See Use externally managed Secrets for the Kix pattern.

Install beside an existing prometheus-operator

Section titled “Install beside an existing prometheus-operator”

The operator package installs four CRDs from the monitoring.coreos.com group: ServiceMonitor, PodMonitor, PrometheusRule and Probe. Its converter watches them and turns each object into the VictoriaMetrics equivalent, which is what lets existing monitoring definitions keep working.

A cluster that already runs prometheus-operator already has those four CRDs, at the upstream schemas. Installing a second version of them fails: the apply conflicts on .spec.versions with the field manager that owns them, usually helm or kubectl, and the deploy stops.

Tell the operator instance that the CRDs are already there:

vm-operator = {
package = packages.victoria-metrics-operator;
config.prometheusOperatorCRDs = "reference";
};

Kix then waits for each of the four CRDs to be Established and moves on, rather than declaring them. The out.mkServiceMonitor, out.mkPodMonitor, out.mkPrometheusRule and out.mkProbe builders are unchanged, so packages that create those objects need no edit.

Leave the setting at its default, "declare", on a cluster where VictoriaMetrics is the only monitoring stack. Setting "reference" where the CRDs are absent makes the deploy wait for them until it times out.

The VictoriaMetrics CRDs are not affected by this setting. The operator ships its own complete bundle either way, because it crashes at startup if any of the kinds it indexes is missing.

The operator converts four prometheus-operator kinds into VictoriaMetrics equivalents, and all four are on by default. That is what makes the package a drop-in: existing ServiceMonitors keep working.

Each converter writes its output into the namespace the source object lives in. A converter whose output nothing reads therefore spreads objects across the cluster for no benefit. The clearest case is running VictoriaMetrics next to Prometheus without a vmalert: every PrometheusRule becomes a VMRule that nothing evaluates.

There are six converters. serviceMonitors, podMonitors, prometheusRules and probes are on by default. scrapeConfigs and alertmanagerConfigs are off by default, because turning one on also installs its CRD and a cluster with no ScrapeConfigs has no use for it.

vm-operator = {
package = packages.victoria-metrics-operator;
config.convert = {
scrapeConfigs = true; # this cluster has ScrapeConfig objects
prometheusRules = false; # no vmalert here to evaluate the result
probes = false;
};
};

serviceMonitors, podMonitors and scrapeConfigs are what feed a VMAgent. Leave those on wherever the agent is meant to scrape what Prometheus scrapes. A Prometheus that reads ScrapeConfig objects has targets that no ServiceMonitor describes, and with the converter off the agent misses them without saying so.

Turning a converter off does not remove what it already converted. The operator stops reconciling that kind and leaves the objects it made.

Evaluate the cluster before deploying it:

Run in kix-examples/
❱ kix check how-to-package-victoria-metrics
 TOOL         RESULT  DETAILS                                                       
 eval         pass    65 manifests evaluated                                        
 kubeconform  pass    skipped (this validation tool is not yet integrated with Kix) 
 pluto        pass    skipped (this validation tool is not yet integrated with Kix) 
 kyverno      pass    skipped (this validation tool is not yet integrated with Kix) 
 scorecard    pass    0 errors, 18 warnings, 3 info

List the resolved package instances:

Run in kix-examples/
❱ kix list packages --cluster how-to-package-victoria-metrics
 NAME              VERSION   OWNER     STATUS                        
 grafana           13.1.1    platform  installed [monitoring-system] 
 platform-dns      -         -         import [kube-system]          
 platform-storage  -         -         import [kube-system]          
 vm-agent          v1.150.0  platform  installed [monitoring-system] 
 vm-alert          v1.150.0  platform  installed [monitoring-system] 
 vm-cluster        v1.150.0  platform  installed [monitoring-system] 
 vm-operator       v0.74.1   platform  installed [monitoring-system]

Inspect the custom resources that configure storage, scraping, and alerting:

Run in kix-examples/ Output excerpt
❱ kix build how-to-package-victoria-metrics --output json Show output
[
  {
    "apiVersion": "operator.victoriametrics.com/v1beta1",
    "kind": "VMAgent",
    "metadata": {
      "name": "vm-agent",
      "namespace": "monitoring-system"
    },
    "spec": {
      "replicaCount": 1,
      "scrapeInterval": "30s",
      "remoteWrite": [
        {
          "url": "http://vminsert-vm-cluster.monitoring-system.svc.cluster.local:8480/insert/0/prometheus/api/v1/write"
        }
      ],
      "externalLabels": {
        "cluster": "example"
      }
    }
  },
  {
    "apiVersion": "operator.victoriametrics.com/v1beta1",
    "kind": "VMAlert",
    "metadata": {
      "name": "vm-alert",
      "namespace": "monitoring-system"
    },
    "spec": {
      "replicaCount": 1,
      "datasource": {
        "url": "http://vmselect-vm-cluster.monitoring-system.svc.cluster.local:8481/select/0/prometheus"
      },
      "remoteRead": {
        "url": "http://vmselect-vm-cluster.monitoring-system.svc.cluster.local:8481/select/0/prometheus"
      },
      "remoteWrite": {
        "url": "http://vminsert-vm-cluster.monitoring-system.svc.cluster.local:8480/insert/0/prometheus"
      }
    }
  },
  {
    "apiVersion": "operator.victoriametrics.com/v1beta1",
    "kind": "VMAlertmanager",
    "metadata": {
      "name": "vm-alert",
      "namespace": "monitoring-system"
    },
    "spec": {
      "replicaCount": 1,
      "retention": "120h"
    }
  },
  {
    "apiVersion": "operator.victoriametrics.com/v1beta1",
    "kind": "VMCluster",
    "metadata": {
      "name": "vm-cluster",
      "namespace": "monitoring-system"
    },
    "spec": {
      "retentionPeriod": "30d",
      "replicationFactor": 1,
      "vmstorage": {
        "replicaCount": 1,
        "storage": {
          "volumeClaimTemplate": {
            "spec": {
              "resources": {
                "requests": {
                  "storage": "20Gi"
                }
              },
              "storageClassName": "standard"
            }
          }
        },
        "podDisruptionBudget": {
          "maxUnavailable": 1
        }
      },
      "vmselect": {
        "replicaCount": 1,
        "podDisruptionBudget": {
          "maxUnavailable": 1
        }
      },
      "vminsert": {
        "replicaCount": 1,
        "podDisruptionBudget": {
          "maxUnavailable": 1
        }
      }
    }
  }
]

The output excerpt shows the HA replica counts, persistent vmstorage request, and URLs carried from the cluster package into the agent and alerting packages. The command itself emits the complete generated manifest set.

Deploy the stack, then check the package and custom resource status:

Run in kix-examples/
❱ kix deploy how-to-package-victoria-metrics
❱ kix status how-to-package-victoria-metrics
❱ kubectl get vmcluster,vmagent,vmalert,vmalertmanager -n monitoring-system

The operator must become ready before it can create the operand workloads. Image pulls and the initial volume provisioning can make the first deployment take several minutes.

To open Grafana locally, forward its Service port:

Run in kix-examples/
❱ kix pf how-to-package-victoria-metrics grafana 3000:3000

Open http://127.0.0.1:3000 while the forward is running. Stop it with Ctrl-C.

If an operand remains pending, inspect its custom resource and the operator logs, then check node capacity and the vmstorage PVCs:

❱ kubectl describe vmcluster -n monitoring-system vm-cluster
❱ kubectl logs -n monitoring-system deployment/vm-operator
❱ kubectl get pvc -n monitoring-system