# Install the VictoriaMetrics HA monitoring stack

Use the VictoriaMetrics packages to install clustered metrics storage, an HA
scrape agent, HA alerting, and Grafana.

This guide assumes Kubernetes 1.25 or later and a default StorageClass. The
example requests three 20 GiB vmstorage volumes. Check that the target cluster
has enough storage, CPU, and memory before deploying it.

## Add the stack

Declare the five package instances in one namespace:

<Snippet {...victoriaMetricsHaStack} />

The package dependencies connect the components:

* `vm-agent` sends samples to `vm-cluster` and uses its scrape interval.
* `vm-alert` reads from and writes recording rules to `vm-cluster`. It also
  creates a two-replica VMAlertmanager.
* Grafana discovers `vm-cluster` through the `prometheusServer` alias and adds
  its Prometheus-compatible query endpoint as a datasource.
* The operator supplies the custom resource definitions and reconciles the
  VictoriaMetrics workloads.

The catalogue defaults create three vmstorage replicas, two vmselect replicas,
two vminsert replicas, two VMAgent replicas, and two replicas for each alerting
component. Pod disruption budgets allow one unavailable replica per component,
and preferred pod anti-affinity spreads replicas across nodes when possible.

## Set retention and storage

Set `retentionPeriod` to the amount of metrics history you need. Each vmstorage
replica gets the volume request under
`vmstorage.storage.resources.requests.storage`, so account for all three
volumes when estimating capacity.

The example leaves Alertmanager state on ephemeral storage. For a long-lived
cluster, set `vm-alert.config.alertmanager.storage` to a persistent volume claim
spec. This preserves silences and the notification log across a simultaneous
restart of both replicas.

Before making Grafana reachable outside the cluster, configure its admin
credentials with an existing Secret. See [Use externally managed
Secrets](/docs/v0.1/how-to/platform-capabilities/use-externally-managed-secrets/)
for the Kix pattern.

## Install beside an existing prometheus-operator

The operator package installs four CRDs from the `monitoring.coreos.com`
group: ServiceMonitor, PodMonitor, PrometheusRule and Probe. Its converter
watches them and turns each object into the VictoriaMetrics equivalent, which
is what lets existing monitoring definitions keep working.

A cluster that already runs prometheus-operator already has those four CRDs,
at the upstream schemas. Installing a second version of them fails: the apply
conflicts on `.spec.versions` with the field manager that owns them, usually
`helm` or `kubectl`, and the deploy stops.

Tell the operator instance that the CRDs are already there:

```nix
vm-operator = {
  package = packages.victoria-metrics-operator;
  config.prometheusOperatorCRDs = "reference";
};
```

Kix then waits for each of the four CRDs to be Established and moves on,
rather than declaring them. The `out.mkServiceMonitor`, `out.mkPodMonitor`,
`out.mkPrometheusRule` and `out.mkProbe` builders are unchanged, so packages
that create those objects need no edit.

Leave the setting at its default, `"declare"`, on a cluster where
VictoriaMetrics is the only monitoring stack. Setting `"reference"` where the
CRDs are absent makes the deploy wait for them until it times out.

The VictoriaMetrics CRDs are not affected by this setting. The operator ships
its own complete bundle either way, because it crashes at startup if any of
the kinds it indexes is missing.

## Choose what the converter converts

The operator converts four prometheus-operator kinds into VictoriaMetrics
equivalents, and all four are on by default. That is what makes the package a
drop-in: existing ServiceMonitors keep working.

Each converter writes its output into the namespace the source object lives in.
A converter whose output nothing reads therefore spreads objects across the
cluster for no benefit. The clearest case is running VictoriaMetrics next to
Prometheus without a vmalert: every PrometheusRule becomes a VMRule that
nothing evaluates.

There are six converters. `serviceMonitors`, `podMonitors`, `prometheusRules`
and `probes` are on by default. `scrapeConfigs` and `alertmanagerConfigs` are
off by default, because turning one on also installs its CRD and a cluster with
no ScrapeConfigs has no use for it.

```nix
vm-operator = {
  package = packages.victoria-metrics-operator;
  config.convert = {
    scrapeConfigs = true;      # this cluster has ScrapeConfig objects
    prometheusRules = false;   # no vmalert here to evaluate the result
    probes = false;
  };
};
```

`serviceMonitors`, `podMonitors` and `scrapeConfigs` are what feed a VMAgent.
Leave those on wherever the agent is meant to scrape what Prometheus scrapes.
A Prometheus that reads ScrapeConfig objects has targets that no ServiceMonitor
describes, and with the converter off the agent misses them without saying so.

Turning a converter off does not remove what it already converted. The operator
stops reconciling that kind and leaves the objects it made.

## Check the generated stack

Evaluate the cluster before deploying it:

<Command {...check} />

List the resolved package instances:

<Command {...packages} />

Inspect the custom resources that configure storage, scraping, and alerting:

<Command expandable {...resources} />

The output excerpt shows the HA replica counts, persistent vmstorage request,
and URLs carried from the cluster package into the agent and alerting packages.
The command itself emits the complete generated manifest set.

## Deploy and inspect

Deploy the stack, then check the package and custom resource status:

<Command
  commands={[
  "kix deploy how-to-package-victoria-metrics",
  "kix status how-to-package-victoria-metrics",
  "kubectl get vmcluster,vmagent,vmalert,vmalertmanager -n monitoring-system",
]}
  cwd="kix-examples/"
/>

The operator must become ready before it can create the operand workloads.
Image pulls and the initial volume provisioning can make the first deployment
take several minutes.

To open Grafana locally, forward its Service port:

<Command commands={["kix pf how-to-package-victoria-metrics grafana 3000:3000"]} cwd="kix-examples/" />

Open `http://127.0.0.1:3000` while the forward is running. Stop it with
`Ctrl-C`.

If an operand remains pending, inspect its custom resource and the operator
logs, then check node capacity and the vmstorage PVCs:

<Command
  commands={[
  "kubectl describe vmcluster -n monitoring-system vm-cluster",
  "kubectl logs -n monitoring-system deployment/vm-operator",
  "kubectl get pvc -n monitoring-system",
]}
/>