How much memory is left on your nodes? Which pods keep restarting? And does an alert actually reach the team on call? With kube-prometheus-stack, you set up metrics, dashboards and alerting for your Kubernetes cluster in one go.
This guide takes you from the Helm installation and persistent storage to monitoring your own application. Finally, you verify alert delivery with a test alert. The result is a foundation you can adapt to the requirements of your production environment.
What is kube-prometheus-stack?
kube-prometheus-stack is a Helm chart maintained by the prometheus-community. By default, it installs Prometheus Operator, Prometheus, Alertmanager, Grafana, node-exporter and kube-state-metrics. It also ships preconfigured Kubernetes dashboards and alert rules.
| Component | Purpose |
|---|---|
| Prometheus Operator | Manages monitoring resources and generates the Prometheus configuration from them |
| Prometheus | Scrapes metrics, stores time series and evaluates rules |
| Alertmanager | Groups and deduplicates alerts and delivers notifications |
| Grafana | Visualizes the metrics in dashboards |
| node-exporter | Provides system metrics of the Linux nodes, such as CPU, RAM and file systems |
| kube-state-metrics | Provides metrics about the state of Kubernetes objects |
The chart also installs Custom Resource Definitions (CRDs). They extend Kubernetes with resource types such as ServiceMonitor, PodMonitor and PrometheusRule. You create resources of these types; the operator processes the resources selected by your Prometheus instance. Chart documentation.
graph LR SM["ServiceMonitor / PodMonitor"] --> OP["Prometheus Operator"] PR["PrometheusRule"] --> OP OP -->|"configures and manages"| P["Prometheus"] OP -->|"manages"| AM["Alertmanager"] P -->|"scrapes metrics"| NE["node-exporter"] P -->|"scrapes metrics"| KSM["kube-state-metrics"] P -->|"scrapes /metrics"| APP["Your application"] P -->|"Alerts"| AM AM --> N["Email, Slack, webhook"] G["Grafana"] -->|"PromQL queries"| P
How does kube-prometheus-stack differ from Prometheus Operator?
kube-prometheus-stack is a Helm chart that installs the core components of the kube-prometheus project, whereas the Prometheus Operator is just one component within it that manages Prometheus and Alertmanager based on Kubernetes resources.
The chart used to be called prometheus-operator and was renamed because it installs more than the operator: alongside it, the chart ships the other components from the table above. Even so, it does not cover the complete kube-prometheus stack. According to the chart documentation, Prometheus Adapter and Blackbox Exporter are not included.
Prerequisites
You need:
- A Kubernetes cluster with Linux worker nodes and a Kubernetes version supported by the chart you choose, for example centron Managed Kubernetes. If a Prometheus Operator is already running in the cluster, for instance as part of monitoring bundled by your provider, clarify before the installation how a second operator will be kept separate.
kubectl get crd | grep monitoring.coreos.comshows whether the CRDs already exist. kubectlwith permissions to create cluster-wide resources (CRDs, ClusterRoles, ClusterRoleBindings, webhook configurations) as well as resources in the monitoring namespace and inkube-system. An administrative account is often used for the initial installation.- A Helm version that matches the chart. The commands below use Helm 3.
- A StorageClass with dynamic provisioning for Prometheus, Alertmanager and Grafana.
- Sufficient CPU and RAM on suitable nodes. In the example, the Prometheus container alone reserves
2Giof RAM; sidecars and the other components need additional resources. - A Bash environment for the commands, for example on Linux or WSL, plus
curl,jqandbase64. - For email notifications, an SMTP server reachable from the cluster and matching credentials.
Check cluster access and the available StorageClasses:
$ kubectl cluster-info
$ kubectl get nodes
$ kubectl get storageclassWithout an explicit storageClassName, Kubernetes uses the default StorageClass for new PVCs. If there is none, add a suitable class to the configuration. For the local Prometheus database, a block volume with a supported file system is a good fit, for example; Prometheus does not support NFS. StorageClasses, Prometheus storage.
Add the Helm repository and pin the version
$ helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
$ helm repo update
$ helm search repo prometheus-community/kube-prometheus-stackThe chart documentation also lists the chart as the OCI artifact oci://ghcr.io/prometheus-community/charts/kube-prometheus-stack and uses this source in its installation examples. This guide sticks with the classic Helm repository; its index contains version 91.9.0.
This guide refers to the configuration of chart 91.9.0. Set the version explicitly so that later runs do not accidentally install a different release:
$ CHART_VERSION='91.9.0'
$ helm show chart prometheus-community/kube-prometheus-stack --version "$CHART_VERSION"In particular, check kubeVersion against your cluster. Before moving to a newer chart version, read the upgrade notes.
Next, create your working directory. All further file paths are relative to it:
$ mkdir -p "$HOME/monitoring"
$ cd "$HOME/monitoring"
$ umask 077Create values.yaml
The following example enables persistent storage for all three stateful components. On a fresh installation without a password setting, the Grafana chart generates a random administrator password. The well-known default password of earlier chart versions was removed in version 79. Change to the password default.
Save the configuration as values.yaml and replace the SMTP placeholders:
prometheus:
prometheusSpec:
retention: 15d
retentionSize: 40GB
# Only select resources that carry the label of this release.
serviceMonitorSelectorNilUsesHelmValues: true
podMonitorSelectorNilUsesHelmValues: true
ruleSelectorNilUsesHelmValues: true
serviceMonitorSelector: {}
podMonitorSelector: {}
ruleSelector: {}
# Find matching labeled resources in all namespaces.
serviceMonitorNamespaceSelector: {}
podMonitorNamespaceSelector: {}
ruleNamespaceSelector: {}
resources:
requests:
cpu: 500m
memory: 2Gi
storageSpec:
volumeClaimTemplate:
spec:
# storageClassName: your-block-storage-class
accessModes: ["ReadWriteOnce"]
resources:
requests:
storage: 50Gi
alertmanager:
alertmanagerSpec:
storage:
volumeClaimTemplate:
spec:
# storageClassName: your-block-storage-class
accessModes: ["ReadWriteOnce"]
resources:
requests:
storage: 2Gi
config:
global:
resolve_timeout: 5m
route:
receiver: team-mail
group_by: ["alertname", "namespace"]
group_wait: 30s
group_interval: 5m
repeat_interval: 4h
routes:
- receiver: "null"
matchers:
- 'alertname = "Watchdog"'
receivers:
- name: "null"
- name: team-mail
email_configs:
- to: "<your-team-address>"
from: "<sender-address>"
smarthost: "<smtp-server>:587"
auth_username: "<smtp-user>"
auth_password: "<smtp-password>"
require_tls: true
send_resolved: true
grafana:
persistence:
enabled: true
# storageClassName: your-block-storage-class
size: 10GiThe storage sizes and resources are starting values for the example. Size them according to cluster size, number of time series and retention requirements. A single suitable node must have enough capacity for the Prometheus pod.
Retention: retention: 15d limits the retention period, retentionSize: 40GB the size of the retained data. Whichever limit is reached first applies. Prometheus recommends allocating at most 80–85% of the assigned disk space for this. The headroom mainly covers the space that running compactions temporarily occupy on top. WAL and head chunks already count toward retentionSize. So monitor the free volume space as well. Storage planning.
Resource selection: Here, the *SelectorNilUsesHelmValues options make the chart restrict empty object selectors to the label release: kube-prometheus-stack. Our own resources will carry this label later as well. The namespace selectors allow them to be discovered in all namespaces; you can narrow the search scope if needed. Chart default values.
Watchdog and email: The permanently firing Watchdog alert is routed to the empty receiver here. To continuously verify the alerting pipeline, you can later send it to an external service that raises an alarm when the signal stops arriving. send_resolved: true additionally enables resolved notifications.
The SMTP details in this file are a simple starter configuration. Protect the file and do not store it in a repository with real credentials. For long-term operation, manage the Alertmanager configuration through a Kubernetes Secret, for example with External Secrets or Sealed Secrets. Helm also stores the values used in its release data; file permissions alone are therefore no substitute for secret management. You can reference an existing Grafana Secret via grafana.admin.existingSecret.
Adapt control plane monitoring to your cluster
With managed Kubernetes, the metrics endpoints of etcd, the scheduler and the controller manager may not be accessible, depending on the provider. Only if this applies to your cluster, add these values to the same file:
kubeEtcd:
enabled: false
kubeScheduler:
enabled: false
kubeControllerManager:
enabled: falseThis disables their monitoring, not the Kubernetes components themselves. Scraping the Kubernetes API server is not affected. If your cluster does not use kube-proxy, add kubeProxy.enabled: false accordingly.
Matching infrastructure at centron
From container to cluster: managed Kubernetes from centron with control plane and traffic included. Explore managed Kubernetes →
Install kube-prometheus-stack
$ helm upgrade --install kube-prometheus-stack prometheus-community/kube-prometheus-stack \
--namespace monitoring --create-namespace \
--version "$CHART_VERSION" \
--values values.yaml \
--wait --timeout 10mThe remaining examples use the release name kube-prometheus-stack. With a different name, Service names and release labels may change.
After the installation, check both the pods and the PersistentVolumeClaims:
$ kubectl get pods -n monitoring
$ kubectl get pvc -n monitoringThe started pods should be ready and the PVCs should reach the status Bound. How long this takes depends on image downloads and storage provisioning, among other things. Because the operator creates further resources asynchronously, this check remains useful even after a successful Helm command.
node-exporter runs as a DaemonSet on the suitable Linux nodes. The number of its pods therefore depends on node selection and scheduling constraints.
Open Grafana and Prometheus
Start a port forward for Grafana:
$ kubectl port-forward -n monitoring svc/kube-prometheus-stack-grafana 3000:80Keep this terminal open. In a second terminal, read the generated password:
$ kubectl get secret -n monitoring kube-prometheus-stack-grafana \
-o jsonpath='{.data.admin-password}' | base64 --decode
$ printf '\n'Open http://localhost:3000 and log in as admin. Under Dashboards, you will find the bundled views for nodes, namespaces and workloads. If you use grafana.admin.existingSecret instead, read the credentials from that Secret.
For Prometheus, start another port forward in a separate terminal:
$ kubectl port-forward -n monitoring svc/kube-prometheus-stack-prometheus 9090:9090The web interface is then available at http://localhost:9090. For permanent team access, you can use an Ingress with TLS and authentication or internal access via VPN, for example. Prometheus and Alertmanager also support TLS and basic auth themselves. Do not expose their web interfaces without access protection. Prometheus authentication, Alertmanager HTTPS.
Connect your own application with a ServiceMonitor
A ServiceMonitor describes which Services Prometheus should scrape. The example assumes an application that is already running and has these properties:
| Property | Value in the example |
|---|---|
| Namespace | shop |
| Service name | shop-api |
| Label on the Service | app: shop-api |
| Name of the Service port | metrics |
| Metrics path | /metrics |
The application must expose metrics at that endpoint in a format supported by Prometheus. A ServiceMonitor does not add instrumentation to the application.
Save this resource as shop-api-servicemonitor.yaml:
apiVersion: monitoring.coreos.com/v1
kind: ServiceMonitor
metadata:
name: shop-api
namespace: shop
labels:
release: kube-prometheus-stack
spec:
selector:
matchLabels:
app: shop-api
endpoints:
- port: metrics
path: /metrics
interval: 30s$ kubectl apply -f shop-api-servicemonitor.yamlTwo selection steps apply here: The release label makes the ServiceMonitor visible to our Prometheus instance. spec.selector.matchLabels then selects the Services. Because no namespace selector of its own is set in the ServiceMonitor, it searches the shop namespace here.
port: metrics refers to the name of the Service port, not a port number. Without a different jobLabel, the resulting Prometheus label job equals the Service name, here shop-api. ServiceMonitor reference.
Create an alert rule for the application
The following rule fires when Prometheus has not recorded a successful scrape of the shop-api metrics endpoint for five minutes. It also covers the case where all matching targets disappear from discovery.
Save it as shop-api-rules.yaml:
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: shop-api-alerts
namespace: shop
labels:
release: kube-prometheus-stack
spec:
groups:
- name: shop-api
rules:
- alert: ShopApiMetricsUnavailable
expr: (sum(up{namespace="shop", job="shop-api", endpoint="metrics"}) or vector(0)) == 0
for: 5m
labels:
severity: critical
namespace: shop
service: shop-api
annotations:
summary: "No reachable metrics targets for shop-api"
description: "Successful scrapes of the shop-api metrics endpoint in namespace shop have been missing for at least 5 minutes."$ kubectl apply -f shop-api-rules.yamlFor each scrape, up has the value 1 on success or 0 on failure. sum(...) therefore counts the successfully scraped targets. or vector(0) returns a zero value even when there are no matching time series at all. The namespace filter prevents a service with the same name in another namespace from masking the outage.
The rule assumes a service that is expected to run permanently: intentional scale-to-zero, a wrong selector or a missing ServiceMonitor will trigger it as well. Also, a successful scrape does not prove that the API works functionally. Add availability checks or alerts on error rate and response time as needed.
Verify the installation
While the Prometheus port forward is running, you can query the active targets in another terminal:
$ curl -fsS 'http://localhost:9090/api/v1/targets' \
| jq -r '.data.activeTargets[] | [.labels.namespace, .labels.job, .labels.endpoint, .health, .lastError] | @tsv'Check whether the expected cluster targets and the shop-api targets are present and report up. A list without down is not enough: a completely missing target does not appear in it at all. If there are errors, the lastError column helps.
The rules API shows whether the application rule has been loaded:
$ curl -fsS 'http://localhost:9090/api/v1/rules' \
| jq '.data.groups[] | select(.name == "shop-api") | .rules[] | {name, state, health, lastError}'No output means the rule has not been loaded yet or was not selected. In case of an error, check health and lastError. A loaded, healthy rule does not yet confirm working email delivery.
Verify alert delivery with a test alert
A temporary test alert verifies the path from Prometheus through Alertmanager to the mailbox without shutting down your application. Inform the recipients before the test.
Save as monitoring-testalert.yaml:
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: monitoring-notification-test
namespace: monitoring
labels:
release: kube-prometheus-stack
spec:
groups:
- name: monitoring-notification-test
rules:
- alert: MonitoringNotificationTest
expr: vector(1)
for: 1m
labels:
severity: warning
namespace: monitoring
annotations:
summary: "Scheduled test of monitoring notifications"
description: "This is a test alert. There is no reported application outage."$ kubectl apply -f monitoring-testalert.yamlAs soon as the rule is loaded and evaluated, the alert first becomes pending and after one minute firing. Add to that the evaluation intervals, group_wait: 30s and mail delivery. Check the status in Prometheus and then the arrival in the mailbox.
Remove the test rule when you are done:
$ kubectl delete -f monitoring-testalert.yamlWith send_resolved: true, a resolved notification is also expected; it may arrive late because of the grouping intervals. If the first message does not arrive, check the alert in Prometheus, its arrival and any silences in Alertmanager, and SMTP reachability and credentials, in that order.
Additionally, test the actual application rule in a test environment: once with metrics targets that exist but are unreachable, and once with completely missing targets.
Why does a ServiceMonitor not show up in Prometheus?
Typical causes are a non-matching release label, an excluded namespace, or a Service without matching labels and port names.
First, check the selection of the Prometheus resource:
$ kubectl get prometheus -n monitoring kube-prometheus-stack-prometheus -o json \
| jq '.spec | {serviceMonitorSelector, serviceMonitorNamespaceSelector}'With the example configuration, you expect release: kube-prometheus-stack in the object selector and {} in the namespace selector. An empty object selector {} would match all ServiceMonitors within the selected namespaces; it does not lift a namespace restriction.
Then check the ServiceMonitor and the selected Service:
$ kubectl get servicemonitor -n shop shop-api -o yaml
$ kubectl get svc -n shop -l app=shop-api -o yaml
$ kubectl get endpointslice -n shop -l kubernetes.io/service-name=shop-apiCheck the Service labels, the port name metrics and the existing endpoints. If the target appears but its scrape fails, check network access, NetworkPolicies, the path and, if applicable, TLS or authentication. You can also trace configuration problems in the operator logs:
$ kubectl logs -n monitoring deployment/kube-prometheus-stack-operator --tail=100Other common errors
PVCs stay Pending
Check the events of the affected claim and the StorageClass configuration:
$ kubectl describe pvc -n monitoring
$ kubectl get storageclassPossible causes are a missing default class, storage quotas or problems with the CSI provisioner. With WaitForFirstConsumer, a claim can wait until the scheduler has found a suitable node for the pod.
Outdated CRDs after a chart update
By default, Helm 3 does not update CRDs from the crds directory during an upgrade. The chart offers an optional CRD upgrade job (crds.upgradeJob.enabled), which it marks as a preview feature itself; alternatively, you update the CRDs manually according to the notes for the target version.
For chart 91.9.0, the files are located under charts/crds/crds/ after unpacking. In a new shell, set CHART_VERSION to the target version again, otherwise helm pull downloads the latest release. For other versions, check the path in the unpacked archive. In a new working directory, the manual step is:
$ helm pull prometheus-community/kube-prometheus-stack \
--version "$CHART_VERSION" --untar
$ kubectl apply --server-side \
-f kube-prometheus-stack/charts/crds/crds/Server-side apply avoids the size problems of the annotation that client-side apply would create for large CRDs. If field conflicts occur, investigate their cause. --force-conflicts takes over the affected fields and therefore does not belong in every update command without a second look. Helm and CRDs, Upgrade guide.
Persistent alerts for control plane components
Check whether the metrics endpoints in question are available in the cluster. Depending on discovery and the error, alerts such as TargetDown or component-specific alerts like KubeControllerManagerDown can occur. If your provider does not expose the endpoints, adjust the monitoring selection as described above.
Prometheus is terminated with OOMKilled
Distinguish between an exceeded container memory limit and memory pressure on the node. The pod configuration and events give first clues:
$ kubectl describe pod -n monitoring prometheus-kube-prometheus-stack-prometheus-0
$ kubectl get limitrange -n monitoringrequests.memory reserves resources for scheduling. limits.memory caps the container's consumption. Raising only the request does not fix a limit that is too low; the example does not set a limit of its own. Also check whether cluster defaults add one. Kubernetes resource management.
This PromQL query shows metric names with many active time series:
topk(10, count by (__name__)({__name__=~".+"}))With high cardinality, limit variable label values in the application if possible, or drop dispensable series selectively via metricRelabelings. Do not remove labels across the board: different series can become identical as a result, without their values being aggregated. Prometheus relabeling.
Remove the installation
If you no longer need the installation, first remove the example resources from this guide. You already deleted the test alert after the check. If it still exists, also remove it with kubectl delete -f monitoring-testalert.yaml.
$ kubectl delete -f shop-api-rules.yaml -f shop-api-servicemonitor.yamlThen remove the release:
$ helm uninstall kube-prometheus-stack --namespace monitoringhelm uninstall removes the resources of the release and its release history. Some parts nevertheless remain in the cluster.
Kubelet Service: The Service kube-prometheus-stack-kubelet in the kube-system namespace is created by the operator at runtime, not by the chart. helm uninstall therefore does not cover it. Check it with kubectl get service -n kube-system kube-prometheus-stack-kubelet and remove it with kubectl delete service -n kube-system kube-prometheus-stack-kubelet. Otherwise, a later installation under a different release name scrapes the kubelets twice.
Volume claims of Prometheus and Alertmanager: The operator runs both as a StatefulSet and creates their claims via volumeClaimTemplates. By default, such claims are retained when the StatefulSet is deleted (whenDeleted: Retain). Check which claims are left, and delete them only when you no longer need the stored data:
$ kubectl get pvc -n monitoring
$ kubectl delete pvc -n monitoring <pvc-name><pvc-name> stands for the name of a claim from the list returned by the first command.
Grafana volume claim: By default, the Grafana chart runs Grafana as a Deployment and creates the claim as a separate resource of the release. helm uninstall therefore removes it as well. Before that, back up any dashboards you created in Grafana yourself and still need.
Persistent volumes: For dynamically provisioned volumes, the reclaimPolicy of the StorageClass determines whether the volume is deleted along with the claim. If the StorageClass does not specify one, Delete applies. Reclaim policy.
CRDs: helm uninstall leaves the CRDs of the chart in place. The chart documentation provides for deleting them manually if needed. When a CRD is deleted, Kubernetes also removes all resources of that type across the entire cluster, including ServiceMonitors and rules that do not come from this guide. So delete the CRDs only if no other operator uses them. This applies in particular to the case mentioned in the prerequisites, where a Prometheus Operator is already present:
$ kubectl delete crd \
alertmanagerconfigs.monitoring.coreos.com \
alertmanagers.monitoring.coreos.com \
podmonitors.monitoring.coreos.com \
probes.monitoring.coreos.com \
prometheusagents.monitoring.coreos.com \
prometheuses.monitoring.coreos.com \
prometheusrules.monitoring.coreos.com \
scrapeconfigs.monitoring.coreos.com \
servicemonitors.monitoring.coreos.com \
thanosrulers.monitoring.coreos.comOnce you have checked the remaining claims and need nothing else from the namespace, remove it last:
$ kubectl delete namespace monitoringExternal systems and long-term operation
You can also monitor services outside the cluster with Prometheus. If a system already exposes a compatible metrics endpoint, you do not need an additional exporter. For other systems, suitable exporters provide the metrics. You can connect external targets via the ScrapeConfig resource, for example.
Whether you monitor surrounding VMs in the same Prometheus or separately from cluster monitoring depends on responsibilities, the checks you need and who responds to alerts.
For long-term operation, also define how access protection, secret management, backup and restore work. Persistent volumes retain data across pod restarts, but they replace neither backups nor an appropriate failure strategy. Whether you need multiple replicas or additional long-term storage such as Thanos or Grafana Mimir depends on your availability, scaling and retention requirements.
Conclusion
kube-prometheus-stack gives you a common foundation for Kubernetes metrics, dashboards and alerting. A usable installation requires suitable persistent storage, a pinned chart version and a traceable selection of your monitoring resources.
Then check three things separately: Are all expected targets present and reachable? Do your rules capture the failure cases you want? Do notifications reach the team? Only this complete path turns collected metrics into monitoring you can rely on in operation.
You can connect further workloads in the same cluster, such as a GitLab Runner on Kubernetes, using the same pattern, provided they expose a metrics endpoint: a Service with a named port, a ServiceMonitor with the release label and a matching rule.
Testen Sie Ihr Setup auf ccloud³
Registrieren Sie sich in der ccloud³ und erhalten Sie 200 € Startguthaben für Ihr Projekt – z. B. für eine PostgreSQL-VM mit automatischen Backups.