说实话,刚开始接触 Kubernetes 监控的时候,我也懵过。那么多 Pod、Service、Deployment,节点还时不时抽风,日志满天飞,光靠 kubectl logs 根本抓瞎。后来我折腾了一套 Prometheus + Grafana 的方案,真·香。今天这篇,我就像给朋友讲代码一样,把这事儿掰开揉碎了说清楚,保证你看完就能上手,哪怕你是第一次碰监控,也能搞定 CPU 飙高、内存爆仓这种经典“坑”。
别急,先搞清楚我们在监控啥
K8s 里,一个 Pod 跑着容器,容器吃 CPU 和内存,节点上有多个 Pod。我们要盯的,主要是这几类数据:
- 节点级别:CPU 使用率、内存可用/已用、磁盘 I/O
- 容器级别:每个容器的 CPU 核心占用、内存 RSS(物理内存)、CPU 限额/请求比例
- Pod 级别:重启次数、运行状态、网络流量
- 应用层面:QPS、延迟、错误率(这个得看你业务埋点,先不说)
新手最容易踩的坑是:只看节点 CPU,不看容器 CPU。比如节点整体 CPU 50%,但某个 Pod 可能已经 95% 了,因为它独占了 3 个核。所以得深入容器层看。
第一步:部署 Prometheus 和 Kube-Prometheus-Stack
别自己瞎装 Prometheus,太麻烦。推荐用 kube-prometheus-stack,这是 Prometheus 官方社区维护的 Helm Chart,把 Prometheus、Grafana、Alertmanager 全打包了,一键搞定。
先装 Helm
# Ubuntu/Debian
apt-get update && apt-get install -y helm
# 或者 macOS
brew install helm
添加 Helm Repo
helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
helm repo update
安装 Stack
# 创建专属 namespace
kubectl create namespace monitoring
# 安装,带自定义 values
helm install prometheus prometheus-community/kube-prometheus-stack \
--namespace monitoring \
--set grafana.adminPassword=admin123 \
--set prometheus.prometheusSpec.retention=15d \
--set alertmanager.alertmanagerSpec.retention=15d
参数解释:
grafana.adminPassword:Grafana 管理员密码,记得改prometheus.prometheusSpec.retention:指标保留天数,15d 够用- 默认会安装 Node Exporter(节点监控)、Kube State Metrics(K8s 资源监控)、Prometheus Operator(自动发现服务)
等几分钟,Pods 起来:
kubectl get pods -n monitoring
你应该看到类似这些:
alertmanager-prometheus-alertmanager-0 1/1 Running
kube-state-metrics-xxx 1/1 Running
prometheus-kube-prometheus-prometheus-0 2/2 Running
grafana-xxx 1/1 Running
node-exporter-xxx 1/1 Running
暴露 Grafana 和 Prometheus
# 端口转发,本地访问
kubectl port-forward -n monitoring svc/prometheus-kube-prometheus-prometheus 9090:9090 &
kubectl port-forward -n monitoring svc/grafana 3000:3000 &
浏览器打开 http://localhost:3000,账号 admin,密码你设置的 admin123。Prometheus 是 http://localhost:9090。
第二步:Grafana 里加数据源和看面板
添加 Prometheus 数据源
Grafana 左边菜单 → Connections → Add data source → 选 Prometheus → URL 填 http://prometheus-kube-prometheus-prometheus:9090(在集群内)→ Save & Test。
导入现成面板
kube-prometheus-stack 自带一堆面板,不用从零画。在 Grafana 主页,点 + → Import,搜这几个 ID:
- 1860:Kubernetes / Compute Resources / Pod
- 193:Kubernetes / Compute Resources / Node
- 13103:Kubernetes / Compute Resources / Cluster
- 6417:Kubernetes / Networking / Node
导入后,你会看到超详细的面板:CPU 使用率、内存、网络、磁盘等。点进去就能选 namespace、pod、container。
第三步:配置告警——抓 CPU 内存异常
告警是监控的灵魂。没告警,等于没监控。
Alertmanager 配置
Alertmanager 负责把 Prometheus 的告警发出去(邮件、钉钉、企业微信、Slack 等)。默认配置已经能工作,但我们需要自定义规则。
自定义 PrometheusRule
创建一个 YAML 文件 alerts.yaml:
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: k8s-alerts
namespace: monitoring
spec:
groups:
- name: container-alerts
rules:
- alert: ContainerCPUUsageHigh
expr: sum(rate(container_cpu_usage_seconds_total{container!=""}[5m])) by (pod, container, namespace) > 0.8
for: 5m
labels:
severity: warning
annotations:
summary: "CPU 使用率过高: {{ $labels.pod }} ({{ $labels.container }})"
description: "容器 {{ $labels.container }} 在 namespace {{ $labels.namespace }} 中 CPU 使用率超过 80% 持续 5 分钟。当前值: {{ $value | humanizePercentage }}"
- alert: ContainerMemoryUsageHigh
expr: (container_memory_working_set_bytes{container!=""} / container_spec_memory_limit_bytes{container!=""}) * 100 > 85
for: 5m
labels:
severity: warning
annotations:
summary: "内存使用率过高: {{ $labels.pod }} ({{ $labels.container }})"
description: "容器 {{ $labels.container }} 内存使用率超过 85%,即将触发 OOMKilled。当前值: {{ $value | humanize }}%"
- alert: PodCrashLooping
expr: rate(kube_pod_container_status_restarts_total{namespace!=""}[15m]) * 60 * 5 > 0
for: 10m
labels:
severity: critical
annotations:
summary: "Pod 频繁重启: {{ $labels.pod }}"
description: "Pod {{ $labels.pod }} 在最近 15 分钟内重启超过 5 次,可能崩溃循环。"
- alert: NodeCPUUsageHigh
expr: 100 - (avg by(instance) (rate(node_cpu_seconds_total{mode="idle"}[5m])) * 100) > 90
for: 10m
labels:
severity: critical
annotations:
summary: "节点 CPU 使用率过高: {{ $labels.instance }}"
description: "节点 {{ $labels.instance }} CPU 使用率超过 90% 持续 10 分钟。"
- alert: NodeMemoryAvailableLow
expr: (node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes) * 100 < 10
for: 5m
labels:
severity: critical
annotations:
summary: "节点内存不足: {{ $labels.instance }}"
description: "节点 {{ $labels.instance }} 可用内存低于 10%,可能导致 OOM。"
应用规则:
kubectl apply -f alerts.yaml
配置 Alertmanager 路由(发送钉钉/企业微信)
假设你用企业微信机器人。编辑 alertmanager-config.yaml(或者通过 Helm values 覆盖)。
在 Helm values 里加:
alertmanager:
alertmanagerSpec:
config:
global:
wechat_api_url: 'https://qyapi.weixin.qq.com/cgi-bin/'
wechat_api_secret: 'YOUR_SECRET'
wechat_api_corp_id: 'YOUR_CORP_ID'
templates:
- /etc/alertmanager/configmap/templates/*.tmpl
receivers:
- name: 'wechat'
wechat_configs:
- send_resolved: true
to_party: 'YOUR_PARTY_ID'
agent_id: 'YOUR_AGENT_ID'
message: '{{ template "wechat.default.message" . }}'
route:
receiver: 'wechat'
group_by: ['alertname', 'namespace', 'pod']
group_wait: 30s
group_interval: 5m
repeat_interval: 4h
模板文件 templates/default.tmpl:
{{ define "wechat.default.message" }}
{{ range .Alerts }}
【{{ .Labels.severity | upper }}】{{ .Annotations.summary }}
描述:{{ .Annotations.description }}
{{ end }}
{{ end }}
更新 Helm:
helm upgrade prometheus prometheus-community/kube-prometheus-stack \
--namespace monitoring \
--values values.yaml
重启后,告警就会发到企业微信群里。
第四步:实战排查——CPU 飙高怎么查
假设你收到告警:Pod my-app-7f9d4b6c8-x2k9m CPU 使用率 95%。
1. 进 Grafana 确认
打开 Pod 面板,选对应 namespace 和 pod,看 CPU 曲线。是突然 spike 还是持续高?持续高可能是正常业务,spike 可能是 bug 或 DDoS。
2. 登录 Pod 看进程
kubectl exec -it -n default my-app-7f9d4b6c8-x2k9m -- bash
里面用 top 或 htop(如果装了的话):
top -p $(pgrep -f my-app) # 找主进程 PID
看是哪个进程吃 CPU。如果是 Node.js,可能是死循环;如果是 Python,可能是计算密集型任务。
3. 如果是 Java 应用
# 找 JVM PID
jps -l
# 或者
pidstat -p <PID> 1 5
用 jstack <PID> 导出线程栈,看哪个线程占 CPU 高:
jstack <PID> | grep -A 10 "RUNNABLE" | head -50
结合代码定位。
4. 查资源限制
kubectl describe pod my-app-7f9d4b6c8-x2k9m -n default
看 Limits 和 Requests。如果 CPU limit 设了 500m,但实际用到 900m,会被 Throttled,表现为 CPU 使用率图“锯齿状”——其实是被限流了,不是真吃这么多。这时调大 limit 或优化代码。
第五步:内存异常怎么查
告警:Pod memory 使用率 90%,接近 limit。
1. 看内存趋势
Grafana 里看 container_memory_working_set_bytes。Working Set 是实际物理内存,不含 cache。如果 cache 高,不用慌,系统会自动回收。
2. 进 Pod 查
kubectl exec -it -n default my-app-7f9d4b6c8-x2k9m -- bash
free -h # 看整体内存
cat /sys/fs/cgroup/memory/memory.usage_in_bytes # cgroup 实际内存
3. 泄漏排查
Go 程序:用
pprofkubectl port-forward -n default my-app-7f9d4b6c8-x2k9m 6060:6060 # 本地浏览器打开 http://localhost:6060/debug/pprof/heapJava 程序:
jmap -histo <PID>看对象分布,或用 MAT 分析 dump 文件。Python 程序:用
tracemalloc或memory_profiler。
4. OOMKilled 了怎么办
如果 Pod 频繁重启,看事件:
kubectl describe pod my-app-7f9d4b6c8-x2k9m
最后几行如果有 OOMKilled,说明内存真的爆了。调大 memory limit,或者优化代码减少内存占用。注意:limit 设太大,会挤占其他 Pod,可能导致节点整体 OOM。
第六步:一些实用小技巧
用 kubectl top 快速查看
kubectl top pods -A --sort-by=cpu # 按 CPU 排序
kubectl top pods -A --sort-by=memory # 按内存排序
导出 Prometheus 指标到 CSV 分析
在 Prometheus 网页,用 PromQL 查询,点“Export as CSV”,本地用 Excel 或 Python pandas 分析。
用 Loki 配合 Promtail 看日志
监控 + 日志 = 王炸。Loki 和 Prometheus 集成,可以直接从 Grafana 跳转到对应时间点的日志。
helm install loki prometheus-community/loki-distributed \
--namespace monitoring
自动化:用 Argo CD 管理监控配置
把上面的 YAML 都放进 Git 仓库,用 Argo CD 同步,改告警规则直接推代码,别人审计也方便。
最后说两句
监控这事儿,一开始觉得复杂,其实就三步:采集 → 展示 → 告警。Prometheus + Grafana 这套,采集靠 Kube State Metrics 和 Node Exporter,展示靠现成面板,告警靠自定义规则。你按我说的顺序来,半小时就能跑起来。
记住,告警不要设太多,否则会被“告警疲劳”。先抓 CPU 内存这类硬指标,再慢慢加业务指标。排查时,从 Grafana 看趋势,再用 kubectl 进容器查细节,基本能解决 90% 的问题。
有问题随时问,监控这块我折腾过不少坑,帮你避坑。祝你集群稳稳的,半夜不用被告警吵醒 😄
