Kubernetes集群CPU飙升90导致服务宕机 资深运维用Prometheus加Grafana搭建监控体系 从Pod资源水位到容器异常告警 手把手教你用开源工具搞定生产环境容器监控
说实话,我第一次遇到K8s集群CPU飙升到90%的时候,整个人都懵了。
那是凌晨两点,钉钉突然狂震,业务方电话打爆。上线才三天的服务,CPU直接从40%跳到90%以上,几个核心Pod疯狂重启,用户端开始大面积报错。等我们冲进战况室的时候,集群已经处于半瘫痪状态。
复盘那晚的”事故”,最痛的一个点就是:我们完全不知道发生了什么,直到业务方来找我们。
没有实时监控,没有预警,没有可视化的资源水位数据。等到发现问题的时侯,已经晚了。
从那之后,我下定决心:必须搭建一套完整的、可视化的、能提前预警的监控体系。今天,我就把这个过程,完整分享给你。
一、先搞清楚:为什么容器监控这么重要?
很多刚接触K8s的运维朋友,会觉得”容器不就是跑起来吗?”
但生产环境不是开发环境。在K8s集群里,你有几十甚至几百个Pod在跑,每个Pod里可能有多个容器,每个容器都在消耗CPU、内存、磁盘、网络资源。如果没有监控,你就像在迷雾中开车——完全不知道前面有什么。
具体来说,监控要解决三个核心问题:
第一,资源水位。你的CPU、内存用了多少?还剩多少?什么时候会爆?
第二,服务健康。Pod有没有崩?容器有没有异常退出?服务能不能正常响应?
第三,异常预警。在问题发生之前,你能不能提前感知到?比如CPU持续上升、内存泄漏、接口响应变慢。
这三个问题,缺一不可。
而我们要用的工具,就是两个开源神器:Prometheus 和 Grafana。
二、Prometheus:监控数据的”采集器”
Prometheus是由SoundCloud开发,现在是CNCF(云原生计算基金会)的毕业项目。它的核心思路很简单:把所有需要监控的数据,都变成”指标”(Metrics),然后定时采集、存储、查询。
2.1 核心概念
在深入配置之前,你先要理解Prometheus的几个核心概念:
指标(Metric):这是最基本的数据单元。比如cpu_usage表示CPU使用率,memory_usage表示内存使用量。
时间序列(Time Series):每个指标随时间变化的数据,就叫时间序列。比如”过去一小时,每个Pod的CPU使用率”。
拉取模型(Pull Model):Prometheus主动去目标节点”拉取”数据,而不是目标节点”推送”数据。这个设计有个好处:如果某个节点挂了,Prometheus拉不到数据,自然就知道它出问题了。
PromQL:这是Prometheus的查询语言,专门用来查询和聚合时间序列数据。
2.2 在K8s中部署Prometheus
好,概念理解了,我们开始部署。
首先,我们需要用Helm来安装Prometheus,这是最简单的方式。
# 添加Prometheus社区仓库
helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
helm repo update
# 创建monitoring命名空间
kubectl create namespace monitoring
# 安装Prometheus
helm install prometheus prometheus-community/kube-prometheus-stack \
--namespace monitoring \
--set prometheus.prometheusSpec.retention=7d \
--set grafana.enabled=true \
--set alertmanager.enabled=true
这段命令做了四件事:
- 添加Helm仓库
- 创建
monitoring命名空间 - 安装
kube-prometheus-stack,这个Chart包含了Prometheus、Grafana、Alertmanager等多个组件 - 设置数据保留7天,开启Grafana和Alertmanager
安装完成后,验证一下:
kubectl get pods -n monitoring
你应该能看到类似这样的输出:
NAME READY STATUS RESTARTS AGE
prometheus-kube-prometheus-prometheus-0 2/2 Running 0 5m
grafana-7d9f8b6c5-x2k9m 1/1 Running 0 5m
alertmanager-prometheus-alertmanager-0 2/2 Running 0 5m
如果状态都是Running,说明Prometheus已经部署成功了。
三、Grafana:让数据”看得见”
Prometheus负责采集和存储数据,Grafana负责展示数据。
3.1 访问Grafana
Grafana默认在集群内部运行,你需要通过端口转发或者Ingress来访问它。
最简单的办法是用kubectl port-forward:
kubectl port-forward -n monitoring svc/grafana 3000:80
然后在浏览器打开 http://localhost:3000。
默认账号是:
- 用户名:
admin - 密码:在
admin-password这个Secret里
kubectl get secret -n monitoring grafana-admin-password -o jsonpath='{.data.password}' | base64 -d
3.2 配置数据源
进入Grafana后,首先需要配置数据源,告诉Grafana去哪里拿数据。
点击左侧菜单的 Connections → Data Sources → Add data source → Prometheus。
填入Prometheus的URL:
http://prometheus-kube-prometheus-prometheus.monitoring.svc.cluster.local:9090
然后点 Save & Test,看到”Data source is working”就说明连接成功了。
3.3 导入现成Dashboard
Grafana最强大的地方,就是有很多现成的Dashboard可以导入,不用从零开始画。
点击 Dashboards → Browse → Import,输入ID 315(这是Kubernetes集群监控的经典Dashboard)。
点击 Load,选择Prometheus作为数据源,然后点击 Import。
等几秒钟,你会看到一个非常专业的集群监控面板,包含:
- 集群整体CPU使用率
- 集群整体内存使用率
- 各Node的资源使用情况
- 各Namespace的Pod数量
- 网络流量
- 存储使用情况
- 等等
这些图表,基本覆盖了一个运维人员需要关注的所有核心指标。
四、核心监控指标:你应该关注什么?
有了Dashboard,你可能会问:到底哪些指标最关键?
我根据你的实战经验,给你梳理一下:
4.1 CPU相关指标
| 指标 | 含义 | 告警阈值建议 |
|---|---|---|
container_cpu_usage_seconds_total |
容器累计CPU使用量 | - |
container_cpu_usage_seconds_total / limit |
CPU使用率 | 超过80%预警 |
rate(container_cpu_usage_seconds_total[5m]) |
过去5分钟CPU平均使用率 | - |
4.2 内存相关指标
| 指标 | 含义 | 告警阈值建议 |
|---|---|---|
container_memory_working_set_bytes |
容器实际使用的内存 | - |
container_memory_working_set_bytes / limit |
内存使用率 | 超过85%预警 |
container_memory_failures_total |
内存失败次数(OOM) | 大于0就告警 |
4.3 网络相关指标
| 指标 | 含义 |
|---|---|
container_network_transmit_bytes_total |
网络发送字节数 |
container_network_receive_bytes_total |
网络接收字节数 |
container_network_transmit_packets_dropped_total |
发送丢包数 |
container_network_receive_packets_dropped_total |
接收丢包数 |
4.4 服务健康相关指标
| 指标 | 含义 |
|---|---|
kube_pod_status_phase |
Pod状态(Running、Pending、Failed等) |
kube_pod_container_status_restarts_total |
容器重启次数 |
kube_deployment_status_replicas_available |
可用副本数 |
kube_pod_init_container_status_restarts_total |
Init容器重启次数 |
五、配置告警:让问题在爆发前被发现
监控的最终目的,不是为了看数据,而是为了在问题发生之前发现它。
Alertmanager就是负责告警的核心组件。
5.1 创建告警规则
首先,我们需要定义”什么情况下应该告警”。这个定义,写在PrometheusRule里。
创建一个文件alert-rules.yaml:
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: cluster-alerts
namespace: monitoring
labels:
prometheus: kube-prometheus
role: alert-rules
spec:
groups:
- name: cluster-cpu-alerts
rules:
- alert: HighCPUUsage
expr: sum(rate(container_cpu_usage_seconds_total{container!=""}[5m])) by (namespace, pod) > 0.8
for: 5m
labels:
severity: warning
team: infrastructure
annotations:
summary: "Pod {{ $labels.pod }} in namespace {{ $labels.namespace }} is using high CPU"
description: "Pod {{ $labels.pod }} CPU usage is {{ $value | humanizePercentage }}, threshold is 80%"
- alert: CriticalCPUUsage
expr: sum(rate(container_cpu_usage_seconds_total{container!=""}[5m])) by (namespace, pod) > 0.9
for: 3m
labels:
severity: critical
team: infrastructure
annotations:
summary: "Pod {{ $labels.pod }} in namespace {{ $labels.namespace }} is in CPU critical state"
description: "Pod {{ $labels.pod }} CPU usage is {{ $value | humanizePercentage }}, threshold is 90%"
- name: cluster-memory-alerts
rules:
- alert: HighMemoryUsage
expr: sum(container_memory_working_set_bytes{container!=""}) by (namespace, pod) / sum(container_spec_memory_limit_bytes{container!=""}) by (namespace, pod) > 0.85
for: 5m
labels:
severity: warning
team: infrastructure
annotations:
summary: "Pod {{ $labels.pod }} in namespace {{ $labels.namespace }} is using high memory"
description: "Pod {{ $labels.pod }} memory usage is {{ $value | humanizePercentage }}, threshold is 85%"
- alert: ContainerOOMKilled
expr: increase(container_memory_failures_total{action="oom"}[1m]) > 0
for: 0m
labels:
severity: critical
team: infrastructure
annotations:
summary: "Pod {{ $labels.pod }} in namespace {{ $labels.namespace }} was OOM killed"
description: "Pod {{ $labels.pod }} exceeded its memory limit and was killed by the kernel"
- name: cluster-pod-alerts
rules:
- alert: PodCrashLooping
expr: rate(kube_pod_container_status_restarts_total[15m]) * 60 * 15 > 0
for: 5m
labels:
severity: warning
team: infrastructure
annotations:
summary: "Pod {{ $labels.pod }} in namespace {{ $labels.namespace }} is crash looping"
description: "Pod {{ $labels.pod }} has been restarting frequently"
- alert: PodNotRunning
expr: kube_pod_status_phase{phase!="Running"} == 1
for: 5m
labels:
severity: warning
team: infrastructure
annotations:
summary: "Pod {{ $labels.pod }} in namespace {{ $labels.namespace }} is not running"
description: "Pod {{ $labels.pod }} is in state {{ $labels.phase }}"
- name: cluster-node-alerts
rules:
- alert: NodeNotReady
expr: kube_node_status_condition{condition="Ready",status="true"} == 0
for: 5m
labels:
severity: critical
team: infrastructure
annotations:
summary: "Node {{ $labels.node }} is not ready"
description: "Node {{ $labels.node }} has been in NotReady state for more than 5 minutes"
- alert: HighNodeCPU
expr: 1 - sum(rate(node_cpu_seconds_total{mode="idle"}[5m])) by (instance) / sum(node_cpu_seconds_total) by (instance) > 0.85
for: 10m
labels:
severity: warning
team: infrastructure
annotations:
summary: "Node {{ $labels.instance }} is under high CPU load"
description: "Node {{ $labels.instance }} CPU usage is {{ $value | humanizePercentage }}, threshold is 85%"
然后应用这个规则:
kubectl apply -f alert-rules.yaml
5.2 配置告警通知渠道
告警规则定义好了,还需要告诉Alertmanager”告警来了往哪里发”。
我们创建一个配置文件alertmanager-config.yaml:
apiVersion: monitoring.coreos.com/v1alpha1
kind: AlertmanagerConfig
metadata:
name: cluster-alerts
namespace: monitoring
labels:
alertmanager: kube-prometheus-stack
spec:
route:
receiver: wechat-work
group_by: ['alertname', 'namespace', 'pod']
group_wait: 30s
group_interval: 5m
repeat_interval: 4h
routes:
- match:
severity: critical
receiver: wechat-work-critical
continue: true
- match:
severity: warning
receiver: wechat-work
continue: false
receivers:
- name: wechat-work
wechat_configs:
- corp_id: 'your-corp-id'
to_user: '@all'
agent_id: 'your-agent-id'
api_secret: 'your-api-secret'
message: |
🚨 集群告警
告警名称: {{ .CommonLabels.alertname }}
严重程度: {{ .CommonLabels.severity }}
命名空间: {{ .CommonLabels.namespace }}
Pod: {{ .CommonLabels.pod }}
详情: {{ range .Alerts }}{{ .Annotations.description }}{{ end }}
时间: {{ .StartsAt }}
- name: wechat-work-critical
wechat_configs:
- corp_id: 'your-corp-id'
to_user: '你的手机号'
agent_id: 'your-agent-id'
api_secret: 'your-api-secret'
message: |
🔴 严重告警
告警名称: {{ .CommonLabels.alertname }}
严重程度: {{ .CommonLabels.severity }}
命名空间: {{ .CommonLabels.namespace }}
Pod: {{ .CommonLabels.pod }}
详情: {{ range .Alerts }}{{ .Annotations.description }}{{ end }}
时间: {{ .StartsAt }}
⚠️ 请立即处理!
注意:上面的企业微信配置,你需要替换成你自己的:
corp_id:企业IDagent_id:应用IDapi_secret:应用密钥to_user:接收人的手机号
如果你用的是钉钉,也可以换成钉钉的webhook配置:
- name: dingtalk
dingtalk_configs:
- send_resolved: true
webhook: 'https://oapi.dingtalk.com/robot/send?access_token=your-token'
message: |
【集群告警】
告警名称: {{ .CommonLabels.alertname }}
严重程度: {{ .CommonLabels.severity }}
详情: {{ range .Alerts }}{{ .Annotations.description }}{{ end }}
时间: {{ .StartsAt }}
应用这个配置:
kubectl apply -f alertmanager-config.yaml
5.3 验证告警是否生效
要验证告警是否正常工作,最快的方法是制造一个告警。
比如,我们可以把某个Pod的CPU请求调得很低,让它更容易触发CPU告警:
# 编辑某个Deployment,给容器加上更严格的CPU限制
kubectl edit deployment your-app -n your-namespace
在spec里找到containers,加上:
resources:
limits:
cpu: "100m"
memory: "128Mi"
requests:
cpu: "50m"
memory: "64Mi"
保存后,如果Pod实际CPU使用超过了100m,就会触发告警。
你也可以直接查看Alertmanager的状态:
# 查看Prometheus的告警规则
kubectl port-forward -n monitoring svc/prometheus-kube-prometheus-prometheus 9090:9090
然后访问 http://localhost:9090/alerts,可以看到所有告警规则的状态。
六、实战:从监控到排障的完整流程
现在,我们已经有了监控和告警,但真正考验运维能力的,是告警来了之后,怎么快速定位和解决问题。
让我模拟一次真实的排障过程。
场景:CPU告警触发
凌晨3点,你的钉钉疯狂响起:
🔴 严重告警 告警名称: HighCPUUsage 严重程度: critical 命名空间: production Pod: api-gateway-7d9f8b6c5-x2k9m 详情: Pod api-gateway-7d9f8b6c5-x2k9m CPU usage is 95%, threshold is 90% 时间: 2024-01-15T03:00:00Z ⚠️ 请立即处理!
第一步:在Grafana中查看趋势
你打开Grafana,找到这个Pod的CPU使用率图表:
# 在Prometheus中查询
sum(rate(container_cpu_usage_seconds_total{pod="api-gateway-7d9f8b6c5-x2k9m",container!="",namespace="production"}[5m])) by (pod)
你会看到,CPU是从凌晨2点半开始,慢慢爬升到95%的。这意味着不是突发的流量攻击,而是一个渐进的过程。
第二步:查看相关指标
接下来,你需要看看其他指标有没有异常:
内存使用情况:
sum(container_memory_working_set_bytes{pod="api-gateway-7d9f8b6c5-x2k9m",container!="",namespace="production"}) by (pod)
如果内存也在持续上升,可能是内存泄漏导致频繁的GC,进而消耗CPU。
接口响应时间:
histogram_quantile(0.99, sum(rate(http_request_duration_seconds_bucket{exported_pod="api-gateway-7d9f8b6c5-x2k9m"}[5m])) by (le))
如果P99响应时间也在飙升,说明服务已经受到了影响。
Pod重启次数:
kube_pod_container_status_restarts_total{pod="api-gateway-7d9f8b6c5-x2k9m"}
如果重启次数也在增加,说明问题可能在恶化。
第三步:进入Pod内部排查
在Grafana中,你可以直接点击进入Pod的详情页,或者SSH到Node上执行命令:
# 进入Pod内部
kubectl exec -it api-gateway-7d9f8b6c5-x2k9m -n production -- /bin/sh
# 查看容器内进程
top -p $(pgrep -f api-gateway)
# 查看CPU使用最高的进程
ps aux | sort -nrk 3 | head -10
# 查看CPU详细的系统调用
strace -p $(pgrep -f api-gateway) -c
通过这个过程中,你可能会发现:某个函数调用次数异常多,或者某个定时任务执行频率异常高。
第四步:定位代码问题
如果进入Pod后发现是某个进程异常,你可以通过以下方式定位到代码:
查看日志:
kubectl logs api-gateway-7d9f8b6c5-x2k9m -n production --tail=1000
查看慢查询日志(如果是数据库操作):
kubectl exec -it api-gateway-7d9f8b6c5-x2k9m -n production -- cat /var/log/app/slow-query.log
查看APM追踪(如果你接入了Jaeger或SkyWalking):
# Jaeger追踪
kubectl port-forward svc/jaeger-query 16686:80 -n monitoring
第五步:临时止血
在定位到根本原因之前,你可能需要先”止血”:
方案一:水平扩容
kubectl scale deployment api-gateway -n production --replicas=6
通过增加Pod数量,分散CPU压力。
方案二:紧急重启
kubectl rollout restart deployment api-gateway -n production
如果是内存泄漏导致的CPU飙升,重启可以暂时缓解问题。
方案三:降级服务
如果某些非核心功能可以暂时下线,可以通过Feature Flag关闭它们,减少CPU消耗。
第六步:根因分析和长期解决
止血之后,你需要真正解决这个问题:
- 分析代码:找到CPU高的原因(死循环、频繁GC、不必要的计算等)
- 修复代码:提交修复,重新构建镜像
- 灰度发布:先灰度发布到一小部分Pod,观察CPU是否恢复正常
- 全量发布:确认无误后,全量发布
- 复盘:写一份事故复盘文档,记录问题原因、处理过程、后续改进措施
七、高级技巧:让监控更智能
上面的内容,已经能解决大部分问题了。但如果你想要更智能的监控,还有几个进阶玩法:
7.1 使用黑盒监控探测服务可用性
Prometheus默认是”拉取”模型,但有些服务是在内网里的,外网探测不到。这时候可以用黑盒监控(Blackbox Exporter)。
# 安装blackbox-exporter
helm install blackbox prometheus-community/blackbox-exporter --namespace monitoring
然后创建一个Probe资源:
apiVersion: monitoring.coreos.com/v1
kind: Probe
metadata:
name: api-gateway-probe
namespace: monitoring
spec:
module: http_2xx
prober:
url: blackbox-exporter.monitoring.svc.cluster.local:9115
targets:
staticConfig:
static:
- https://api.example.com/health
interval: 30s
这样,Prometheus就可以定期探测你的服务是否可用,而不仅仅是在集群内部监控。
7.2 使用Node Exporter监控节点级指标
除了容器级别的监控,你还需要监控节点本身:
# Node Exporter通常已经通过kube-prometheus-stack部署了
# 如果没有,可以单独安装
helm install node-exporter prometheus-community/node-exporter --namespace monitoring
Node Exporter会收集节点级的指标:CPU、内存、磁盘、网络、负载等。
在Grafana中,你可以导入 Node Exporter 的Dashboard(ID 1860),就能看到非常详细的节点监控信息。
7.3 使用cAdvisor监控容器详情
cAdvisor是Google开发的容器监控工具,K8s默认已经集成。它会收集每个容器的:
- CPU、内存、磁盘、网络使用量
- 文件系统使用量
- 网络吞吐量
- 进程数
- 文件系统I/O
- 等等
在Prometheus中,你可以查询:
# 某个容器的CPU使用率
rate(container_cpu_usage_seconds_total{container="api-gateway",pod="api-gateway-7d9f8b6c5-x2k9m"}[5m])
# 某个容器的内存使用量
container_memory_working_set_bytes{container="api-gateway",pod="api-gateway-7d9f8b6c5-x2k9m"}
# 某个容器的网络接收字节数
rate(container_network_receive_bytes_total{container="api-gateway",pod="api-gateway-7d9f8b6c5-x2k9m"}[5m])
7.4 使用Pod Autoscaler实现自动扩缩容
有了监控数据,你甚至可以配置自动扩缩容,让系统在CPU高的时候自动增加Pod数量:
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: api-gateway-hpa
namespace: production
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: api-gateway
minReplicas: 3
maxReplicas: 20
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 70
这个配置的意思是:当API网关Pod的平均CPU使用率超过70%时,自动增加Pod数量,最多增加到20个。
7.5 使用日志监控配合
除了指标监控,日志监控也很重要。你可以用EFK(Elasticsearch + Fluentd + Kibana)或者Loki(轻量级日志系统)来收集和分析日志。
Loki的配置相对简单:
helm install loki grafana/loki-distributed --namespace monitoring
然后通过Promtail采集日志,发送到Loki:
apiVersion: v1
kind: ConfigMap
metadata:
name: promtail-config
namespace: monitoring
data:
promtail.yaml: |
clients:
- url: http://loki-gateway/loki/api/v1/push
scrape_configs:
- job_name: kubernetes-pods
kubernetes_sd_configs:
- role: pod
relabel_configs:
- source_labels: [__meta_kubernetes_pod_label_app]
target_label: app
八、一些实战中的坑和注意事项
写了这么多,我再说几个实际踩过的坑:
8.1 告警风暴
刚开始配置告警的时候,最容易遇到的问题就是告警风暴:一条根因触发了一堆子告警,你的手机从早响到晚。
解决方法:
- 使用
group_by把相关告警合并 - 设置
group_wait和repeat_interval,避免短时间内重复告警 - 配置告警静默,比如凌晨时段只发严重告警
8.2 指标丢失
有时候你会发现,某些Pod的指标在Grafana里消失了。这通常是因为:
- Pod被删除了(正常情况)
- Node出了问题,Kubelet挂了(需要排查Node)
- Prometheus的抓取配置有问题(检查ServiceMonitor或PodMonitor)
8.3 存储不够
Prometheus是时序数据库,数据量会持续增长。如果数据保留时间太长,磁盘会爆。
解决方法:
- 设置合适的
retention时间(比如7天或30天) - 使用对象存储(S3、OSS)做长期存储
- 使用Thanos或Cortex做分布式存储
8.4 误报和漏报
告警阈值设置得太敏感,会有大量误报;设置得太宽松,又会漏报。
解决方法:
- 根据历史数据,设置合理的阈值
- 使用动态阈值(比如过去7天的平均值加2倍标准差)
- 持续优化告警规则,根据实际运行情况调整
九、总结
回到最开始的那个夜晚。
如果当时我们有这套监控体系,凌晨两点就能在钉钉上收到告警,看到CPU在3点半开始爬升,然后在4点之前就介入处理,而不是等到业务方打爆电话才发现。
监控不是为了”看数据”,而是为了提前发现问题、快速定位问题、及时解决问题。
Prometheus + Grafana 这套组合,已经是云原生时代的事实标准。虽然学习曲线不算平缓,但只要掌握了核心概念和常用操作,就能快速上手。
希望这篇文章,能帮你搭建起属于自己的监控体系。如果有什么疑问,欢迎随时交流。
