From 3ca030c235cb148e82d915b1db042c5922e75480 Mon Sep 17 00:00:00 2001 From: howieyuen Date: Thu, 6 May 2021 11:59:12 +0800 Subject: [PATCH] [zh] Resync tasks files[3] --- .../monitor-node-health.md | 360 +++++++++--------- 1 file changed, 181 insertions(+), 179 deletions(-) diff --git a/content/zh/docs/tasks/debug-application-cluster/monitor-node-health.md b/content/zh/docs/tasks/debug-application-cluster/monitor-node-health.md index b34d1040d9..4e81fd8567 100644 --- a/content/zh/docs/tasks/debug-application-cluster/monitor-node-health.md +++ b/content/zh/docs/tasks/debug-application-cluster/monitor-node-health.md @@ -3,159 +3,131 @@ content_type: task title: 节点健康监测 --- -*节点问题探测器* 是一个 [DaemonSet](/zh/docs/concepts/workloads/controllers/daemonset/), -用来监控节点健康。它从各种守护进程收集节点问题,并以 +*节点问题检测器(Node Problem Detector)*是一个守护程序,用于监视和报告节点的健康状况。 +你可以将节点问题探测器以 `DaemonSet` 或独立守护程序运行。 +节点问题检测器从各种守护进程收集节点问题,并以 [NodeCondition](/zh/docs/concepts/architecture/nodes/#condition) 和 [Event](/docs/reference/generated/kubernetes-api/{{< param "version" >}}/#event-v1-core) 的形式报告给 API 服务器。 - -它现在支持一些已知的内核问题检测,并将随着时间的推移,检测更多节点问题。 - - -目前,Kubernetes 不会对节点问题检测器监测到的节点状态和事件采取任何操作。 -将来可能会引入一个补救系统来处理这些节点问题。 - - -更多信息请参阅 [这里](https://github.com/kubernetes/node-problem-detector)。 +要了解如何安装和使用节点问题检测器,请参阅 +[节点问题探测器项目文档](https://github.com/kubernetes/node-problem-detector)。 ## {{% heading "prerequisites" %}} -{{< include "task-tutorial-prereqs.md" >}} {{< version-check >}} +{{< include "task-tutorial-prereqs.md" >}} ## 局限性 {#limitations} -* 节点问题检测器的内核问题检测现在只支持基于文件类型的内核日志。 +* 节点问题检测器只支持基于文件类型的内核日志。 它不支持像 journald 这样的命令行日志工具。 +* 节点问题检测器使用内核日志格式来报告内核问题。 + 要了解如何扩展内核日志格式,请参阅[添加对另一个日志格式的支持](#support-other-log-format)。 -* 节点问题检测器的内核问题检测对内核日志格式有一定要求,现在它只适用于 Ubuntu 和 Debian。 - 不过将其扩展为[支持其它日志格式](#support-other-log-format) 也很容易。 +## 启用节点问题检测器 + +一些云供应商将节点问题检测器以{{< glossary_tooltip text="插件" term_id="addons" >}}形式启用。 +你还可以使用 `kubectl` 或创建插件 Pod 来启用节点问题探测器。 -## 在 GCE 集群中启用/禁用 +## 使用 kubectl 启用节点问题检测器 {#using-kubectl} -节点问题检测器在 gce 集群中以 -[集群插件的形式](/zh/docs/setup/best-practices/cluster-large/#addon-resources) -默认启用。 +`kubectl` 提供了节点问题探测器最灵活的管理。 +你可以覆盖默认配置使其适合你的环境或检测自定义节点问题。例如: -你可以在运行 `kube-up.sh` 之前,以设置环境变量 `KUBE_ENABLE_NODE_PROBLEM_DETECTOR` 的形式启用/禁用它。 +1. 创建类似于 `node-strought-detector.yaml` 的节点问题检测器配置: + {{< codenew file="debug/node-problem-detector.yaml" >}} + + {{< note >}} + 你应该检查系统日志目录是否适用于操作系统发行版本。 + {{< /note >}} + +1. 使用 `kubectl` 启动节点问题检测器: + + ```shell + kubectl apply -f https://k8s.io/examples/debug/node-problem-detector.yaml + ``` -## 在其它环境中使用 {#use-in-other-environment} +### 使用插件 pod 启用节点问题检测器 {#using-addon-pod} -要在 GCE 之外的其他环境中启用节点问题检测器,你可以使用 `kubectl` 或插件 pod。 +如果你使用的是自定义集群引导解决方案,不需要覆盖默认配置, +可以利用插件 Pod 进一步自动化部署。 - -### Kubectl - -这是在 GCE 之外启动节点问题检测器的推荐方法。 -它的管理更加灵活,例如覆盖默认配置以使其适合你的环境或检测自定义节点问题。 - - -* **步骤 1:** `node-problem-detector.yaml`: - -{{< codenew file="debug/node-problem-detector.yaml" >}} - - -***请注意保证你的系统日志路径与你的 OS 发行版相对应。*** - - -* **步骤 2:** 执行 `kubectl` 来启动节点问题检测器: - -```shell - kubectl create -f https://k8s.io/examples/debug/node-problem-detector.yaml -``` - - -### 插件 Pod {#addon-pod} - -这适用于拥有自己的集群引导程序解决方案的用户,并且不需要覆盖默认配置。 -他们可以利用插件 Pod 进一步自动化部署。 - - -只需创建 `node-problem-detector.yaml`,并将其放在主节点上的插件 pod 目录 -`/etc/kubernetes/addons/node-problem-detector` 下。 +创建 `node-strick-detector.yaml`,并在控制平面节点上保存配置到插件 Pod 的目录 +`/etc/kubernetes/addons/node-problem-detector`。 ## 覆盖配置文件 @@ -163,73 +135,97 @@ is embedded when building the docker image of node problem detector. [默认配置](https://github.com/kubernetes/node-problem-detector/tree/v0.1/config)。 -不过,你可以像下面这样使用 [ConfigMap](/zh/docs/tasks/configure-pod-container/configure-pod-configmap/) +不过,你可以像下面这样使用 [`ConfigMap`](/zh/docs/tasks/configure-pod-container/configure-pod-configmap/) 将其覆盖: -* **步骤 1:** 在 `config/` 中更改配置文件。 -* **步骤 2:** 使用 `kubectl create configmap node-problem-detector-config --from-file=config/` 创建 `node-problem-detector-config` 。 -* **步骤 3:** 更改 `node-problem-detector.yaml` 以使用 ConfigMap: +1. 更改 `config/` 中的配置文件 +1. 创建 `ConfigMap` `node-strick-detector-config`: + + ```shell + kubectl create configmap node-problem-detector-config --from-file=config/ + ``` -{{< codenew file="debug/node-problem-detector-configmap.yaml" >}} +1. 更改 `node-problem-detector.yaml` 以使用 ConfigMap: + + {{< codenew file="debug/node-problem-detector-configmap.yaml" >}} - -* **步骤 4:** 使用新的 yaml 文件重新创建节点问题检测器: +{{< note >}} +此方法仅适用于通过 `kubectl` 启动的节点问题检测器。 +{{< /note >}} -```shell - kubectl delete -f https://k8s.io/examples/debug/node-problem-detector.yaml # If you have a node-problem-detector running - kubectl create -f https://k8s.io/examples/debug/node-problem-detector-configmap.yaml -``` - - -***请注意,此方法仅适用于通过 `kubectl` 启动的节点问题检测器。*** - - -由于插件管理器不支持ConfigMap,因此现在不支持对于作为集群插件运行的节点问题检测器的配置进行覆盖。 +如果节点问题检测器作为集群插件运行,则不支持覆盖配置。 +插件管理器不支持 `ConfigMap`。 ## 内核监视器 -*内核监视器* 是节点问题检测器中的问题守护进程。它监视内核日志并按照预定义规则检测已知内核问题。 +*内核监视器(Kernel Monitor)*是节点问题检测器中支持的系统日志监视器守护进程。 +内核监视器观察内核日志并根据预定义规则检测已知的内核问题。 -内核监视器根据 [`config/kernel-monitor.json`](https://github.com/kubernetes/node-problem-detector/blob/v0.1/config/kernel-monitor.json) 中的一组预定义规则列表匹配内核问题。 +内核监视器根据 [`config/kernel-monitor.json`](https://github.com/kubernetes/node-problem-detector/blob/v0.1/config/kernel-monitor.json) +中的一组预定义规则列表匹配内核问题。 规则列表是可扩展的,你始终可以通过覆盖配置来扩展它。 ### 添加新的 NodeCondition -你可以使用新的状态描述来扩展 `config/kernel-monitor.json` 中的 `conditions` 字段以支持新的节点状态。 +要支持新的 `NodeCondition`,请在 `config/kernel-monitor.json` 中的 +`conditions` 字段中创建一个条件定义: ```json { @@ -240,14 +236,14 @@ To support new node conditions, you can extend the `conditions` field in ``` ### 检测新的问题 -你可以使用新的规则描述来扩展 `config/kernel-monitor.json` 中的 `rules` 字段以检测新问题。 +你可以使用新的规则描述来扩展 `config/kernel-monitor.json` 中的 `rules` 字段以检测新问题: ```json { @@ -259,51 +255,57 @@ with new rule definition: ``` -### 更改日志路径 +### 配置内核日志设备的路径 {#kernel-log-device-path} -不同操作系统发行版的内核日志的可能不同。 `config/kernel-monitor.json` 中的 `log` 字段是容器内的日志路径。你始终可以修改配置使其与你的 OS 发行版匹配。 +检查你的操作系统(OS)发行版本中的内核日志路径位置。 +Linux 内核[日志设备](https://www.kernel.org/doc/documentation/abi/testing/dev-kmsg) +通常呈现为 `/dev/kmsg`。 +但是,日志路径位置因 OS 发行版本而异。 +`config/kernel-monitor.json` 中的 `log` 字段表示容器内的日志路径。 +你可以配置 `log` 字段以匹配节点问题检测器所示的设备路径。 -### 支持其它日志格式 {#support-other-log-format} +### 添加对其它日志格式的支持 {#support-other-log-format} -内核监视器使用 [`Translator`] 插件将内核日志转换为内部数据结构。 -我们可以很容易为新的日志格式实现新的翻译器。 +内核监视器使用 +[`Translator`](https://github.com/kubernetes/node-problem-detector/blob/v0.1/pkg/kernelmonitor/translator.go) +插件转换内核日志的内部数据结构。 +你可以为新的日志格式实现新的转换器。 -## 注意事项 {#caveats} - -我们建议在集群中运行节点问题检测器来监视节点运行状况。 -但是,你应该知道这将在每个节点上引入额外的资源开销。一般情况下没有影响,因为: - - -* 内核日志生成相对较慢。 -* 节点问题检测器有资源限制。 -* 即使在高负载下,资源使用也是可以接受的。 -(参阅 [基准测试结果](https://github.com/kubernetes/node-problem-detector/issues/2#issuecomment-220255629)) +## 建议和限制 +建议在集群中运行节点问题检测器以监控节点运行状况。 +运行节点问题检测器时,你可以预期每个节点上的额外资源开销。 +通常这是可接受的,因为: +* 内核日志增长相对缓慢。 +* 已经为节点问题检测器设置了资源限制。 +* 即使在高负载下,资源使用也是可接受的。有关更多信息,请参阅节点问题检测器 + [基准结果](https://github.com/kubernetes/node-problem-detector/issues/2.suecomment-220255629)。