Merge pull request #31372 from sftim/20220117_tidy_kubeadm_ha
Tidy kubeadm HA guide
This commit is contained in:
@@ -1,7 +1,7 @@
|
|||||||
---
|
---
|
||||||
reviewers:
|
reviewers:
|
||||||
- sig-cluster-lifecycle
|
- sig-cluster-lifecycle
|
||||||
title: Options for Highly Available topology
|
title: Options for Highly Available Topology
|
||||||
content_type: concept
|
content_type: concept
|
||||||
weight: 50
|
weight: 50
|
||||||
---
|
---
|
||||||
|
|||||||
@@ -1,7 +1,7 @@
|
|||||||
---
|
---
|
||||||
reviewers:
|
reviewers:
|
||||||
- sig-cluster-lifecycle
|
- sig-cluster-lifecycle
|
||||||
title: Creating Highly Available clusters with kubeadm
|
title: Creating Highly Available Clusters with kubeadm
|
||||||
content_type: task
|
content_type: task
|
||||||
weight: 60
|
weight: 60
|
||||||
---
|
---
|
||||||
@@ -17,12 +17,12 @@ and control plane nodes are co-located.
|
|||||||
control plane nodes and etcd members are separated.
|
control plane nodes and etcd members are separated.
|
||||||
|
|
||||||
Before proceeding, you should carefully consider which approach best meets the needs of your applications
|
Before proceeding, you should carefully consider which approach best meets the needs of your applications
|
||||||
and environment. [This comparison topic](/docs/setup/production-environment/tools/kubeadm/ha-topology/) outlines the advantages and disadvantages of each.
|
and environment. [Options for Highly Available topology](/docs/setup/production-environment/tools/kubeadm/ha-topology/) outlines the advantages and disadvantages of each.
|
||||||
|
|
||||||
If you encounter issues with setting up the HA cluster, please provide us with feedback
|
If you encounter issues with setting up the HA cluster, please report these
|
||||||
in the kubeadm [issue tracker](https://github.com/kubernetes/kubeadm/issues/new).
|
in the kubeadm [issue tracker](https://github.com/kubernetes/kubeadm/issues/new).
|
||||||
|
|
||||||
See also [The upgrade documentation](/docs/tasks/administer-cluster/kubeadm/kubeadm-upgrade/).
|
See also the [upgrade documentation](/docs/tasks/administer-cluster/kubeadm/kubeadm-upgrade/).
|
||||||
|
|
||||||
{{< caution >}}
|
{{< caution >}}
|
||||||
This page does not address running your cluster on a cloud provider. In a cloud
|
This page does not address running your cluster on a cloud provider. In a cloud
|
||||||
@@ -32,22 +32,80 @@ LoadBalancer, or with dynamic PersistentVolumes.
|
|||||||
|
|
||||||
## {{% heading "prerequisites" %}}
|
## {{% heading "prerequisites" %}}
|
||||||
|
|
||||||
|
The prerequisites depend on which topology you have selected for your cluster's
|
||||||
|
control plane:
|
||||||
|
|
||||||
For both methods you need this infrastructure:
|
{{< tabs name="prerequisite_tabs" >}}
|
||||||
|
{{% tab name="Stacked etcd" %}}
|
||||||
|
<!--
|
||||||
|
note to reviewers: these prerequisites should match the start of the
|
||||||
|
external etc tab
|
||||||
|
-->
|
||||||
|
|
||||||
- Three machines that meet [kubeadm's minimum requirements](/docs/setup/production-environment/tools/kubeadm/install-kubeadm/#before-you-begin) for
|
You need:
|
||||||
the control-plane nodes
|
|
||||||
- Three machines that meet [kubeadm's minimum
|
- Three or more machines that meet [kubeadm's minimum requirements](/docs/setup/production-environment/tools/kubeadm/install-kubeadm/#before-you-begin) for
|
||||||
|
the control-plane nodes. Having an odd number of control plane nodes can help
|
||||||
|
with leader selection in the case of machine or zone failure.
|
||||||
|
- including a {{< glossary_tooltip text="container runtime" term_id="container-runtime" >}}, already set up and working
|
||||||
|
- Three or more machines that meet [kubeadm's minimum
|
||||||
requirements](/docs/setup/production-environment/tools/kubeadm/install-kubeadm/#before-you-begin) for the workers
|
requirements](/docs/setup/production-environment/tools/kubeadm/install-kubeadm/#before-you-begin) for the workers
|
||||||
|
- including a container runtime, already set up and working
|
||||||
- Full network connectivity between all machines in the cluster (public or
|
- Full network connectivity between all machines in the cluster (public or
|
||||||
private network)
|
private network)
|
||||||
- sudo privileges on all machines
|
- Superuser privileges on all machines using `sudo`
|
||||||
|
- You can use a different tool; this guide uses `sudo` in the examples.
|
||||||
- SSH access from one device to all nodes in the system
|
- SSH access from one device to all nodes in the system
|
||||||
- `kubeadm` and `kubelet` installed on all machines. `kubectl` is optional.
|
- `kubeadm` and `kubelet` already installed on all machines.
|
||||||
|
|
||||||
For the external etcd cluster only, you also need:
|
_See [Stacked etcd topology](/docs/setup/production-environment/tools/kubeadm/ha-topology/#stacked-etcd-topology) for context._
|
||||||
|
|
||||||
- Three additional machines for etcd members
|
{{% /tab %}}
|
||||||
|
{{% tab name="External etcd" %}}
|
||||||
|
<!--
|
||||||
|
note to reviewers: these prerequisites should match the start of the
|
||||||
|
stacked etc tab
|
||||||
|
-->
|
||||||
|
You need:
|
||||||
|
|
||||||
|
- Three or more machines that meet [kubeadm's minimum requirements](/docs/setup/production-environment/tools/kubeadm/install-kubeadm/#before-you-begin) for
|
||||||
|
the control-plane nodes. Having an odd number of control plane nodes can help
|
||||||
|
with leader selection in the case of machine or zone failure.
|
||||||
|
- including a {{< glossary_tooltip text="container runtime" term_id="container-runtime" >}}, already set up and working
|
||||||
|
- Three or more machines that meet [kubeadm's minimum
|
||||||
|
requirements](/docs/setup/production-environment/tools/kubeadm/install-kubeadm/#before-you-begin) for the workers
|
||||||
|
- including a container runtime, already set up and working
|
||||||
|
- Full network connectivity between all machines in the cluster (public or
|
||||||
|
private network)
|
||||||
|
- Superuser privileges on all machines using `sudo`
|
||||||
|
- You can use a different tool; this guide uses `sudo` in the examples.
|
||||||
|
- SSH access from one device to all nodes in the system
|
||||||
|
- `kubeadm` and `kubelet` already installed on all machines.
|
||||||
|
|
||||||
|
<!-- end of shared prerequisites -->
|
||||||
|
|
||||||
|
And you also need:
|
||||||
|
- Three or more additional machines, that will become etcd cluster members.
|
||||||
|
Having an odd number of members in the etcd cluster is a requirement for achieving
|
||||||
|
optimal voting quorum.
|
||||||
|
- These machines again need to have `kubeadm` and `kubelet` installed.
|
||||||
|
- These machines also require a container runtime, that is already set up and working.
|
||||||
|
|
||||||
|
_See [External etcd topology](/docs/setup/production-environment/tools/kubeadm/ha-topology/#external-etcd-topology) for context._
|
||||||
|
{{% /tab %}}
|
||||||
|
{{< /tabs >}}
|
||||||
|
|
||||||
|
### Container images
|
||||||
|
|
||||||
|
Each host should have access read and fetch images from the Kubernetes container image registry, `k8s.gcr.io`.
|
||||||
|
If you want to deploy a highly-available cluster where the hosts do not have access to pull images, this is possible. You must ensure by some other means that the correct container images are already available on the relevant hosts.
|
||||||
|
|
||||||
|
### Command line interface {#kubectl}
|
||||||
|
|
||||||
|
To manage Kubernetes once your cluster is set up, you should
|
||||||
|
[install kubectl](/docs/tasks/tools/#kubectl) on your PC. It is also useful
|
||||||
|
to install the `kubectl` tool on each control plane node, as this can be
|
||||||
|
helpful for troubleshooting.
|
||||||
|
|
||||||
<!-- steps -->
|
<!-- steps -->
|
||||||
|
|
||||||
@@ -80,14 +138,14 @@ option. Your cluster requirements may need a different configuration.
|
|||||||
- Read the [Options for Software Load Balancing](https://git.k8s.io/kubeadm/docs/ha-considerations.md#options-for-software-load-balancing)
|
- Read the [Options for Software Load Balancing](https://git.k8s.io/kubeadm/docs/ha-considerations.md#options-for-software-load-balancing)
|
||||||
guide for more details.
|
guide for more details.
|
||||||
|
|
||||||
1. Add the first control plane nodes to the load balancer and test the
|
1. Add the first control plane node to the load balancer, and test the
|
||||||
connection:
|
connection:
|
||||||
|
|
||||||
```sh
|
```shell
|
||||||
nc -v LOAD_BALANCER_IP PORT
|
nc -v <LOAD_BALANCER_IP> <PORT>
|
||||||
```
|
```
|
||||||
|
|
||||||
- A connection refused error is expected because the apiserver is not yet
|
A connection refused error is expected because the API server is not yet
|
||||||
running. A timeout, however, means the load balancer cannot communicate
|
running. A timeout, however, means the load balancer cannot communicate
|
||||||
with the control plane node. If a timeout occurs, reconfigure the load
|
with the control plane node. If a timeout occurs, reconfigure the load
|
||||||
balancer to communicate with the control plane node.
|
balancer to communicate with the control plane node.
|
||||||
@@ -127,7 +185,7 @@ option. Your cluster requirements may need a different configuration.
|
|||||||
set the `podSubnet` field under the `networking` object of `ClusterConfiguration`.
|
set the `podSubnet` field under the `networking` object of `ClusterConfiguration`.
|
||||||
{{< /note >}}
|
{{< /note >}}
|
||||||
|
|
||||||
- The output looks similar to:
|
The output looks similar to:
|
||||||
|
|
||||||
```sh
|
```sh
|
||||||
...
|
...
|
||||||
@@ -141,10 +199,12 @@ option. Your cluster requirements may need a different configuration.
|
|||||||
kubeadm join 192.168.0.200:6443 --token 9vr73a.a8uxyaju799qwdjv --discovery-token-ca-cert-hash sha256:7c2e69131a36ae2a042a339b33381c6d0d43887e2de83720eff5359e26aec866
|
kubeadm join 192.168.0.200:6443 --token 9vr73a.a8uxyaju799qwdjv --discovery-token-ca-cert-hash sha256:7c2e69131a36ae2a042a339b33381c6d0d43887e2de83720eff5359e26aec866
|
||||||
```
|
```
|
||||||
|
|
||||||
- Copy this output to a text file. You will need it later to join control plane and worker nodes to the cluster.
|
- Copy this output to a text file. You will need it later to join control plane and worker nodes to
|
||||||
|
the cluster.
|
||||||
- When `--upload-certs` is used with `kubeadm init`, the certificates of the primary control plane
|
- When `--upload-certs` is used with `kubeadm init`, the certificates of the primary control plane
|
||||||
are encrypted and uploaded in the `kubeadm-certs` Secret.
|
are encrypted and uploaded in the `kubeadm-certs` Secret.
|
||||||
- To re-upload the certificates and generate a new decryption key, use the following command on a control plane
|
- To re-upload the certificates and generate a new decryption key, use the following command on a
|
||||||
|
control plane
|
||||||
node that is already joined to the cluster:
|
node that is already joined to the cluster:
|
||||||
|
|
||||||
```sh
|
```sh
|
||||||
@@ -168,7 +228,8 @@ option. Your cluster requirements may need a different configuration.
|
|||||||
|
|
||||||
1. Apply the CNI plugin of your choice:
|
1. Apply the CNI plugin of your choice:
|
||||||
[Follow these instructions](/docs/setup/production-environment/tools/kubeadm/create-cluster-kubeadm/#pod-network)
|
[Follow these instructions](/docs/setup/production-environment/tools/kubeadm/create-cluster-kubeadm/#pod-network)
|
||||||
to install the CNI provider. Make sure the configuration corresponds to the Pod CIDR specified in the kubeadm configuration file if applicable.
|
to install the CNI provider. Make sure the configuration corresponds to the Pod CIDR specified in the
|
||||||
|
kubeadm configuration file (if applicable).
|
||||||
|
|
||||||
{{< note >}}
|
{{< note >}}
|
||||||
You must pick a network plugin that suits your use case and deploy it before you move on to next step.
|
You must pick a network plugin that suits your use case and deploy it before you move on to next step.
|
||||||
@@ -183,12 +244,6 @@ option. Your cluster requirements may need a different configuration.
|
|||||||
|
|
||||||
### Steps for the rest of the control plane nodes
|
### Steps for the rest of the control plane nodes
|
||||||
|
|
||||||
{{< note >}}
|
|
||||||
Since kubeadm version 1.15 you can join multiple control-plane nodes in parallel.
|
|
||||||
Prior to this version, you must join new control plane nodes sequentially, only after
|
|
||||||
the first node has finished initializing.
|
|
||||||
{{< /note >}}
|
|
||||||
|
|
||||||
For each additional control plane node you should:
|
For each additional control plane node you should:
|
||||||
|
|
||||||
1. Execute the join command that was previously given to you by the `kubeadm init` output on the first node.
|
1. Execute the join command that was previously given to you by the `kubeadm init` output on the first node.
|
||||||
@@ -202,6 +257,8 @@ For each additional control plane node you should:
|
|||||||
- The `--certificate-key ...` will cause the control plane certificates to be downloaded
|
- The `--certificate-key ...` will cause the control plane certificates to be downloaded
|
||||||
from the `kubeadm-certs` Secret in the cluster and be decrypted using the given key.
|
from the `kubeadm-certs` Secret in the cluster and be decrypted using the given key.
|
||||||
|
|
||||||
|
You can join multiple control-plane nodes in parallel.
|
||||||
|
|
||||||
## External etcd nodes
|
## External etcd nodes
|
||||||
|
|
||||||
Setting up a cluster with external etcd nodes is similar to the procedure used for stacked etcd
|
Setting up a cluster with external etcd nodes is similar to the procedure used for stacked etcd
|
||||||
@@ -210,7 +267,7 @@ in the kubeadm config file.
|
|||||||
|
|
||||||
### Set up the etcd cluster
|
### Set up the etcd cluster
|
||||||
|
|
||||||
1. Follow [these instructions](/docs/setup/production-environment/tools/kubeadm/setup-ha-etcd-with-kubeadm/) to set up the etcd cluster.
|
1. Follow these [instructions](/docs/setup/production-environment/tools/kubeadm/setup-ha-etcd-with-kubeadm/) to set up the etcd cluster.
|
||||||
|
|
||||||
1. Setup SSH as described [here](#manual-certs).
|
1. Setup SSH as described [here](#manual-certs).
|
||||||
|
|
||||||
@@ -229,24 +286,28 @@ in the kubeadm config file.
|
|||||||
|
|
||||||
1. Create a file called `kubeadm-config.yaml` with the following contents:
|
1. Create a file called `kubeadm-config.yaml` with the following contents:
|
||||||
|
|
||||||
|
|
||||||
|
```yaml
|
||||||
|
---
|
||||||
apiVersion: kubeadm.k8s.io/v1beta3
|
apiVersion: kubeadm.k8s.io/v1beta3
|
||||||
kind: ClusterConfiguration
|
kind: ClusterConfiguration
|
||||||
kubernetesVersion: stable
|
kubernetesVersion: stable
|
||||||
controlPlaneEndpoint: "LOAD_BALANCER_DNS:LOAD_BALANCER_PORT"
|
controlPlaneEndpoint: "LOAD_BALANCER_DNS:LOAD_BALANCER_PORT" # change this (see below)
|
||||||
etcd:
|
etcd:
|
||||||
external:
|
external:
|
||||||
endpoints:
|
endpoints:
|
||||||
- https://ETCD_0_IP:2379
|
- https://ETCD_0_IP:2379 # change ETCD_0_IP appropriately
|
||||||
- https://ETCD_1_IP:2379
|
- https://ETCD_1_IP:2379 # change ETCD_1_IP appropriately
|
||||||
- https://ETCD_2_IP:2379
|
- https://ETCD_2_IP:2379 # change ETCD_2_IP appropriately
|
||||||
caFile: /etc/kubernetes/pki/etcd/ca.crt
|
caFile: /etc/kubernetes/pki/etcd/ca.crt
|
||||||
certFile: /etc/kubernetes/pki/apiserver-etcd-client.crt
|
certFile: /etc/kubernetes/pki/apiserver-etcd-client.crt
|
||||||
keyFile: /etc/kubernetes/pki/apiserver-etcd-client.key
|
keyFile: /etc/kubernetes/pki/apiserver-etcd-client.key
|
||||||
|
```
|
||||||
|
|
||||||
{{< note >}}
|
{{< note >}}
|
||||||
The difference between stacked etcd and external etcd here is that the external etcd setup requires
|
The difference between stacked etcd and external etcd here is that the external etcd setup requires
|
||||||
a configuration file with the etcd endpoints under the `external` object for `etcd`.
|
a configuration file with the etcd endpoints under the `external` object for `etcd`.
|
||||||
In the case of the stacked etcd topology this is managed automatically.
|
In the case of the stacked etcd topology, this is managed automatically.
|
||||||
{{< /note >}}
|
{{< /note >}}
|
||||||
|
|
||||||
- Replace the following variables in the config template with the appropriate values for your cluster:
|
- Replace the following variables in the config template with the appropriate values for your cluster:
|
||||||
@@ -263,11 +324,12 @@ The following steps are similar to the stacked etcd setup:
|
|||||||
|
|
||||||
1. Write the output join commands that are returned to a text file for later use.
|
1. Write the output join commands that are returned to a text file for later use.
|
||||||
|
|
||||||
1. Apply the CNI plugin of your choice. The given example is for Weave Net:
|
1. Apply the CNI plugin of your choice.
|
||||||
|
|
||||||
```sh
|
{{< note >}}
|
||||||
kubectl apply -f "https://cloud.weave.works/k8s/net?k8s-version=$(kubectl version | base64 | tr -d '\n')"
|
You must pick a network plugin that suits your use case and deploy it before you move on to next step.
|
||||||
```
|
If you don't do this, you will not be able to launch your cluster properly.
|
||||||
|
{{< /note >}}
|
||||||
|
|
||||||
### Steps for the rest of the control plane nodes
|
### Steps for the rest of the control plane nodes
|
||||||
|
|
||||||
@@ -295,7 +357,7 @@ If you choose to not use `kubeadm init` with the `--upload-certs` flag this mean
|
|||||||
you are going to have to manually copy the certificates from the primary control plane node to the
|
you are going to have to manually copy the certificates from the primary control plane node to the
|
||||||
joining control plane nodes.
|
joining control plane nodes.
|
||||||
|
|
||||||
There are many ways to do this. In the following example we are using `ssh` and `scp`:
|
There are many ways to do this. The following example uses `ssh` and `scp`:
|
||||||
|
|
||||||
SSH is required if you want to control all nodes from a single machine.
|
SSH is required if you want to control all nodes from a single machine.
|
||||||
|
|
||||||
@@ -314,7 +376,9 @@ SSH is required if you want to control all nodes from a single machine.
|
|||||||
|
|
||||||
1. SSH between nodes to check that the connection is working correctly.
|
1. SSH between nodes to check that the connection is working correctly.
|
||||||
|
|
||||||
- When you SSH to any node, make sure to add the `-A` flag:
|
- When you SSH to any node, add the `-A` flag. This flag allows the node that you
|
||||||
|
have logged into via SSH to access the SSH agent on your PC. Consider alternative
|
||||||
|
methods if you do not fully trust the security of your user session on the node.
|
||||||
|
|
||||||
```
|
```
|
||||||
ssh -A 10.0.0.7
|
ssh -A 10.0.0.7
|
||||||
@@ -327,9 +391,9 @@ SSH is required if you want to control all nodes from a single machine.
|
|||||||
sudo -E -s
|
sudo -E -s
|
||||||
```
|
```
|
||||||
|
|
||||||
1. After configuring SSH on all the nodes you should run the following script on the first control plane node after
|
1. After configuring SSH on all the nodes you should run the following script on the first
|
||||||
running `kubeadm init`. This script will copy the certificates from the first control plane node to the other
|
control plane node after running `kubeadm init`. This script will copy the certificates from
|
||||||
control plane nodes:
|
the first control plane node to the other control plane nodes:
|
||||||
|
|
||||||
In the following example, replace `CONTROL_PLANE_IPS` with the IP addresses of the
|
In the following example, replace `CONTROL_PLANE_IPS` with the IP addresses of the
|
||||||
other control plane nodes.
|
other control plane nodes.
|
||||||
@@ -344,7 +408,7 @@ SSH is required if you want to control all nodes from a single machine.
|
|||||||
scp /etc/kubernetes/pki/front-proxy-ca.crt "${USER}"@$host:
|
scp /etc/kubernetes/pki/front-proxy-ca.crt "${USER}"@$host:
|
||||||
scp /etc/kubernetes/pki/front-proxy-ca.key "${USER}"@$host:
|
scp /etc/kubernetes/pki/front-proxy-ca.key "${USER}"@$host:
|
||||||
scp /etc/kubernetes/pki/etcd/ca.crt "${USER}"@$host:etcd-ca.crt
|
scp /etc/kubernetes/pki/etcd/ca.crt "${USER}"@$host:etcd-ca.crt
|
||||||
# Quote this line if you are using external etcd
|
# Skip the next line if you are using external etcd
|
||||||
scp /etc/kubernetes/pki/etcd/ca.key "${USER}"@$host:etcd-ca.key
|
scp /etc/kubernetes/pki/etcd/ca.key "${USER}"@$host:etcd-ca.key
|
||||||
done
|
done
|
||||||
```
|
```
|
||||||
@@ -368,6 +432,6 @@ SSH is required if you want to control all nodes from a single machine.
|
|||||||
mv /home/${USER}/front-proxy-ca.crt /etc/kubernetes/pki/
|
mv /home/${USER}/front-proxy-ca.crt /etc/kubernetes/pki/
|
||||||
mv /home/${USER}/front-proxy-ca.key /etc/kubernetes/pki/
|
mv /home/${USER}/front-proxy-ca.key /etc/kubernetes/pki/
|
||||||
mv /home/${USER}/etcd-ca.crt /etc/kubernetes/pki/etcd/ca.crt
|
mv /home/${USER}/etcd-ca.crt /etc/kubernetes/pki/etcd/ca.crt
|
||||||
# Quote this line if you are using external etcd
|
# Skip the next line if you are using external etcd
|
||||||
mv /home/${USER}/etcd-ca.key /etc/kubernetes/pki/etcd/ca.key
|
mv /home/${USER}/etcd-ca.key /etc/kubernetes/pki/etcd/ca.key
|
||||||
```
|
```
|
||||||
|
|||||||
+3
-2
@@ -1,7 +1,7 @@
|
|||||||
---
|
---
|
||||||
reviewers:
|
reviewers:
|
||||||
- sig-cluster-lifecycle
|
- sig-cluster-lifecycle
|
||||||
title: Set up a High Availability etcd cluster with kubeadm
|
title: Set up a High Availability etcd Cluster with kubeadm
|
||||||
content_type: task
|
content_type: task
|
||||||
weight: 70
|
weight: 70
|
||||||
---
|
---
|
||||||
@@ -19,7 +19,8 @@ aspects.
|
|||||||
By default, kubeadm runs a local etcd instance on each control plane node.
|
By default, kubeadm runs a local etcd instance on each control plane node.
|
||||||
It is also possible to treat the etcd cluster as external and provision
|
It is also possible to treat the etcd cluster as external and provision
|
||||||
etcd instances on separate hosts. The differences between the two approaches are covered in the
|
etcd instances on separate hosts. The differences between the two approaches are covered in the
|
||||||
[Options for Highly Available topology][/docs/setup/production-environment/tools/kubeadm/ha-topology] page.
|
[Options for Highly Available topology](/docs/setup/production-environment/tools/kubeadm/ha-topology) page.
|
||||||
|
|
||||||
This task walks through the process of creating a high availability external
|
This task walks through the process of creating a high availability external
|
||||||
etcd cluster of three members that can be used by kubeadm during cluster creation.
|
etcd cluster of three members that can be used by kubeadm during cluster creation.
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user