Move a batch of cluster admin topics. (#2813)
This commit is contained in:
@@ -3,6 +3,10 @@ abstract: "Detailed explanations of Kubernetes system concepts and abstractions.
|
|||||||
toc:
|
toc:
|
||||||
- docs/concepts/index.md
|
- docs/concepts/index.md
|
||||||
|
|
||||||
|
- title: Overview
|
||||||
|
section:
|
||||||
|
- docs/concepts/overview/components.md
|
||||||
|
|
||||||
- title: Kubernetes Objects
|
- title: Kubernetes Objects
|
||||||
section:
|
section:
|
||||||
- docs/concepts/abstractions/overview.md
|
- docs/concepts/abstractions/overview.md
|
||||||
@@ -29,7 +33,10 @@ toc:
|
|||||||
- title: Cluster Administration
|
- title: Cluster Administration
|
||||||
section:
|
section:
|
||||||
- docs/concepts/cluster-administration/logging.md
|
- docs/concepts/cluster-administration/logging.md
|
||||||
|
- docs/concepts/cluster-administration/multiple-clusters.md
|
||||||
- docs/concepts/cluster-administration/federation.md
|
- docs/concepts/cluster-administration/federation.md
|
||||||
|
- docs/concepts/cluster-administration/guaranteed-scheduling-critical-addon-pods.md
|
||||||
|
- docs/concepts/cluster-administration/sysctl-cluster.md
|
||||||
|
|
||||||
- title: Configuration
|
- title: Configuration
|
||||||
section:
|
section:
|
||||||
|
|||||||
@@ -4,133 +4,6 @@ assignees:
|
|||||||
title: Kubernetes Components
|
title: Kubernetes Components
|
||||||
---
|
---
|
||||||
|
|
||||||
This document outlines the various binary components that need to run to
|
{% include user-guide-content-moved.md %}
|
||||||
deliver a functioning Kubernetes cluster.
|
|
||||||
|
|
||||||
## Master Components
|
[Kubernetes Components](/docs/concepts/overview/components/)
|
||||||
|
|
||||||
Master components are those that provide the cluster's control plane. For
|
|
||||||
example, master components are responsible for making global decisions about the
|
|
||||||
cluster (e.g., scheduling), and detecting and responding to cluster events
|
|
||||||
(e.g., starting up a new pod when a replication controller's 'replicas' field is
|
|
||||||
unsatisfied).
|
|
||||||
|
|
||||||
In theory, Master components can be run on any node in the cluster. However,
|
|
||||||
for simplicity, current set up scripts typically start all master components on
|
|
||||||
the same VM, and does not run user containers on this VM. See
|
|
||||||
[high-availability.md](/docs/admin/high-availability) for an example multi-master-VM setup.
|
|
||||||
|
|
||||||
Even in the future, when Kubernetes is fully self-hosting, it will probably be
|
|
||||||
wise to only allow master components to schedule on a subset of nodes, to limit
|
|
||||||
co-running with user-run pods, reducing the possible scope of a
|
|
||||||
node-compromising security exploit.
|
|
||||||
|
|
||||||
### kube-apiserver
|
|
||||||
|
|
||||||
[kube-apiserver](/docs/admin/kube-apiserver) exposes the Kubernetes API; it is the front-end for the
|
|
||||||
Kubernetes control plane. It is designed to scale horizontally (i.e., one scales
|
|
||||||
it by running more of them-- [high-availability.md](/docs/admin/high-availability)).
|
|
||||||
|
|
||||||
### etcd
|
|
||||||
|
|
||||||
[etcd](/docs/admin/etcd) is used as Kubernetes' backing store. All cluster data is stored here.
|
|
||||||
Proper administration of a Kubernetes cluster includes a backup plan for etcd's
|
|
||||||
data.
|
|
||||||
|
|
||||||
### kube-controller-manager
|
|
||||||
|
|
||||||
[kube-controller-manager](/docs/admin/kube-controller-manager) is a binary that runs controllers, which are the
|
|
||||||
background threads that handle routine tasks in the cluster. Logically, each
|
|
||||||
controller is a separate process, but to reduce the number of moving pieces in
|
|
||||||
the system, they are all compiled into a single binary and run in a single
|
|
||||||
process.
|
|
||||||
|
|
||||||
These controllers include:
|
|
||||||
|
|
||||||
* Node Controller: Responsible for noticing & responding when nodes go down.
|
|
||||||
* Replication Controller: Responsible for maintaining the correct number of pods for every replication
|
|
||||||
controller object in the system.
|
|
||||||
* Endpoints Controller: Populates the Endpoints object (i.e., join Services & Pods).
|
|
||||||
* Service Account & Token Controllers: Create default accounts and API access tokens for new namespaces.
|
|
||||||
* ... and others.
|
|
||||||
|
|
||||||
### kube-scheduler
|
|
||||||
|
|
||||||
[kube-scheduler](/docs/admin/kube-scheduler) watches newly created pods that have no node assigned, and
|
|
||||||
selects a node for them to run on.
|
|
||||||
|
|
||||||
### addons
|
|
||||||
|
|
||||||
Addons are pods and services that implement cluster features. The pods may be managed
|
|
||||||
by Deployments, ReplicationContollers, etc. Namespaced addon objects are created in
|
|
||||||
the "kube-system" namespace.
|
|
||||||
|
|
||||||
Addon manager takes the responsibility for creating and maintaining addon resources.
|
|
||||||
See [here](http://releases.k8s.io/HEAD/cluster/addons) for more details.
|
|
||||||
|
|
||||||
#### DNS
|
|
||||||
|
|
||||||
While the other addons are not strictly required, all Kubernetes
|
|
||||||
clusters should have [cluster DNS](/docs/admin/dns/), as many examples rely on it.
|
|
||||||
|
|
||||||
Cluster DNS is a DNS server, in addition to the other DNS server(s) in your
|
|
||||||
environment, which serves DNS records for Kubernetes services.
|
|
||||||
|
|
||||||
Containers started by Kubernetes automatically include this DNS server
|
|
||||||
in their DNS searches.
|
|
||||||
|
|
||||||
#### User interface
|
|
||||||
|
|
||||||
The kube-ui provides a read-only overview of the cluster state. Access
|
|
||||||
[the UI using kubectl proxy](/docs/user-guide/connecting-to-applications-proxy/#connecting-to-the-kube-ui-service-from-your-local-workstation)
|
|
||||||
|
|
||||||
#### Container Resource Monitoring
|
|
||||||
|
|
||||||
[Container Resource Monitoring](/docs/user-guide/monitoring) records generic time-series metrics
|
|
||||||
about containers in a central database, and provides a UI for browsing that data.
|
|
||||||
|
|
||||||
#### Cluster-level Logging
|
|
||||||
|
|
||||||
A [Cluster-level logging](/docs/user-guide/logging/overview) mechanism is responsible for
|
|
||||||
saving container logs to a central log store with search/browsing interface.
|
|
||||||
|
|
||||||
## Node components
|
|
||||||
|
|
||||||
Node components run on every node, maintaining running pods and providing them
|
|
||||||
the Kubernetes runtime environment.
|
|
||||||
|
|
||||||
### kubelet
|
|
||||||
|
|
||||||
[kubelet](/docs/admin/kubelet) is the primary node agent. It:
|
|
||||||
|
|
||||||
* Watches for pods that have been assigned to its node (either by apiserver
|
|
||||||
or via local configuration file) and:
|
|
||||||
* Mounts the pod's required volumes
|
|
||||||
* Downloads the pod's secrets
|
|
||||||
* Runs the pod's containers via docker (or, experimentally, rkt).
|
|
||||||
* Periodically executes any requested container liveness probes.
|
|
||||||
* Reports the status of the pod back to the rest of the system, by creating a
|
|
||||||
"mirror pod" if necessary.
|
|
||||||
* Reports the status of the node back to the rest of the system.
|
|
||||||
|
|
||||||
### kube-proxy
|
|
||||||
|
|
||||||
[kube-proxy](/docs/admin/kube-proxy) enables the Kubernetes service abstraction by maintaining
|
|
||||||
network rules on the host and performing connection forwarding.
|
|
||||||
|
|
||||||
### docker
|
|
||||||
|
|
||||||
`docker` is of course used for actually running containers.
|
|
||||||
|
|
||||||
### rkt
|
|
||||||
|
|
||||||
`rkt` is supported experimentally as an alternative to docker.
|
|
||||||
|
|
||||||
### supervisord
|
|
||||||
|
|
||||||
`supervisord` is a lightweight process babysitting system for keeping kubelet and docker
|
|
||||||
running.
|
|
||||||
|
|
||||||
### fluentd
|
|
||||||
|
|
||||||
`fluentd` is a daemon which helps provide [cluster-level logging](#cluster-level-logging).
|
|
||||||
|
|||||||
@@ -4,63 +4,6 @@ assignees:
|
|||||||
title: Using Multiple Clusters
|
title: Using Multiple Clusters
|
||||||
---
|
---
|
||||||
|
|
||||||
You may want to set up multiple Kubernetes clusters, both to
|
{% include user-guide-content-moved.md %}
|
||||||
have clusters in different regions to be nearer to your users, and to tolerate failures and/or invasive maintenance.
|
|
||||||
This document describes some of the issues to consider when making a decision about doing so.
|
|
||||||
|
|
||||||
If you decide to have multiple clusters, Kubernetes provides a way to [federate them](/docs/admin/federation/).
|
[Using Multiple Clusters](/docs/concepts/cluster-administration/multiple-clusters/)
|
||||||
|
|
||||||
## Scope of a single cluster
|
|
||||||
|
|
||||||
On IaaS providers such as Google Compute Engine or Amazon Web Services, a VM exists in a
|
|
||||||
[zone](https://cloud.google.com/compute/docs/zones) or [availability
|
|
||||||
zone](http://docs.aws.amazon.com/AWSEC2/latest/UserGuide/using-regions-availability-zones.html).
|
|
||||||
We suggest that all the VMs in a Kubernetes cluster should be in the same availability zone, because:
|
|
||||||
|
|
||||||
- compared to having a single global Kubernetes cluster, there are fewer single-points of failure
|
|
||||||
- compared to a cluster that spans availability zones, it is easier to reason about the availability properties of a
|
|
||||||
single-zone cluster.
|
|
||||||
- when the Kubernetes developers are designing the system (e.g. making assumptions about latency, bandwidth, or
|
|
||||||
correlated failures) they are assuming all the machines are in a single data center, or otherwise closely connected.
|
|
||||||
|
|
||||||
It is okay to have multiple clusters per availability zone, though on balance we think fewer is better.
|
|
||||||
Reasons to prefer fewer clusters are:
|
|
||||||
|
|
||||||
- improved bin packing of Pods in some cases with more nodes in one cluster (less resource fragmentation)
|
|
||||||
- reduced operational overhead (though the advantage is diminished as ops tooling and processes matures)
|
|
||||||
- reduced costs for per-cluster fixed resource costs, e.g. apiserver VMs (but small as a percentage
|
|
||||||
of overall cluster cost for medium to large clusters).
|
|
||||||
|
|
||||||
Reasons to have multiple clusters include:
|
|
||||||
|
|
||||||
- strict security policies requiring isolation of one class of work from another (but, see Partitioning Clusters
|
|
||||||
below).
|
|
||||||
- test clusters to canary new Kubernetes releases or other cluster software.
|
|
||||||
|
|
||||||
## Selecting the right number of clusters
|
|
||||||
|
|
||||||
The selection of the number of Kubernetes clusters may be a relatively static choice, only revisited occasionally.
|
|
||||||
By contrast, the number of nodes in a cluster and the number of pods in a service may change frequently according to
|
|
||||||
load and growth.
|
|
||||||
|
|
||||||
To pick the number of clusters, first, decide which regions you need to be in to have adequate latency to all your end users, for services that will run
|
|
||||||
on Kubernetes (if you use a Content Distribution Network, the latency requirements for the CDN-hosted content need not
|
|
||||||
be considered). Legal issues might influence this as well. For example, a company with a global customer base might decide to have clusters in US, EU, AP, and SA regions.
|
|
||||||
Call the number of regions to be in `R`.
|
|
||||||
|
|
||||||
Second, decide how many clusters should be able to be unavailable at the same time, while still being available. Call
|
|
||||||
the number that can be unavailable `U`. If you are not sure, then 1 is a fine choice.
|
|
||||||
|
|
||||||
If it is allowable for load-balancing to direct traffic to any region in the event of a cluster failure, then
|
|
||||||
you need at least the larger of `R` or `U + 1` clusters. If it is not (e.g. you want to ensure low latency for all
|
|
||||||
users in the event of a cluster failure), then you need to have `R * (U + 1)` clusters
|
|
||||||
(`U + 1` in each of `R` regions). In any case, try to put each cluster in a different zone.
|
|
||||||
|
|
||||||
Finally, if any of your clusters would need more than the maximum recommended number of nodes for a Kubernetes cluster, then
|
|
||||||
you may need even more clusters. Kubernetes v1.3 supports clusters up to 1000 nodes in size.
|
|
||||||
|
|
||||||
## Working with multiple clusters
|
|
||||||
|
|
||||||
When you have multiple clusters, you would typically create services with the same config in each cluster and put each of those
|
|
||||||
service instances behind a load balancer (AWS Elastic Load Balancer, GCE Forwarding Rule or HTTP Load Balancer) spanning all of them, so that
|
|
||||||
failures of a single cluster are not visible to end users.
|
|
||||||
|
|||||||
@@ -6,52 +6,6 @@ assignees:
|
|||||||
title: Guaranteed Scheduling For Critical Add-On Pods
|
title: Guaranteed Scheduling For Critical Add-On Pods
|
||||||
---
|
---
|
||||||
|
|
||||||
* TOC
|
{% include user-guide-content-moved.md %}
|
||||||
{:toc}
|
|
||||||
|
|
||||||
## Overview
|
[Guaranteed Scheduling for Critical Add-On Pods](/docs/concepts/cluster-administration/guaranteed-scheduling-critical-addon-pods/)
|
||||||
|
|
||||||
In addition to Kubernetes core components like api-server, scheduler, controller-manager running on a master machine
|
|
||||||
there are a number of add-ons which, for various reasons, must run on a regular cluster node (rather than the Kubernetes master).
|
|
||||||
Some of these add-ons are critical to a fully functional cluster, such as Heapster, DNS, and UI.
|
|
||||||
A cluster may stop working properly if a critical add-on is evicted (either manually or as a side effect of another operation like upgrade)
|
|
||||||
and becomes pending (for example when the cluster is highly utilized and either there are other pending pods that schedule into the space
|
|
||||||
vacated by the evicted critical add-on pod or the amount of resources available on the node changed for some other reason).
|
|
||||||
|
|
||||||
## Rescheduler: guaranteed scheduling of critical add-ons
|
|
||||||
|
|
||||||
Rescheduler ensures that critical add-ons are always scheduled
|
|
||||||
(assuming the cluster has enough resources to run the critical add-on pods in the absence of regular pods).
|
|
||||||
If the scheduler determines that no node has enough free resources to run the critical add-on pod
|
|
||||||
given the pods that are already running in the cluster
|
|
||||||
(indicated by critical add-on pod's pod condition PodScheduled set to false, the reason set to Unschedulable)
|
|
||||||
the rescheduler tries to free up space for the add-on by evicting some pods; then the scheduler will schedule the add-on pod.
|
|
||||||
|
|
||||||
To avoid situation when another pod is scheduled into the space prepared for the critical add-on,
|
|
||||||
the chosen node gets a temporary taint "CriticalAddonsOnly" before the eviction(s)
|
|
||||||
(see [more details](https://github.com/kubernetes/kubernetes/blob/master/docs/design/taint-toleration-dedicated.md)).
|
|
||||||
Each critical add-on has to tolerate it,
|
|
||||||
while the other pods shouldn't tolerate the taint. The taint is removed once the add-on is successfully scheduled.
|
|
||||||
|
|
||||||
*Warning:* currently there is no guarantee which node is chosen and which pods are being killed
|
|
||||||
in order to schedule critical pods, so if rescheduler is enabled your pods might be occasionally
|
|
||||||
killed for this purpose.
|
|
||||||
|
|
||||||
## Config
|
|
||||||
|
|
||||||
Rescheduler doesn't have any user facing configuration (component config) or API.
|
|
||||||
It's enabled by default. It can be disabled:
|
|
||||||
|
|
||||||
* during cluster setup by setting `ENABLE_RESCHEDULER` flag to `false`
|
|
||||||
* on running cluster by deleting its manifest from master node
|
|
||||||
(default path `/etc/kubernetes/manifests/rescheduler.manifest`)
|
|
||||||
|
|
||||||
### Marking add-on as critical
|
|
||||||
|
|
||||||
To be critical an add-on has to run in `kube-system` namespace (configurable via flag)
|
|
||||||
and have the following annotations specified:
|
|
||||||
|
|
||||||
* `scheduler.alpha.kubernetes.io/critical-pod` set to empty string
|
|
||||||
* `scheduler.alpha.kubernetes.io/tolerations` set to `[{"key":"CriticalAddonsOnly", "operator":"Exists"}]`
|
|
||||||
|
|
||||||
The first one marks a pod a critical. The second one is required by Rescheduler algorithm.
|
|
||||||
|
|||||||
+2
-115
@@ -4,119 +4,6 @@ assignees:
|
|||||||
title: Using Sysctls in a Kubernetes Cluster
|
title: Using Sysctls in a Kubernetes Cluster
|
||||||
---
|
---
|
||||||
|
|
||||||
* TOC
|
{% include user-guide-content-moved.md %}
|
||||||
{:toc}
|
|
||||||
|
|
||||||
This document describes how sysctls are used within a Kubernetes cluster.
|
[Using Sysctls in a Kubernetes Cluster](/docs/concepts/cluster-administration/sysctl-cluster/)
|
||||||
|
|
||||||
## What is a Sysctl?
|
|
||||||
|
|
||||||
In Linux, the sysctl interface allows an administrator to modify kernel
|
|
||||||
parameters at runtime. Parameters are available via the `/proc/sys/` virtual
|
|
||||||
process file system. The parameters cover various subsystems such as:
|
|
||||||
|
|
||||||
- kernel (common prefix: `kernel.`)
|
|
||||||
- networking (common prefix: `net.`)
|
|
||||||
- virtual memory (common prefix: `vm.`)
|
|
||||||
- MDADM (common prefix: `dev.`)
|
|
||||||
- More subsystems are described in [Kernel docs](https://www.kernel.org/doc/Documentation/sysctl/README).
|
|
||||||
|
|
||||||
To get a list of all parameters, you can run
|
|
||||||
|
|
||||||
```
|
|
||||||
$ sudo sysctl -a
|
|
||||||
```
|
|
||||||
|
|
||||||
## Namespaced vs. Node-Level Sysctls
|
|
||||||
|
|
||||||
A number of sysctls are _namespaced_ in today's Linux kernels. This means that
|
|
||||||
they can be set independently for each pod on a node. Being namespaced is a
|
|
||||||
requirement for sysctls to be accessible in a pod context within Kubernetes.
|
|
||||||
|
|
||||||
The following sysctls are known to be _namespaced_:
|
|
||||||
|
|
||||||
- `kernel.shm*`,
|
|
||||||
- `kernel.msg*`,
|
|
||||||
- `kernel.sem`,
|
|
||||||
- `fs.mqueue.*`,
|
|
||||||
- `net.*`.
|
|
||||||
|
|
||||||
Sysctls which are not namespaced are called _node-level_ and must be set
|
|
||||||
manually by the cluster admin, either by means of the underlying Linux
|
|
||||||
distribution of the nodes (e.g. via `/etc/sysctls.conf`) or using a DaemonSet
|
|
||||||
with privileged containers.
|
|
||||||
|
|
||||||
**Note**: it is good practice to consider nodes with special sysctl settings as
|
|
||||||
_tainted_ within a cluster, and only schedule pods onto them which need those
|
|
||||||
sysctl settings. It is suggested to use the Kubernetes [_taints and toleration_
|
|
||||||
feature](/docs/user-guide/kubectl/kubectl_taint.md) to implement this.
|
|
||||||
|
|
||||||
## Safe vs. Unsafe Sysctls
|
|
||||||
|
|
||||||
Sysctls are grouped into _safe_ and _unsafe_ sysctls. In addition to proper
|
|
||||||
namespacing a _safe_ sysctl must be properly _isolated_ between pods on the same
|
|
||||||
node. This means that setting a _safe_ sysctl for one pod
|
|
||||||
|
|
||||||
- must not have any influence on any other pod on the node
|
|
||||||
- must not allow to harm the node's health
|
|
||||||
- must not allow to gain CPU or memory resources outside of the resource limits
|
|
||||||
of a pod.
|
|
||||||
|
|
||||||
By far, most of the _namespaced_ sysctls are not necessarily considered _safe_.
|
|
||||||
|
|
||||||
For Kubernetes 1.4, the following sysctls are supported in the _safe_ set:
|
|
||||||
|
|
||||||
- `kernel.shm_rmid_forced`,
|
|
||||||
- `net.ipv4.ip_local_port_range`,
|
|
||||||
- `net.ipv4.tcp_syncookies`.
|
|
||||||
|
|
||||||
This list will be extended in future Kubernetes versions when the kubelet
|
|
||||||
supports better isolation mechanisms.
|
|
||||||
|
|
||||||
All _safe_ sysctls are enabled by default.
|
|
||||||
|
|
||||||
All _unsafe_ sysctls are disabled by default and must be allowed manually by the
|
|
||||||
cluster admin on a per-node basis. Pods with disabled unsafe sysctls will be
|
|
||||||
scheduled, but will fail to launch.
|
|
||||||
|
|
||||||
**Warning**: Due to their nature of being _unsafe_, the use of _unsafe_ sysctls
|
|
||||||
is at-your-own-risk and can lead to severe problems like wrong behavior of
|
|
||||||
containers, resource shortage or complete breakage of a node.
|
|
||||||
|
|
||||||
## Enabling Unsafe Sysctls
|
|
||||||
|
|
||||||
With the warning above in mind, the cluster admin can allow certain _unsafe_
|
|
||||||
sysctls for very special situations like e.g. high-performance or real-time
|
|
||||||
application tuning. _Unsafe_ sysctls are enabled on a node-by-node basis with a
|
|
||||||
flag of the kubelet, e.g.:
|
|
||||||
|
|
||||||
```shell
|
|
||||||
$ kubelet --experimental-allowed-unsafe-sysctls 'kernel.msg*,net.ipv4.route.min_pmtu' ...
|
|
||||||
```
|
|
||||||
|
|
||||||
Only _namespaced_ sysctls can be enabled this way.
|
|
||||||
|
|
||||||
## Setting Sysctls for a Pod
|
|
||||||
|
|
||||||
The sysctl feature is an alpha API in Kubernetes 1.4. Therefore, sysctls are set
|
|
||||||
using annotations on pods. They apply to all containers in the same pod.
|
|
||||||
|
|
||||||
Here is an example, with different annotations for _safe_ and _unsafe_ sysctls:
|
|
||||||
|
|
||||||
```yaml
|
|
||||||
apiVersion: v1
|
|
||||||
kind: Pod
|
|
||||||
metadata:
|
|
||||||
name: sysctl-example
|
|
||||||
annotations:
|
|
||||||
security.alpha.kubernetes.io/sysctls: kernel.shm_rmid_forced=1
|
|
||||||
security.alpha.kubernetes.io/unsafe-sysctls: net.ipv4.route.min_pmtu=1000,kernel.msgmax=1 2 3
|
|
||||||
spec:
|
|
||||||
...
|
|
||||||
```
|
|
||||||
|
|
||||||
**Note**: a pod with the _unsafe_ sysctls specified above will fail to launch on
|
|
||||||
any node which has not enabled those two _unsafe_ sysctls explicitly. As with
|
|
||||||
_node-level_ sysctls it is recommended to use [_taints and toleration_
|
|
||||||
feature](/docs/user-guide/kubectl/kubectl_taint.md) or [labels on nodes](/docs
|
|
||||||
/user-guide/labels.md) to schedule those pods onto the right nodes.
|
|
||||||
|
|||||||
@@ -0,0 +1,57 @@
|
|||||||
|
---
|
||||||
|
assignees:
|
||||||
|
- davidopp
|
||||||
|
- filipg
|
||||||
|
- piosz
|
||||||
|
title: Guaranteed Scheduling For Critical Add-On Pods
|
||||||
|
---
|
||||||
|
|
||||||
|
* TOC
|
||||||
|
{:toc}
|
||||||
|
|
||||||
|
## Overview
|
||||||
|
|
||||||
|
In addition to Kubernetes core components like api-server, scheduler, controller-manager running on a master machine
|
||||||
|
there are a number of add-ons which, for various reasons, must run on a regular cluster node (rather than the Kubernetes master).
|
||||||
|
Some of these add-ons are critical to a fully functional cluster, such as Heapster, DNS, and UI.
|
||||||
|
A cluster may stop working properly if a critical add-on is evicted (either manually or as a side effect of another operation like upgrade)
|
||||||
|
and becomes pending (for example when the cluster is highly utilized and either there are other pending pods that schedule into the space
|
||||||
|
vacated by the evicted critical add-on pod or the amount of resources available on the node changed for some other reason).
|
||||||
|
|
||||||
|
## Rescheduler: guaranteed scheduling of critical add-ons
|
||||||
|
|
||||||
|
Rescheduler ensures that critical add-ons are always scheduled
|
||||||
|
(assuming the cluster has enough resources to run the critical add-on pods in the absence of regular pods).
|
||||||
|
If the scheduler determines that no node has enough free resources to run the critical add-on pod
|
||||||
|
given the pods that are already running in the cluster
|
||||||
|
(indicated by critical add-on pod's pod condition PodScheduled set to false, the reason set to Unschedulable)
|
||||||
|
the rescheduler tries to free up space for the add-on by evicting some pods; then the scheduler will schedule the add-on pod.
|
||||||
|
|
||||||
|
To avoid situation when another pod is scheduled into the space prepared for the critical add-on,
|
||||||
|
the chosen node gets a temporary taint "CriticalAddonsOnly" before the eviction(s)
|
||||||
|
(see [more details](https://github.com/kubernetes/kubernetes/blob/master/docs/design/taint-toleration-dedicated.md)).
|
||||||
|
Each critical add-on has to tolerate it,
|
||||||
|
while the other pods shouldn't tolerate the taint. The taint is removed once the add-on is successfully scheduled.
|
||||||
|
|
||||||
|
*Warning:* currently there is no guarantee which node is chosen and which pods are being killed
|
||||||
|
in order to schedule critical pods, so if rescheduler is enabled your pods might be occasionally
|
||||||
|
killed for this purpose.
|
||||||
|
|
||||||
|
## Config
|
||||||
|
|
||||||
|
Rescheduler doesn't have any user facing configuration (component config) or API.
|
||||||
|
It's enabled by default. It can be disabled:
|
||||||
|
|
||||||
|
* during cluster setup by setting `ENABLE_RESCHEDULER` flag to `false`
|
||||||
|
* on running cluster by deleting its manifest from master node
|
||||||
|
(default path `/etc/kubernetes/manifests/rescheduler.manifest`)
|
||||||
|
|
||||||
|
### Marking add-on as critical
|
||||||
|
|
||||||
|
To be critical an add-on has to run in `kube-system` namespace (configurable via flag)
|
||||||
|
and have the following annotations specified:
|
||||||
|
|
||||||
|
* `scheduler.alpha.kubernetes.io/critical-pod` set to empty string
|
||||||
|
* `scheduler.alpha.kubernetes.io/tolerations` set to `[{"key":"CriticalAddonsOnly", "operator":"Exists"}]`
|
||||||
|
|
||||||
|
The first one marks a pod a critical. The second one is required by Rescheduler algorithm.
|
||||||
@@ -0,0 +1,66 @@
|
|||||||
|
---
|
||||||
|
assignees:
|
||||||
|
- davidopp
|
||||||
|
title: Using Multiple Clusters
|
||||||
|
---
|
||||||
|
|
||||||
|
You may want to set up multiple Kubernetes clusters, both to
|
||||||
|
have clusters in different regions to be nearer to your users, and to tolerate failures and/or invasive maintenance.
|
||||||
|
This document describes some of the issues to consider when making a decision about doing so.
|
||||||
|
|
||||||
|
If you decide to have multiple clusters, Kubernetes provides a way to [federate them](/docs/admin/federation/).
|
||||||
|
|
||||||
|
## Scope of a single cluster
|
||||||
|
|
||||||
|
On IaaS providers such as Google Compute Engine or Amazon Web Services, a VM exists in a
|
||||||
|
[zone](https://cloud.google.com/compute/docs/zones) or [availability
|
||||||
|
zone](http://docs.aws.amazon.com/AWSEC2/latest/UserGuide/using-regions-availability-zones.html).
|
||||||
|
We suggest that all the VMs in a Kubernetes cluster should be in the same availability zone, because:
|
||||||
|
|
||||||
|
- compared to having a single global Kubernetes cluster, there are fewer single-points of failure
|
||||||
|
- compared to a cluster that spans availability zones, it is easier to reason about the availability properties of a
|
||||||
|
single-zone cluster.
|
||||||
|
- when the Kubernetes developers are designing the system (e.g. making assumptions about latency, bandwidth, or
|
||||||
|
correlated failures) they are assuming all the machines are in a single data center, or otherwise closely connected.
|
||||||
|
|
||||||
|
It is okay to have multiple clusters per availability zone, though on balance we think fewer is better.
|
||||||
|
Reasons to prefer fewer clusters are:
|
||||||
|
|
||||||
|
- improved bin packing of Pods in some cases with more nodes in one cluster (less resource fragmentation)
|
||||||
|
- reduced operational overhead (though the advantage is diminished as ops tooling and processes matures)
|
||||||
|
- reduced costs for per-cluster fixed resource costs, e.g. apiserver VMs (but small as a percentage
|
||||||
|
of overall cluster cost for medium to large clusters).
|
||||||
|
|
||||||
|
Reasons to have multiple clusters include:
|
||||||
|
|
||||||
|
- strict security policies requiring isolation of one class of work from another (but, see Partitioning Clusters
|
||||||
|
below).
|
||||||
|
- test clusters to canary new Kubernetes releases or other cluster software.
|
||||||
|
|
||||||
|
## Selecting the right number of clusters
|
||||||
|
|
||||||
|
The selection of the number of Kubernetes clusters may be a relatively static choice, only revisited occasionally.
|
||||||
|
By contrast, the number of nodes in a cluster and the number of pods in a service may change frequently according to
|
||||||
|
load and growth.
|
||||||
|
|
||||||
|
To pick the number of clusters, first, decide which regions you need to be in to have adequate latency to all your end users, for services that will run
|
||||||
|
on Kubernetes (if you use a Content Distribution Network, the latency requirements for the CDN-hosted content need not
|
||||||
|
be considered). Legal issues might influence this as well. For example, a company with a global customer base might decide to have clusters in US, EU, AP, and SA regions.
|
||||||
|
Call the number of regions to be in `R`.
|
||||||
|
|
||||||
|
Second, decide how many clusters should be able to be unavailable at the same time, while still being available. Call
|
||||||
|
the number that can be unavailable `U`. If you are not sure, then 1 is a fine choice.
|
||||||
|
|
||||||
|
If it is allowable for load-balancing to direct traffic to any region in the event of a cluster failure, then
|
||||||
|
you need at least the larger of `R` or `U + 1` clusters. If it is not (e.g. you want to ensure low latency for all
|
||||||
|
users in the event of a cluster failure), then you need to have `R * (U + 1)` clusters
|
||||||
|
(`U + 1` in each of `R` regions). In any case, try to put each cluster in a different zone.
|
||||||
|
|
||||||
|
Finally, if any of your clusters would need more than the maximum recommended number of nodes for a Kubernetes cluster, then
|
||||||
|
you may need even more clusters. Kubernetes v1.3 supports clusters up to 1000 nodes in size.
|
||||||
|
|
||||||
|
## Working with multiple clusters
|
||||||
|
|
||||||
|
When you have multiple clusters, you would typically create services with the same config in each cluster and put each of those
|
||||||
|
service instances behind a load balancer (AWS Elastic Load Balancer, GCE Forwarding Rule or HTTP Load Balancer) spanning all of them, so that
|
||||||
|
failures of a single cluster are not visible to end users.
|
||||||
@@ -0,0 +1,122 @@
|
|||||||
|
---
|
||||||
|
assignees:
|
||||||
|
- sttts
|
||||||
|
title: Using Sysctls in a Kubernetes Cluster
|
||||||
|
---
|
||||||
|
|
||||||
|
* TOC
|
||||||
|
{:toc}
|
||||||
|
|
||||||
|
This document describes how sysctls are used within a Kubernetes cluster.
|
||||||
|
|
||||||
|
## What is a Sysctl?
|
||||||
|
|
||||||
|
In Linux, the sysctl interface allows an administrator to modify kernel
|
||||||
|
parameters at runtime. Parameters are available via the `/proc/sys/` virtual
|
||||||
|
process file system. The parameters cover various subsystems such as:
|
||||||
|
|
||||||
|
- kernel (common prefix: `kernel.`)
|
||||||
|
- networking (common prefix: `net.`)
|
||||||
|
- virtual memory (common prefix: `vm.`)
|
||||||
|
- MDADM (common prefix: `dev.`)
|
||||||
|
- More subsystems are described in [Kernel docs](https://www.kernel.org/doc/Documentation/sysctl/README).
|
||||||
|
|
||||||
|
To get a list of all parameters, you can run
|
||||||
|
|
||||||
|
```
|
||||||
|
$ sudo sysctl -a
|
||||||
|
```
|
||||||
|
|
||||||
|
## Namespaced vs. Node-Level Sysctls
|
||||||
|
|
||||||
|
A number of sysctls are _namespaced_ in today's Linux kernels. This means that
|
||||||
|
they can be set independently for each pod on a node. Being namespaced is a
|
||||||
|
requirement for sysctls to be accessible in a pod context within Kubernetes.
|
||||||
|
|
||||||
|
The following sysctls are known to be _namespaced_:
|
||||||
|
|
||||||
|
- `kernel.shm*`,
|
||||||
|
- `kernel.msg*`,
|
||||||
|
- `kernel.sem`,
|
||||||
|
- `fs.mqueue.*`,
|
||||||
|
- `net.*`.
|
||||||
|
|
||||||
|
Sysctls which are not namespaced are called _node-level_ and must be set
|
||||||
|
manually by the cluster admin, either by means of the underlying Linux
|
||||||
|
distribution of the nodes (e.g. via `/etc/sysctls.conf`) or using a DaemonSet
|
||||||
|
with privileged containers.
|
||||||
|
|
||||||
|
**Note**: it is good practice to consider nodes with special sysctl settings as
|
||||||
|
_tainted_ within a cluster, and only schedule pods onto them which need those
|
||||||
|
sysctl settings. It is suggested to use the Kubernetes [_taints and toleration_
|
||||||
|
feature](/docs/user-guide/kubectl/kubectl_taint.md) to implement this.
|
||||||
|
|
||||||
|
## Safe vs. Unsafe Sysctls
|
||||||
|
|
||||||
|
Sysctls are grouped into _safe_ and _unsafe_ sysctls. In addition to proper
|
||||||
|
namespacing a _safe_ sysctl must be properly _isolated_ between pods on the same
|
||||||
|
node. This means that setting a _safe_ sysctl for one pod
|
||||||
|
|
||||||
|
- must not have any influence on any other pod on the node
|
||||||
|
- must not allow to harm the node's health
|
||||||
|
- must not allow to gain CPU or memory resources outside of the resource limits
|
||||||
|
of a pod.
|
||||||
|
|
||||||
|
By far, most of the _namespaced_ sysctls are not necessarily considered _safe_.
|
||||||
|
|
||||||
|
For Kubernetes 1.4, the following sysctls are supported in the _safe_ set:
|
||||||
|
|
||||||
|
- `kernel.shm_rmid_forced`,
|
||||||
|
- `net.ipv4.ip_local_port_range`,
|
||||||
|
- `net.ipv4.tcp_syncookies`.
|
||||||
|
|
||||||
|
This list will be extended in future Kubernetes versions when the kubelet
|
||||||
|
supports better isolation mechanisms.
|
||||||
|
|
||||||
|
All _safe_ sysctls are enabled by default.
|
||||||
|
|
||||||
|
All _unsafe_ sysctls are disabled by default and must be allowed manually by the
|
||||||
|
cluster admin on a per-node basis. Pods with disabled unsafe sysctls will be
|
||||||
|
scheduled, but will fail to launch.
|
||||||
|
|
||||||
|
**Warning**: Due to their nature of being _unsafe_, the use of _unsafe_ sysctls
|
||||||
|
is at-your-own-risk and can lead to severe problems like wrong behavior of
|
||||||
|
containers, resource shortage or complete breakage of a node.
|
||||||
|
|
||||||
|
## Enabling Unsafe Sysctls
|
||||||
|
|
||||||
|
With the warning above in mind, the cluster admin can allow certain _unsafe_
|
||||||
|
sysctls for very special situations like e.g. high-performance or real-time
|
||||||
|
application tuning. _Unsafe_ sysctls are enabled on a node-by-node basis with a
|
||||||
|
flag of the kubelet, e.g.:
|
||||||
|
|
||||||
|
```shell
|
||||||
|
$ kubelet --experimental-allowed-unsafe-sysctls 'kernel.msg*,net.ipv4.route.min_pmtu' ...
|
||||||
|
```
|
||||||
|
|
||||||
|
Only _namespaced_ sysctls can be enabled this way.
|
||||||
|
|
||||||
|
## Setting Sysctls for a Pod
|
||||||
|
|
||||||
|
The sysctl feature is an alpha API in Kubernetes 1.4. Therefore, sysctls are set
|
||||||
|
using annotations on pods. They apply to all containers in the same pod.
|
||||||
|
|
||||||
|
Here is an example, with different annotations for _safe_ and _unsafe_ sysctls:
|
||||||
|
|
||||||
|
```yaml
|
||||||
|
apiVersion: v1
|
||||||
|
kind: Pod
|
||||||
|
metadata:
|
||||||
|
name: sysctl-example
|
||||||
|
annotations:
|
||||||
|
security.alpha.kubernetes.io/sysctls: kernel.shm_rmid_forced=1
|
||||||
|
security.alpha.kubernetes.io/unsafe-sysctls: net.ipv4.route.min_pmtu=1000,kernel.msgmax=1 2 3
|
||||||
|
spec:
|
||||||
|
...
|
||||||
|
```
|
||||||
|
|
||||||
|
**Note**: a pod with the _unsafe_ sysctls specified above will fail to launch on
|
||||||
|
any node which has not enabled those two _unsafe_ sysctls explicitly. As with
|
||||||
|
_node-level_ sysctls it is recommended to use [_taints and toleration_
|
||||||
|
feature](/docs/user-guide/kubectl/kubectl_taint.md) or [labels on nodes](/docs
|
||||||
|
/user-guide/labels.md) to schedule those pods onto the right nodes.
|
||||||
@@ -0,0 +1,136 @@
|
|||||||
|
---
|
||||||
|
assignees:
|
||||||
|
- lavalamp
|
||||||
|
title: Kubernetes Components
|
||||||
|
---
|
||||||
|
|
||||||
|
This document outlines the various binary components that need to run to
|
||||||
|
deliver a functioning Kubernetes cluster.
|
||||||
|
|
||||||
|
## Master Components
|
||||||
|
|
||||||
|
Master components are those that provide the cluster's control plane. For
|
||||||
|
example, master components are responsible for making global decisions about the
|
||||||
|
cluster (e.g., scheduling), and detecting and responding to cluster events
|
||||||
|
(e.g., starting up a new pod when a replication controller's 'replicas' field is
|
||||||
|
unsatisfied).
|
||||||
|
|
||||||
|
In theory, Master components can be run on any node in the cluster. However,
|
||||||
|
for simplicity, current set up scripts typically start all master components on
|
||||||
|
the same VM, and does not run user containers on this VM. See
|
||||||
|
[high-availability.md](/docs/admin/high-availability) for an example multi-master-VM setup.
|
||||||
|
|
||||||
|
Even in the future, when Kubernetes is fully self-hosting, it will probably be
|
||||||
|
wise to only allow master components to schedule on a subset of nodes, to limit
|
||||||
|
co-running with user-run pods, reducing the possible scope of a
|
||||||
|
node-compromising security exploit.
|
||||||
|
|
||||||
|
### kube-apiserver
|
||||||
|
|
||||||
|
[kube-apiserver](/docs/admin/kube-apiserver) exposes the Kubernetes API; it is the front-end for the
|
||||||
|
Kubernetes control plane. It is designed to scale horizontally (i.e., one scales
|
||||||
|
it by running more of them-- [high-availability.md](/docs/admin/high-availability)).
|
||||||
|
|
||||||
|
### etcd
|
||||||
|
|
||||||
|
[etcd](/docs/admin/etcd) is used as Kubernetes' backing store. All cluster data is stored here.
|
||||||
|
Proper administration of a Kubernetes cluster includes a backup plan for etcd's
|
||||||
|
data.
|
||||||
|
|
||||||
|
### kube-controller-manager
|
||||||
|
|
||||||
|
[kube-controller-manager](/docs/admin/kube-controller-manager) is a binary that runs controllers, which are the
|
||||||
|
background threads that handle routine tasks in the cluster. Logically, each
|
||||||
|
controller is a separate process, but to reduce the number of moving pieces in
|
||||||
|
the system, they are all compiled into a single binary and run in a single
|
||||||
|
process.
|
||||||
|
|
||||||
|
These controllers include:
|
||||||
|
|
||||||
|
* Node Controller: Responsible for noticing & responding when nodes go down.
|
||||||
|
* Replication Controller: Responsible for maintaining the correct number of pods for every replication
|
||||||
|
controller object in the system.
|
||||||
|
* Endpoints Controller: Populates the Endpoints object (i.e., join Services & Pods).
|
||||||
|
* Service Account & Token Controllers: Create default accounts and API access tokens for new namespaces.
|
||||||
|
* ... and others.
|
||||||
|
|
||||||
|
### kube-scheduler
|
||||||
|
|
||||||
|
[kube-scheduler](/docs/admin/kube-scheduler) watches newly created pods that have no node assigned, and
|
||||||
|
selects a node for them to run on.
|
||||||
|
|
||||||
|
### addons
|
||||||
|
|
||||||
|
Addons are pods and services that implement cluster features. The pods may be managed
|
||||||
|
by Deployments, ReplicationContollers, etc. Namespaced addon objects are created in
|
||||||
|
the "kube-system" namespace.
|
||||||
|
|
||||||
|
Addon manager takes the responsibility for creating and maintaining addon resources.
|
||||||
|
See [here](http://releases.k8s.io/HEAD/cluster/addons) for more details.
|
||||||
|
|
||||||
|
#### DNS
|
||||||
|
|
||||||
|
While the other addons are not strictly required, all Kubernetes
|
||||||
|
clusters should have [cluster DNS](/docs/admin/dns/), as many examples rely on it.
|
||||||
|
|
||||||
|
Cluster DNS is a DNS server, in addition to the other DNS server(s) in your
|
||||||
|
environment, which serves DNS records for Kubernetes services.
|
||||||
|
|
||||||
|
Containers started by Kubernetes automatically include this DNS server
|
||||||
|
in their DNS searches.
|
||||||
|
|
||||||
|
#### User interface
|
||||||
|
|
||||||
|
The kube-ui provides a read-only overview of the cluster state. Access
|
||||||
|
[the UI using kubectl proxy](/docs/user-guide/connecting-to-applications-proxy/#connecting-to-the-kube-ui-service-from-your-local-workstation)
|
||||||
|
|
||||||
|
#### Container Resource Monitoring
|
||||||
|
|
||||||
|
[Container Resource Monitoring](/docs/user-guide/monitoring) records generic time-series metrics
|
||||||
|
about containers in a central database, and provides a UI for browsing that data.
|
||||||
|
|
||||||
|
#### Cluster-level Logging
|
||||||
|
|
||||||
|
A [Cluster-level logging](/docs/user-guide/logging/overview) mechanism is responsible for
|
||||||
|
saving container logs to a central log store with search/browsing interface.
|
||||||
|
|
||||||
|
## Node components
|
||||||
|
|
||||||
|
Node components run on every node, maintaining running pods and providing them
|
||||||
|
the Kubernetes runtime environment.
|
||||||
|
|
||||||
|
### kubelet
|
||||||
|
|
||||||
|
[kubelet](/docs/admin/kubelet) is the primary node agent. It:
|
||||||
|
|
||||||
|
* Watches for pods that have been assigned to its node (either by apiserver
|
||||||
|
or via local configuration file) and:
|
||||||
|
* Mounts the pod's required volumes
|
||||||
|
* Downloads the pod's secrets
|
||||||
|
* Runs the pod's containers via docker (or, experimentally, rkt).
|
||||||
|
* Periodically executes any requested container liveness probes.
|
||||||
|
* Reports the status of the pod back to the rest of the system, by creating a
|
||||||
|
"mirror pod" if necessary.
|
||||||
|
* Reports the status of the node back to the rest of the system.
|
||||||
|
|
||||||
|
### kube-proxy
|
||||||
|
|
||||||
|
[kube-proxy](/docs/admin/kube-proxy) enables the Kubernetes service abstraction by maintaining
|
||||||
|
network rules on the host and performing connection forwarding.
|
||||||
|
|
||||||
|
### docker
|
||||||
|
|
||||||
|
`docker` is of course used for actually running containers.
|
||||||
|
|
||||||
|
### rkt
|
||||||
|
|
||||||
|
`rkt` is supported experimentally as an alternative to docker.
|
||||||
|
|
||||||
|
### supervisord
|
||||||
|
|
||||||
|
`supervisord` is a lightweight process babysitting system for keeping kubelet and docker
|
||||||
|
running.
|
||||||
|
|
||||||
|
### fluentd
|
||||||
|
|
||||||
|
`fluentd` is a daemon which helps provide [cluster-level logging](#cluster-level-logging).
|
||||||
Reference in New Issue
Block a user