diff --git a/_data/concepts.yml b/_data/concepts.yml index c15ad31a62..93c06c68a6 100644 --- a/_data/concepts.yml +++ b/_data/concepts.yml @@ -3,6 +3,10 @@ abstract: "Detailed explanations of Kubernetes system concepts and abstractions. toc: - docs/concepts/index.md +- title: Overview + section: + - docs/concepts/overview/components.md + - title: Kubernetes Objects section: - docs/concepts/abstractions/overview.md @@ -29,7 +33,10 @@ toc: - title: Cluster Administration section: - docs/concepts/cluster-administration/logging.md + - docs/concepts/cluster-administration/multiple-clusters.md - docs/concepts/cluster-administration/federation.md + - docs/concepts/cluster-administration/guaranteed-scheduling-critical-addon-pods.md + - docs/concepts/cluster-administration/sysctl-cluster.md - title: Configuration section: diff --git a/docs/admin/cluster-components.md b/docs/admin/cluster-components.md index 280b9b1f2c..49154b1750 100644 --- a/docs/admin/cluster-components.md +++ b/docs/admin/cluster-components.md @@ -4,133 +4,6 @@ assignees: title: Kubernetes Components --- -This document outlines the various binary components that need to run to -deliver a functioning Kubernetes cluster. +{% include user-guide-content-moved.md %} -## Master Components - -Master components are those that provide the cluster's control plane. For -example, master components are responsible for making global decisions about the -cluster (e.g., scheduling), and detecting and responding to cluster events -(e.g., starting up a new pod when a replication controller's 'replicas' field is -unsatisfied). - -In theory, Master components can be run on any node in the cluster. However, -for simplicity, current set up scripts typically start all master components on -the same VM, and does not run user containers on this VM. See -[high-availability.md](/docs/admin/high-availability) for an example multi-master-VM setup. - -Even in the future, when Kubernetes is fully self-hosting, it will probably be -wise to only allow master components to schedule on a subset of nodes, to limit -co-running with user-run pods, reducing the possible scope of a -node-compromising security exploit. - -### kube-apiserver - -[kube-apiserver](/docs/admin/kube-apiserver) exposes the Kubernetes API; it is the front-end for the -Kubernetes control plane. It is designed to scale horizontally (i.e., one scales -it by running more of them-- [high-availability.md](/docs/admin/high-availability)). - -### etcd - -[etcd](/docs/admin/etcd) is used as Kubernetes' backing store. All cluster data is stored here. -Proper administration of a Kubernetes cluster includes a backup plan for etcd's -data. - -### kube-controller-manager - -[kube-controller-manager](/docs/admin/kube-controller-manager) is a binary that runs controllers, which are the -background threads that handle routine tasks in the cluster. Logically, each -controller is a separate process, but to reduce the number of moving pieces in -the system, they are all compiled into a single binary and run in a single -process. - -These controllers include: - -* Node Controller: Responsible for noticing & responding when nodes go down. -* Replication Controller: Responsible for maintaining the correct number of pods for every replication - controller object in the system. -* Endpoints Controller: Populates the Endpoints object (i.e., join Services & Pods). -* Service Account & Token Controllers: Create default accounts and API access tokens for new namespaces. -* ... and others. - -### kube-scheduler - -[kube-scheduler](/docs/admin/kube-scheduler) watches newly created pods that have no node assigned, and -selects a node for them to run on. - -### addons - -Addons are pods and services that implement cluster features. The pods may be managed -by Deployments, ReplicationContollers, etc. Namespaced addon objects are created in -the "kube-system" namespace. - -Addon manager takes the responsibility for creating and maintaining addon resources. -See [here](http://releases.k8s.io/HEAD/cluster/addons) for more details. - -#### DNS - -While the other addons are not strictly required, all Kubernetes -clusters should have [cluster DNS](/docs/admin/dns/), as many examples rely on it. - -Cluster DNS is a DNS server, in addition to the other DNS server(s) in your -environment, which serves DNS records for Kubernetes services. - -Containers started by Kubernetes automatically include this DNS server -in their DNS searches. - -#### User interface - -The kube-ui provides a read-only overview of the cluster state. Access -[the UI using kubectl proxy](/docs/user-guide/connecting-to-applications-proxy/#connecting-to-the-kube-ui-service-from-your-local-workstation) - -#### Container Resource Monitoring - -[Container Resource Monitoring](/docs/user-guide/monitoring) records generic time-series metrics -about containers in a central database, and provides a UI for browsing that data. - -#### Cluster-level Logging - -A [Cluster-level logging](/docs/user-guide/logging/overview) mechanism is responsible for -saving container logs to a central log store with search/browsing interface. - -## Node components - -Node components run on every node, maintaining running pods and providing them -the Kubernetes runtime environment. - -### kubelet - -[kubelet](/docs/admin/kubelet) is the primary node agent. It: - -* Watches for pods that have been assigned to its node (either by apiserver - or via local configuration file) and: -* Mounts the pod's required volumes -* Downloads the pod's secrets -* Runs the pod's containers via docker (or, experimentally, rkt). -* Periodically executes any requested container liveness probes. -* Reports the status of the pod back to the rest of the system, by creating a - "mirror pod" if necessary. -* Reports the status of the node back to the rest of the system. - -### kube-proxy - -[kube-proxy](/docs/admin/kube-proxy) enables the Kubernetes service abstraction by maintaining -network rules on the host and performing connection forwarding. - -### docker - -`docker` is of course used for actually running containers. - -### rkt - -`rkt` is supported experimentally as an alternative to docker. - -### supervisord - -`supervisord` is a lightweight process babysitting system for keeping kubelet and docker -running. - -### fluentd - -`fluentd` is a daemon which helps provide [cluster-level logging](#cluster-level-logging). +[Kubernetes Components](/docs/concepts/overview/components/) diff --git a/docs/admin/multi-cluster.md b/docs/admin/multi-cluster.md index 085a9afa9f..4fe4e8b6ae 100644 --- a/docs/admin/multi-cluster.md +++ b/docs/admin/multi-cluster.md @@ -4,63 +4,6 @@ assignees: title: Using Multiple Clusters --- -You may want to set up multiple Kubernetes clusters, both to -have clusters in different regions to be nearer to your users, and to tolerate failures and/or invasive maintenance. -This document describes some of the issues to consider when making a decision about doing so. +{% include user-guide-content-moved.md %} -If you decide to have multiple clusters, Kubernetes provides a way to [federate them](/docs/admin/federation/). - -## Scope of a single cluster - -On IaaS providers such as Google Compute Engine or Amazon Web Services, a VM exists in a -[zone](https://cloud.google.com/compute/docs/zones) or [availability -zone](http://docs.aws.amazon.com/AWSEC2/latest/UserGuide/using-regions-availability-zones.html). -We suggest that all the VMs in a Kubernetes cluster should be in the same availability zone, because: - - - compared to having a single global Kubernetes cluster, there are fewer single-points of failure - - compared to a cluster that spans availability zones, it is easier to reason about the availability properties of a - single-zone cluster. - - when the Kubernetes developers are designing the system (e.g. making assumptions about latency, bandwidth, or - correlated failures) they are assuming all the machines are in a single data center, or otherwise closely connected. - -It is okay to have multiple clusters per availability zone, though on balance we think fewer is better. -Reasons to prefer fewer clusters are: - - - improved bin packing of Pods in some cases with more nodes in one cluster (less resource fragmentation) - - reduced operational overhead (though the advantage is diminished as ops tooling and processes matures) - - reduced costs for per-cluster fixed resource costs, e.g. apiserver VMs (but small as a percentage - of overall cluster cost for medium to large clusters). - -Reasons to have multiple clusters include: - - - strict security policies requiring isolation of one class of work from another (but, see Partitioning Clusters - below). - - test clusters to canary new Kubernetes releases or other cluster software. - -## Selecting the right number of clusters - -The selection of the number of Kubernetes clusters may be a relatively static choice, only revisited occasionally. -By contrast, the number of nodes in a cluster and the number of pods in a service may change frequently according to -load and growth. - -To pick the number of clusters, first, decide which regions you need to be in to have adequate latency to all your end users, for services that will run -on Kubernetes (if you use a Content Distribution Network, the latency requirements for the CDN-hosted content need not -be considered). Legal issues might influence this as well. For example, a company with a global customer base might decide to have clusters in US, EU, AP, and SA regions. -Call the number of regions to be in `R`. - -Second, decide how many clusters should be able to be unavailable at the same time, while still being available. Call -the number that can be unavailable `U`. If you are not sure, then 1 is a fine choice. - -If it is allowable for load-balancing to direct traffic to any region in the event of a cluster failure, then -you need at least the larger of `R` or `U + 1` clusters. If it is not (e.g. you want to ensure low latency for all -users in the event of a cluster failure), then you need to have `R * (U + 1)` clusters -(`U + 1` in each of `R` regions). In any case, try to put each cluster in a different zone. - -Finally, if any of your clusters would need more than the maximum recommended number of nodes for a Kubernetes cluster, then -you may need even more clusters. Kubernetes v1.3 supports clusters up to 1000 nodes in size. - -## Working with multiple clusters - -When you have multiple clusters, you would typically create services with the same config in each cluster and put each of those -service instances behind a load balancer (AWS Elastic Load Balancer, GCE Forwarding Rule or HTTP Load Balancer) spanning all of them, so that -failures of a single cluster are not visible to end users. +[Using Multiple Clusters](/docs/concepts/cluster-administration/multiple-clusters/) diff --git a/docs/admin/rescheduler.md b/docs/admin/rescheduler.md index 9e3fc61c39..d8d418ac2b 100644 --- a/docs/admin/rescheduler.md +++ b/docs/admin/rescheduler.md @@ -6,52 +6,6 @@ assignees: title: Guaranteed Scheduling For Critical Add-On Pods --- -* TOC -{:toc} +{% include user-guide-content-moved.md %} -## Overview - -In addition to Kubernetes core components like api-server, scheduler, controller-manager running on a master machine -there are a number of add-ons which, for various reasons, must run on a regular cluster node (rather than the Kubernetes master). -Some of these add-ons are critical to a fully functional cluster, such as Heapster, DNS, and UI. -A cluster may stop working properly if a critical add-on is evicted (either manually or as a side effect of another operation like upgrade) -and becomes pending (for example when the cluster is highly utilized and either there are other pending pods that schedule into the space -vacated by the evicted critical add-on pod or the amount of resources available on the node changed for some other reason). - -## Rescheduler: guaranteed scheduling of critical add-ons - -Rescheduler ensures that critical add-ons are always scheduled -(assuming the cluster has enough resources to run the critical add-on pods in the absence of regular pods). -If the scheduler determines that no node has enough free resources to run the critical add-on pod -given the pods that are already running in the cluster -(indicated by critical add-on pod's pod condition PodScheduled set to false, the reason set to Unschedulable) -the rescheduler tries to free up space for the add-on by evicting some pods; then the scheduler will schedule the add-on pod. - -To avoid situation when another pod is scheduled into the space prepared for the critical add-on, -the chosen node gets a temporary taint "CriticalAddonsOnly" before the eviction(s) -(see [more details](https://github.com/kubernetes/kubernetes/blob/master/docs/design/taint-toleration-dedicated.md)). -Each critical add-on has to tolerate it, -while the other pods shouldn't tolerate the taint. The taint is removed once the add-on is successfully scheduled. - -*Warning:* currently there is no guarantee which node is chosen and which pods are being killed -in order to schedule critical pods, so if rescheduler is enabled your pods might be occasionally -killed for this purpose. - -## Config - -Rescheduler doesn't have any user facing configuration (component config) or API. -It's enabled by default. It can be disabled: - -* during cluster setup by setting `ENABLE_RESCHEDULER` flag to `false` -* on running cluster by deleting its manifest from master node -(default path `/etc/kubernetes/manifests/rescheduler.manifest`) - -### Marking add-on as critical - -To be critical an add-on has to run in `kube-system` namespace (configurable via flag) -and have the following annotations specified: - -* `scheduler.alpha.kubernetes.io/critical-pod` set to empty string -* `scheduler.alpha.kubernetes.io/tolerations` set to `[{"key":"CriticalAddonsOnly", "operator":"Exists"}]` - -The first one marks a pod a critical. The second one is required by Rescheduler algorithm. +[Guaranteed Scheduling for Critical Add-On Pods](/docs/concepts/cluster-administration/guaranteed-scheduling-critical-addon-pods/) diff --git a/docs/admin/sysctls.md b/docs/admin/sysctls.md index aa75c4df2a..4931a8d6bf 100644 --- a/docs/admin/sysctls.md +++ b/docs/admin/sysctls.md @@ -4,119 +4,6 @@ assignees: title: Using Sysctls in a Kubernetes Cluster --- -* TOC -{:toc} +{% include user-guide-content-moved.md %} -This document describes how sysctls are used within a Kubernetes cluster. - -## What is a Sysctl? - -In Linux, the sysctl interface allows an administrator to modify kernel -parameters at runtime. Parameters are available via the `/proc/sys/` virtual -process file system. The parameters cover various subsystems such as: - -- kernel (common prefix: `kernel.`) -- networking (common prefix: `net.`) -- virtual memory (common prefix: `vm.`) -- MDADM (common prefix: `dev.`) -- More subsystems are described in [Kernel docs](https://www.kernel.org/doc/Documentation/sysctl/README). - -To get a list of all parameters, you can run - -``` -$ sudo sysctl -a -``` - -## Namespaced vs. Node-Level Sysctls - -A number of sysctls are _namespaced_ in today's Linux kernels. This means that -they can be set independently for each pod on a node. Being namespaced is a -requirement for sysctls to be accessible in a pod context within Kubernetes. - -The following sysctls are known to be _namespaced_: - -- `kernel.shm*`, -- `kernel.msg*`, -- `kernel.sem`, -- `fs.mqueue.*`, -- `net.*`. - -Sysctls which are not namespaced are called _node-level_ and must be set -manually by the cluster admin, either by means of the underlying Linux -distribution of the nodes (e.g. via `/etc/sysctls.conf`) or using a DaemonSet -with privileged containers. - -**Note**: it is good practice to consider nodes with special sysctl settings as -_tainted_ within a cluster, and only schedule pods onto them which need those -sysctl settings. It is suggested to use the Kubernetes [_taints and toleration_ -feature](/docs/user-guide/kubectl/kubectl_taint.md) to implement this. - -## Safe vs. Unsafe Sysctls - -Sysctls are grouped into _safe_ and _unsafe_ sysctls. In addition to proper -namespacing a _safe_ sysctl must be properly _isolated_ between pods on the same -node. This means that setting a _safe_ sysctl for one pod - -- must not have any influence on any other pod on the node -- must not allow to harm the node's health -- must not allow to gain CPU or memory resources outside of the resource limits - of a pod. - -By far, most of the _namespaced_ sysctls are not necessarily considered _safe_. - -For Kubernetes 1.4, the following sysctls are supported in the _safe_ set: - -- `kernel.shm_rmid_forced`, -- `net.ipv4.ip_local_port_range`, -- `net.ipv4.tcp_syncookies`. - -This list will be extended in future Kubernetes versions when the kubelet -supports better isolation mechanisms. - -All _safe_ sysctls are enabled by default. - -All _unsafe_ sysctls are disabled by default and must be allowed manually by the -cluster admin on a per-node basis. Pods with disabled unsafe sysctls will be -scheduled, but will fail to launch. - -**Warning**: Due to their nature of being _unsafe_, the use of _unsafe_ sysctls -is at-your-own-risk and can lead to severe problems like wrong behavior of -containers, resource shortage or complete breakage of a node. - -## Enabling Unsafe Sysctls - -With the warning above in mind, the cluster admin can allow certain _unsafe_ -sysctls for very special situations like e.g. high-performance or real-time -application tuning. _Unsafe_ sysctls are enabled on a node-by-node basis with a -flag of the kubelet, e.g.: - -```shell -$ kubelet --experimental-allowed-unsafe-sysctls 'kernel.msg*,net.ipv4.route.min_pmtu' ... -``` - -Only _namespaced_ sysctls can be enabled this way. - -## Setting Sysctls for a Pod - -The sysctl feature is an alpha API in Kubernetes 1.4. Therefore, sysctls are set -using annotations on pods. They apply to all containers in the same pod. - -Here is an example, with different annotations for _safe_ and _unsafe_ sysctls: - -```yaml -apiVersion: v1 -kind: Pod -metadata: - name: sysctl-example - annotations: - security.alpha.kubernetes.io/sysctls: kernel.shm_rmid_forced=1 - security.alpha.kubernetes.io/unsafe-sysctls: net.ipv4.route.min_pmtu=1000,kernel.msgmax=1 2 3 -spec: - ... -``` - -**Note**: a pod with the _unsafe_ sysctls specified above will fail to launch on -any node which has not enabled those two _unsafe_ sysctls explicitly. As with -_node-level_ sysctls it is recommended to use [_taints and toleration_ -feature](/docs/user-guide/kubectl/kubectl_taint.md) or [labels on nodes](/docs -/user-guide/labels.md) to schedule those pods onto the right nodes. +[Using Sysctls in a Kubernetes Cluster](/docs/concepts/cluster-administration/sysctl-cluster/) diff --git a/docs/concepts/cluster-administration/guaranteed-scheduling-critical-addon-pods.md b/docs/concepts/cluster-administration/guaranteed-scheduling-critical-addon-pods.md new file mode 100644 index 0000000000..9e3fc61c39 --- /dev/null +++ b/docs/concepts/cluster-administration/guaranteed-scheduling-critical-addon-pods.md @@ -0,0 +1,57 @@ +--- +assignees: +- davidopp +- filipg +- piosz +title: Guaranteed Scheduling For Critical Add-On Pods +--- + +* TOC +{:toc} + +## Overview + +In addition to Kubernetes core components like api-server, scheduler, controller-manager running on a master machine +there are a number of add-ons which, for various reasons, must run on a regular cluster node (rather than the Kubernetes master). +Some of these add-ons are critical to a fully functional cluster, such as Heapster, DNS, and UI. +A cluster may stop working properly if a critical add-on is evicted (either manually or as a side effect of another operation like upgrade) +and becomes pending (for example when the cluster is highly utilized and either there are other pending pods that schedule into the space +vacated by the evicted critical add-on pod or the amount of resources available on the node changed for some other reason). + +## Rescheduler: guaranteed scheduling of critical add-ons + +Rescheduler ensures that critical add-ons are always scheduled +(assuming the cluster has enough resources to run the critical add-on pods in the absence of regular pods). +If the scheduler determines that no node has enough free resources to run the critical add-on pod +given the pods that are already running in the cluster +(indicated by critical add-on pod's pod condition PodScheduled set to false, the reason set to Unschedulable) +the rescheduler tries to free up space for the add-on by evicting some pods; then the scheduler will schedule the add-on pod. + +To avoid situation when another pod is scheduled into the space prepared for the critical add-on, +the chosen node gets a temporary taint "CriticalAddonsOnly" before the eviction(s) +(see [more details](https://github.com/kubernetes/kubernetes/blob/master/docs/design/taint-toleration-dedicated.md)). +Each critical add-on has to tolerate it, +while the other pods shouldn't tolerate the taint. The taint is removed once the add-on is successfully scheduled. + +*Warning:* currently there is no guarantee which node is chosen and which pods are being killed +in order to schedule critical pods, so if rescheduler is enabled your pods might be occasionally +killed for this purpose. + +## Config + +Rescheduler doesn't have any user facing configuration (component config) or API. +It's enabled by default. It can be disabled: + +* during cluster setup by setting `ENABLE_RESCHEDULER` flag to `false` +* on running cluster by deleting its manifest from master node +(default path `/etc/kubernetes/manifests/rescheduler.manifest`) + +### Marking add-on as critical + +To be critical an add-on has to run in `kube-system` namespace (configurable via flag) +and have the following annotations specified: + +* `scheduler.alpha.kubernetes.io/critical-pod` set to empty string +* `scheduler.alpha.kubernetes.io/tolerations` set to `[{"key":"CriticalAddonsOnly", "operator":"Exists"}]` + +The first one marks a pod a critical. The second one is required by Rescheduler algorithm. diff --git a/docs/concepts/cluster-administration/multiple-clusters.md b/docs/concepts/cluster-administration/multiple-clusters.md new file mode 100644 index 0000000000..085a9afa9f --- /dev/null +++ b/docs/concepts/cluster-administration/multiple-clusters.md @@ -0,0 +1,66 @@ +--- +assignees: +- davidopp +title: Using Multiple Clusters +--- + +You may want to set up multiple Kubernetes clusters, both to +have clusters in different regions to be nearer to your users, and to tolerate failures and/or invasive maintenance. +This document describes some of the issues to consider when making a decision about doing so. + +If you decide to have multiple clusters, Kubernetes provides a way to [federate them](/docs/admin/federation/). + +## Scope of a single cluster + +On IaaS providers such as Google Compute Engine or Amazon Web Services, a VM exists in a +[zone](https://cloud.google.com/compute/docs/zones) or [availability +zone](http://docs.aws.amazon.com/AWSEC2/latest/UserGuide/using-regions-availability-zones.html). +We suggest that all the VMs in a Kubernetes cluster should be in the same availability zone, because: + + - compared to having a single global Kubernetes cluster, there are fewer single-points of failure + - compared to a cluster that spans availability zones, it is easier to reason about the availability properties of a + single-zone cluster. + - when the Kubernetes developers are designing the system (e.g. making assumptions about latency, bandwidth, or + correlated failures) they are assuming all the machines are in a single data center, or otherwise closely connected. + +It is okay to have multiple clusters per availability zone, though on balance we think fewer is better. +Reasons to prefer fewer clusters are: + + - improved bin packing of Pods in some cases with more nodes in one cluster (less resource fragmentation) + - reduced operational overhead (though the advantage is diminished as ops tooling and processes matures) + - reduced costs for per-cluster fixed resource costs, e.g. apiserver VMs (but small as a percentage + of overall cluster cost for medium to large clusters). + +Reasons to have multiple clusters include: + + - strict security policies requiring isolation of one class of work from another (but, see Partitioning Clusters + below). + - test clusters to canary new Kubernetes releases or other cluster software. + +## Selecting the right number of clusters + +The selection of the number of Kubernetes clusters may be a relatively static choice, only revisited occasionally. +By contrast, the number of nodes in a cluster and the number of pods in a service may change frequently according to +load and growth. + +To pick the number of clusters, first, decide which regions you need to be in to have adequate latency to all your end users, for services that will run +on Kubernetes (if you use a Content Distribution Network, the latency requirements for the CDN-hosted content need not +be considered). Legal issues might influence this as well. For example, a company with a global customer base might decide to have clusters in US, EU, AP, and SA regions. +Call the number of regions to be in `R`. + +Second, decide how many clusters should be able to be unavailable at the same time, while still being available. Call +the number that can be unavailable `U`. If you are not sure, then 1 is a fine choice. + +If it is allowable for load-balancing to direct traffic to any region in the event of a cluster failure, then +you need at least the larger of `R` or `U + 1` clusters. If it is not (e.g. you want to ensure low latency for all +users in the event of a cluster failure), then you need to have `R * (U + 1)` clusters +(`U + 1` in each of `R` regions). In any case, try to put each cluster in a different zone. + +Finally, if any of your clusters would need more than the maximum recommended number of nodes for a Kubernetes cluster, then +you may need even more clusters. Kubernetes v1.3 supports clusters up to 1000 nodes in size. + +## Working with multiple clusters + +When you have multiple clusters, you would typically create services with the same config in each cluster and put each of those +service instances behind a load balancer (AWS Elastic Load Balancer, GCE Forwarding Rule or HTTP Load Balancer) spanning all of them, so that +failures of a single cluster are not visible to end users. diff --git a/docs/concepts/cluster-administration/sysctl-cluster.md b/docs/concepts/cluster-administration/sysctl-cluster.md new file mode 100644 index 0000000000..aa75c4df2a --- /dev/null +++ b/docs/concepts/cluster-administration/sysctl-cluster.md @@ -0,0 +1,122 @@ +--- +assignees: +- sttts +title: Using Sysctls in a Kubernetes Cluster +--- + +* TOC +{:toc} + +This document describes how sysctls are used within a Kubernetes cluster. + +## What is a Sysctl? + +In Linux, the sysctl interface allows an administrator to modify kernel +parameters at runtime. Parameters are available via the `/proc/sys/` virtual +process file system. The parameters cover various subsystems such as: + +- kernel (common prefix: `kernel.`) +- networking (common prefix: `net.`) +- virtual memory (common prefix: `vm.`) +- MDADM (common prefix: `dev.`) +- More subsystems are described in [Kernel docs](https://www.kernel.org/doc/Documentation/sysctl/README). + +To get a list of all parameters, you can run + +``` +$ sudo sysctl -a +``` + +## Namespaced vs. Node-Level Sysctls + +A number of sysctls are _namespaced_ in today's Linux kernels. This means that +they can be set independently for each pod on a node. Being namespaced is a +requirement for sysctls to be accessible in a pod context within Kubernetes. + +The following sysctls are known to be _namespaced_: + +- `kernel.shm*`, +- `kernel.msg*`, +- `kernel.sem`, +- `fs.mqueue.*`, +- `net.*`. + +Sysctls which are not namespaced are called _node-level_ and must be set +manually by the cluster admin, either by means of the underlying Linux +distribution of the nodes (e.g. via `/etc/sysctls.conf`) or using a DaemonSet +with privileged containers. + +**Note**: it is good practice to consider nodes with special sysctl settings as +_tainted_ within a cluster, and only schedule pods onto them which need those +sysctl settings. It is suggested to use the Kubernetes [_taints and toleration_ +feature](/docs/user-guide/kubectl/kubectl_taint.md) to implement this. + +## Safe vs. Unsafe Sysctls + +Sysctls are grouped into _safe_ and _unsafe_ sysctls. In addition to proper +namespacing a _safe_ sysctl must be properly _isolated_ between pods on the same +node. This means that setting a _safe_ sysctl for one pod + +- must not have any influence on any other pod on the node +- must not allow to harm the node's health +- must not allow to gain CPU or memory resources outside of the resource limits + of a pod. + +By far, most of the _namespaced_ sysctls are not necessarily considered _safe_. + +For Kubernetes 1.4, the following sysctls are supported in the _safe_ set: + +- `kernel.shm_rmid_forced`, +- `net.ipv4.ip_local_port_range`, +- `net.ipv4.tcp_syncookies`. + +This list will be extended in future Kubernetes versions when the kubelet +supports better isolation mechanisms. + +All _safe_ sysctls are enabled by default. + +All _unsafe_ sysctls are disabled by default and must be allowed manually by the +cluster admin on a per-node basis. Pods with disabled unsafe sysctls will be +scheduled, but will fail to launch. + +**Warning**: Due to their nature of being _unsafe_, the use of _unsafe_ sysctls +is at-your-own-risk and can lead to severe problems like wrong behavior of +containers, resource shortage or complete breakage of a node. + +## Enabling Unsafe Sysctls + +With the warning above in mind, the cluster admin can allow certain _unsafe_ +sysctls for very special situations like e.g. high-performance or real-time +application tuning. _Unsafe_ sysctls are enabled on a node-by-node basis with a +flag of the kubelet, e.g.: + +```shell +$ kubelet --experimental-allowed-unsafe-sysctls 'kernel.msg*,net.ipv4.route.min_pmtu' ... +``` + +Only _namespaced_ sysctls can be enabled this way. + +## Setting Sysctls for a Pod + +The sysctl feature is an alpha API in Kubernetes 1.4. Therefore, sysctls are set +using annotations on pods. They apply to all containers in the same pod. + +Here is an example, with different annotations for _safe_ and _unsafe_ sysctls: + +```yaml +apiVersion: v1 +kind: Pod +metadata: + name: sysctl-example + annotations: + security.alpha.kubernetes.io/sysctls: kernel.shm_rmid_forced=1 + security.alpha.kubernetes.io/unsafe-sysctls: net.ipv4.route.min_pmtu=1000,kernel.msgmax=1 2 3 +spec: + ... +``` + +**Note**: a pod with the _unsafe_ sysctls specified above will fail to launch on +any node which has not enabled those two _unsafe_ sysctls explicitly. As with +_node-level_ sysctls it is recommended to use [_taints and toleration_ +feature](/docs/user-guide/kubectl/kubectl_taint.md) or [labels on nodes](/docs +/user-guide/labels.md) to schedule those pods onto the right nodes. diff --git a/docs/concepts/overview/components.md b/docs/concepts/overview/components.md new file mode 100644 index 0000000000..280b9b1f2c --- /dev/null +++ b/docs/concepts/overview/components.md @@ -0,0 +1,136 @@ +--- +assignees: +- lavalamp +title: Kubernetes Components +--- + +This document outlines the various binary components that need to run to +deliver a functioning Kubernetes cluster. + +## Master Components + +Master components are those that provide the cluster's control plane. For +example, master components are responsible for making global decisions about the +cluster (e.g., scheduling), and detecting and responding to cluster events +(e.g., starting up a new pod when a replication controller's 'replicas' field is +unsatisfied). + +In theory, Master components can be run on any node in the cluster. However, +for simplicity, current set up scripts typically start all master components on +the same VM, and does not run user containers on this VM. See +[high-availability.md](/docs/admin/high-availability) for an example multi-master-VM setup. + +Even in the future, when Kubernetes is fully self-hosting, it will probably be +wise to only allow master components to schedule on a subset of nodes, to limit +co-running with user-run pods, reducing the possible scope of a +node-compromising security exploit. + +### kube-apiserver + +[kube-apiserver](/docs/admin/kube-apiserver) exposes the Kubernetes API; it is the front-end for the +Kubernetes control plane. It is designed to scale horizontally (i.e., one scales +it by running more of them-- [high-availability.md](/docs/admin/high-availability)). + +### etcd + +[etcd](/docs/admin/etcd) is used as Kubernetes' backing store. All cluster data is stored here. +Proper administration of a Kubernetes cluster includes a backup plan for etcd's +data. + +### kube-controller-manager + +[kube-controller-manager](/docs/admin/kube-controller-manager) is a binary that runs controllers, which are the +background threads that handle routine tasks in the cluster. Logically, each +controller is a separate process, but to reduce the number of moving pieces in +the system, they are all compiled into a single binary and run in a single +process. + +These controllers include: + +* Node Controller: Responsible for noticing & responding when nodes go down. +* Replication Controller: Responsible for maintaining the correct number of pods for every replication + controller object in the system. +* Endpoints Controller: Populates the Endpoints object (i.e., join Services & Pods). +* Service Account & Token Controllers: Create default accounts and API access tokens for new namespaces. +* ... and others. + +### kube-scheduler + +[kube-scheduler](/docs/admin/kube-scheduler) watches newly created pods that have no node assigned, and +selects a node for them to run on. + +### addons + +Addons are pods and services that implement cluster features. The pods may be managed +by Deployments, ReplicationContollers, etc. Namespaced addon objects are created in +the "kube-system" namespace. + +Addon manager takes the responsibility for creating and maintaining addon resources. +See [here](http://releases.k8s.io/HEAD/cluster/addons) for more details. + +#### DNS + +While the other addons are not strictly required, all Kubernetes +clusters should have [cluster DNS](/docs/admin/dns/), as many examples rely on it. + +Cluster DNS is a DNS server, in addition to the other DNS server(s) in your +environment, which serves DNS records for Kubernetes services. + +Containers started by Kubernetes automatically include this DNS server +in their DNS searches. + +#### User interface + +The kube-ui provides a read-only overview of the cluster state. Access +[the UI using kubectl proxy](/docs/user-guide/connecting-to-applications-proxy/#connecting-to-the-kube-ui-service-from-your-local-workstation) + +#### Container Resource Monitoring + +[Container Resource Monitoring](/docs/user-guide/monitoring) records generic time-series metrics +about containers in a central database, and provides a UI for browsing that data. + +#### Cluster-level Logging + +A [Cluster-level logging](/docs/user-guide/logging/overview) mechanism is responsible for +saving container logs to a central log store with search/browsing interface. + +## Node components + +Node components run on every node, maintaining running pods and providing them +the Kubernetes runtime environment. + +### kubelet + +[kubelet](/docs/admin/kubelet) is the primary node agent. It: + +* Watches for pods that have been assigned to its node (either by apiserver + or via local configuration file) and: +* Mounts the pod's required volumes +* Downloads the pod's secrets +* Runs the pod's containers via docker (or, experimentally, rkt). +* Periodically executes any requested container liveness probes. +* Reports the status of the pod back to the rest of the system, by creating a + "mirror pod" if necessary. +* Reports the status of the node back to the rest of the system. + +### kube-proxy + +[kube-proxy](/docs/admin/kube-proxy) enables the Kubernetes service abstraction by maintaining +network rules on the host and performing connection forwarding. + +### docker + +`docker` is of course used for actually running containers. + +### rkt + +`rkt` is supported experimentally as an alternative to docker. + +### supervisord + +`supervisord` is a lightweight process babysitting system for keeping kubelet and docker +running. + +### fluentd + +`fluentd` is a daemon which helps provide [cluster-level logging](#cluster-level-logging).