From 877c6c224e6d97bbae12b2cb255902c892b08729 Mon Sep 17 00:00:00 2001 From: Solly Ross Date: Mon, 6 Mar 2017 18:18:19 -0500 Subject: [PATCH] Update HPA documentation to cover HPA v2 This updates the horizontal pod autoscaling documention to cover the new autosclaing/v2alpha1 API version. It also notes the removal of the old alpha annotations for autoscaling on custom metrics, and reccomends against using the alpha collection method. --- .../horizontal-pod-autoscaling/index.md | 125 ++++++++------- .../horizontal-pod-autoscaling/walkthrough.md | 150 +++++++++++++++++- 2 files changed, 214 insertions(+), 61 deletions(-) diff --git a/docs/user-guide/horizontal-pod-autoscaling/index.md b/docs/user-guide/horizontal-pod-autoscaling/index.md index 44ad3440e2..6bfc634701 100644 --- a/docs/user-guide/horizontal-pod-autoscaling/index.md +++ b/docs/user-guide/horizontal-pod-autoscaling/index.md @@ -2,6 +2,7 @@ assignees: - fgrzadkowski - jszczepkowski +- directxman12 title: Horizontal Pod Autoscaling --- @@ -22,19 +23,43 @@ to match the observed average CPU utilization to the target specified by user. ![Horizontal Pod Autoscaler diagram](/images/docs/horizontal-pod-autoscaler.svg) -The autoscaler is implemented as a control loop. -It periodically queries CPU utilization for the pods it targets. -(The period of the autoscaler is controlled by `--horizontal-pod-autoscaler-sync-period` flag of controller manager. -The default value is 30 seconds). -Then, it compares the arithmetic mean of the pod's CPU utilization with the target and adjust the number of replicas if needed. +The Horizontal Pod Autoscaler is implemented as a control loop, with a period controlled +by the controller manager's `--horizontal-pod-autoscaler-sync-period` flag (with a default +value of 30 seconds). -CPU utilization is the recent CPU usage of a pod divided by the sum of CPU requested by the pod's containers. -Please note that if some of the pod's containers do not have CPU request set, -CPU utilization for the pod will not be defined and the autoscaler will not take any action. -Further details of the autoscaling algorithm are given [here](https://github.com/kubernetes/kubernetes/blob/{{page.githubbranch}}/docs/design/horizontal-pod-autoscaler.md#autoscaling-algorithm). +During each period, the controller manager queries the resource utiliuzation against the +metrics specified in each HorizontalPodAutoscaler definition. The controller manager +obtains the metrics from either the resource metrics API (for per-pod resource metrics), +or the custom metrics API (for all ofther metrics). -The autoscaler uses heapster to collect CPU utilization. -Therefore, it is required to deploy heapster monitoring in your cluster for autoscaling to work. +* For per-pod resource metrics (like CPU), the controller fetches the metrics + from the resource metrics API for each pod targeted by the HorizontalPodAutoscaler. + Then, if a target utilization value is set, the controller calculates the utilization + value as a percentage of the equivalent resource request on the containers in + each pod. If a target raw value is set, the raw metric values are used directly. + the controller then takes the mean of the utilization or the raw value (depending on the type + of target specified) across all targeted pods, and produces a ratio used to scale + the number of desired replicas. + + Please note that if some of the pod's containers do not have the relevant resource request set, + CPU utilization for the pod will not be defined and the autoscaler will not take any action + for that metric. See the [autoscaling algorithm design document](https://github.com/kubernetes/kubernetes/blob/{{page.githubbranch}}/docs/design/horizontal-pod-autoscaler.md#autoscaling-algorithm) for further + details about how the autoscaling algorithm works. + +* For per-pod custom metrics, the controller functions similarly to per-pod resource metrics, + except that it works with raw values, not utilization values. + +* For object metrics, a single metric is fetched (which describes the object + in question), and compared to the target value, to produce a ratio as above. + +The HorizontalPodAutoscaler controller can fetch metrics in two different ways: direct Heapster +access, and REST client access. + +When using direct Heapster access, the HorizontalPodAutoscaler queries Heapster directly +through the API server's service proxy subresource. Heapster needs to be deployed on the +cluster and running in the kube-system namespace. + +See [Support for custom metrics](#prerequisites) for more details on REST client access. The autoscaler accesses corresponding replication controller, deployment or replica set by scale sub-resource. Scale is an interface which allows to dynamically set the number of replicas and to learn the current state of them. @@ -43,12 +68,13 @@ More details on scale sub-resource can be found [here](https://github.com/kubern ## API Object -Horizontal Pod Autoscaler is a top-level resource in the Kubernetes REST API. -In Kubernetes 1.2 HPA was graduated from beta to stable (more details about [api versioning](/docs/api/#api-versioning)) with compatibility between versions. -The stable version is available in the `autoscaling/v1` api group whereas the beta vesion is available in the `extensions/v1beta1` api group as before. -The transition plan is to deprecate beta version of HPA in Kubernetes 1.3, and get it rid off completely in Kubernetes 1.4. +The Horizontal Pod Autoscaler is an API resource in the Kubernetes `autoscaling` API group. +The current stable version, which only includes support for CPU autoscaling, +can be found in the `autoscaling/v1` API version. -**Warning!** Please have in mind that all Kubernetes components still use HPA in `extensions/v1beta1` in Kubernetes 1.2. +The alpha version, which includes support for scaling on memory and custom metrics, +can be found in `autoscaling/v2alpha1`. The new fields introduced in `autoscaling/v2alpha1` +are preserved as annotations when working with `autoscaling/v1`. More details about the API object can be found at [HorizontalPodAutoscaler Object](https://github.com/kubernetes/kubernetes/blob/{{page.githubbranch}}/docs/design/horizontal-pod-autoscaler.md#horizontalpodautoscaler-object). @@ -79,55 +105,36 @@ i.e. you cannot bind a Horizontal Pod Autoscaler to a replication controller and The reason this doesn't work is that when rolling update creates a new replication controller, the Horizontal Pod Autoscaler will not be bound to the new replication controller. +## Support for multiple metrics + +Kubernetes 1.6 adds support for scaling based on multiple metrics. You can use the `autoscaling/v2alpha1` API +version to specify multiple metrics for the Horizontal Pod Autoscaler to scale on. Then, the Horizontal Pod +Autoscaler controller will evaluate each metric, and propose a new scale based on that metric. The largest of the +proposed scales will be used as the new scale. + ## Support for custom metrics -Kubernetes 1.2 adds alpha support for scaling based on application-specific metrics like QPS (queries per second) or average request latency. +**Note**: Kubernetes 1.2 added alpha support for scaling based on application-specific metrics using special annotations. +Support for these annotations was removed in Kubernetes 1.6 in favor of the `autoscaling/v2alpha1` API. While the old method for collecting +custom metrics is still available, these metrics will not be available for use by the Horizontal Pod Autoscaler, and the former +annotations for specifying which custom metrics to scale on are no longer honored by the Horizontal Pod Autoscaler controller. + +Kubernetes 1.6 adds support for making use of custom metrics in the Horizontal Pod Autoscaler. +You can add custom metrics for the Horizontal Pod Autoscaler to use in the `autoscaling/v2alpha1` API. +Kubernetes then queries the new custom metrics API to fetch the values of the appropriate custom metrics. ### Prerequisites -The cluster has to be started with `ENABLE_CUSTOM_METRICS` environment variable set to `true`. - -### Pod configuration - -The pods to be scaled must have cAdvisor-specific custom (aka application) metrics endpoint configured. The configuration format is described [here](https://github.com/google/cadvisor/blob/master/docs/application_metrics.md). Kubernetes expects the configuration to - be placed in `definition.json` mounted via a [configMap](/docs/user-guide/configmap/) in `/etc/custom-metrics`. A sample config map may look like this: - -```yaml -apiVersion: v1 -kind: ConfigMap -metadata: - name: cm-config -data: - definition.json: "{\"endpoint\" : \"http://localhost:8080/metrics\"}" -``` - -**Warning** -Due to the way cAdvisor currently works `localhost` refers to the node itself, not to the running pod. Thus the appropriate container in the pod must ask for a node port. Example: - -```yaml - ports: - - hostPort: 8080 - containerPort: 8080 -``` - -### Specifying target - -HPA for custom metrics is configured via an annotation. The value in the annotation is interpreted as a target metric value averaged over -all running pods. Example: - -```yaml - annotations: - alpha/target.custom-metrics.podautoscaler.kubernetes.io: '{"items":[{"name":"qps", "value": "10"}]}' -``` - -In this case, if there are four pods running and each pod reports a QPS metric of 15 or higher, horizontal pod autoscaling will start two additional pods (for a total of six pods running). - -If you specify multiple metrics in your annotation or if you set a target CPU utilization, horizontal pod autoscaling will scale to according to the metric that requires the highest number of replicas. - -If you do not specify a target for CPU utilization, Kubernetes defaults to an 80% utilization threshold for horizontal pod autoscaling. - -If you want to ensure that horizontal pod autoscaling calculates the number of required replicas based only on custom metrics, you should set the CPU utilization target to a very large value (such as 100000%). As this level of CPU utilization isn't possible, horizontal pod autoscaling will calculate based only on the custom metrics (and min/max limits). +In order to use custom metrics in the Horizontal Pod Autoscaler, you must deploy your cluster with the +`--horizontal-pod-autoscaler-use-rest-clients` flag on the controller manager set to true. You must then configure +your controller manager to speak to the API server through the API server aggregator, by setting the controller +manager's target API server to the API server aggregator (using the `--apiserver` flag). The resource metrics API and +custom metrics API must also be registered with the API server aggregator, and must be served by API servers running +on the cluster. +You can use Heapster's implementation of the resource metrics API by running Heapster with the`--api-server` flag set +to true. A separate component must provide the custom metrics API (more information on the custom metrics API is +available at [the k8s.io/metrics repository](https://github.com.com/kubernetes/metrics)). ## Further reading diff --git a/docs/user-guide/horizontal-pod-autoscaling/walkthrough.md b/docs/user-guide/horizontal-pod-autoscaling/walkthrough.md index 616061e930..1532c47a76 100644 --- a/docs/user-guide/horizontal-pod-autoscaling/walkthrough.md +++ b/docs/user-guide/horizontal-pod-autoscaling/walkthrough.md @@ -3,6 +3,7 @@ assignees: - fgrzadkowski - jszczepkowski - justinsb +- directxman12 title: Horizontal Pod Autoscaling --- @@ -10,7 +11,7 @@ Horizontal Pod Autoscaling automatically scales the number of pods in a replication controller, deployment or replica set based on observed CPU utilization (or, with alpha support, on some other, application-provided metrics). -This document walks you through an example of enabling Horizontal Pod Autoscaling for the php-apache server. For more information on how Horizontal Pod Autoscaling behaves, see the [Horizontal Pod Autoscaling glossary entry](/docs/user-guide/horizontal-pod-autoscaling/). +This document walks you through an example of enabling Horizontal Pod Autoscaling for the php-apache server. For more information on how Horizontal Pod Autoscaling behaves, see the [Horizontal Pod Autoscaling user guide](/docs/user-guide/horizontal-pod-autoscaling/). ## Prerequisites @@ -20,6 +21,11 @@ as Horizontal Pod Autoscaler uses it to collect metrics (if you followed [getting started on GCE guide](/docs/getting-started-guides/gce), heapster monitoring will be turned-on by default). +To specify multiple resource metrics for a Horizontal Pod Autoscaler, you must have a Kubernetes cluster +and kubectl at version 1.6 or later. Furthermore, in order to make use of custom metrics, your cluster +must be able to communicate with the API server providing the custom metrics API. +See the [Horizontal Pod Autoscaling user guide](/docs/user-guide/horizontal-pod-autoscaling/#support-for-custom-metrics) for more details. + ## Step One: Run & expose php-apache server To demonstrate Horizontal Pod Autoscaler we will use a custom docker image based on the php-apache image. @@ -95,7 +101,7 @@ php-apache 7 7 7 7 19m **Note** Sometimes it may take a few minutes to stabilize the number of replicas. Since the amount of load is not controlled in any way it may happen that the final number of replicas will -differ from this example. +differ from this example. ## Step Four: Stop load @@ -120,6 +126,146 @@ Here CPU utilization dropped to 0, and so HPA autoscaled the number of replicas **Note** autoscaling the replicas may take a few minutes. +## Autoscaling on multiple metrics and custom metrics + +You can introduce additional metrics to use when autoscaling the `php-apache` Deployment +by making use of the `autoscaling/v2alpha1` API version. + +First, get the YAML of your HorizontalPodAutoscaler in the `autoscaling/v2alpha1` form: + +```shell +$ kubectl get hpa.autoscaling.v2alpha1 -o yaml > /tmp/hpa-v2.yaml +``` + +Open the `/tmp/hpa-v2.yaml` file in an editor, and you should see YAML which looks like this: + +```yaml +apiVersion: autoscaling/v2alpha1 +kind: HorizontalPodAutoscaler +metadata: + name: php-apache + namespace: default +spec: + scaleTargetRef: + apiVersion: extensions/v1beta1 + kind: Deployment + name: php-apache + minReplicas: 1 + maxReplicas: 10 + metrics: + - type: Resource + resource: + name: cpu + targetAverageUtilization: 50 +status: + observedGeneration: 1 + lastScaleTime: + currentReplicas: 1 + desiredReplicas: 1 + currentMetrics: + - type: Resource + resource: + name: cpu + currentAverageUtilization: 0 + currentAverageValue: 0 +``` + +Notice that the `targetCPUUtilizationPercentage` field has been replaced with an array called `metrics`. +The CPU utilization metric is a *resource metric*, since it is represented as a percentage of a resource +specified on pod containers. Notice that you can specify other resource metrics besides CPU. By default, +the only other supported resource metric is memory. These resources do not change names from cluster +to cluster, and should always be available, as long as Heapster is deployed. + +You can also specify resource metrics in terms of direct values, instead of as percentages of the +requested value. To do so, use the `targetAverageValue` field insted of the `targetAverageUtilization` +field. + +There are two other types of metrics, both of which are considered *custom metrics*: pod metrics and +object metrics. These metrics may have names which are cluster specific, and require a more +advanced cluster monitoring setup. + +The first of these alternative metric types is *pod metrics*. These metrics describe pods, and +are averaged together across pods and compared with a target value to determine the replica count. +They work much like resource metrics, except that they *only* have the `targetAverageValue` field. + +Pod metrics are specified using a metric block like this: +```yaml +type: Pods +pods: + metricName: packets-per-second + targetAverageValue: 1k +``` + +The second alternative metric type is *object metrics*. These metrics describe a different +object in the same namespace, instead of describing pods. Note that the metrics are not +fetched from the object -- they simply describe it. Object metrics do not involve averaging, +and look like this: + +```yaml +type: Object +object: + metricName: requests-per-second + target: + apiVersion: extensions/v1beta1 + kind: Ingress + name: main-route + targetValue: 2k +``` + +If you provide multiple such metric blocks, the HorizontalPodAutoscaler will consider each metric in turn. +The HorizontalPodAutoscaler will calculate proposed replica counts for each metric, and then choose the +one with the highest replica count. + +For example, if you had your monitoring system collecting metrics about network traffic, +you could update the definition above using `kubectl edit` to look like this: + +```yaml +apiVersion: autoscaling/v2alpha1 +kind: HorizontalPodAutoscaler +metadata: + name: php-apache + namespace: default +spec: + scaleTargetRef: + apiVersion: extensions/v1beta1 + kind: Deployment + name: php-apache + minReplicas: 1 + maxReplicas: 10 + metrics: + - type: Resource + resource: + name: cpu + targetAverageUtilization: 50 + - type: Pods + pods: + metricName: packets-per-second + targetAverageValue: 1k + - type: Object + object: + metricName: requests-per-second + target: + apiVersion: extensions/v1beta1 + kind: Ingress + name: main-route + targetValue: 10k +status: + observedGeneration: 1 + lastScaleTime: + currentReplicas: 1 + desiredReplicas: 1 + currentMetrics: + - type: Resource + resource: + name: cpu + currentAverageUtilization: 0 + currentAverageValue: 0 +``` + +Then, your HorizontalPodAutoscaler would attempt to ensure that each pod was consuming roughly +50% of its requested CPU, serving 1000 packets per second, and that all pods behind the main-route +Ingress were serving a total of 10000 requests per second. + ## Appendix: Other possible scenarios ### Creating the autoscaler from a .yaml file