From a0a51fb032c4c57a7501f8accfb7483033e914f8 Mon Sep 17 00:00:00 2001 From: Andrew Chen Date: Thu, 16 Mar 2017 11:31:57 -0700 Subject: [PATCH] Move Guide topic: Parallel Processing using Expansions (#2867) * Move Guide topic: Parallel Processing using Expansions * fix links to /docs/user-guide/jobs/ --- _data/tasks.yml | 4 + docs/admin/daemons.md | 2 +- docs/concepts/configuration/overview.md | 2 +- .../run-to-completion-finite-workloads.md | 2 +- .../overview/working-with-objects/labels.md | 2 +- docs/tasks/job/job.yaml | 18 ++ .../job/parallel-processing-expansion.md | 195 ++++++++++++++++++ docs/user-guide/cron-jobs.md | 6 +- docs/user-guide/jobs/expansions/index.md | 191 +---------------- docs/user-guide/jobs/work-queue-1/index.md | 4 +- docs/user-guide/jobs/work-queue-2/index.md | 4 +- docs/user-guide/pod-templates.md | 2 +- .../replication-controller/index.md | 2 +- index.html | 2 +- 14 files changed, 233 insertions(+), 203 deletions(-) create mode 100644 docs/tasks/job/job.yaml create mode 100644 docs/tasks/job/parallel-processing-expansion.md diff --git a/_data/tasks.yml b/_data/tasks.yml index c113a1fa87..ca4c1f2200 100644 --- a/_data/tasks.yml +++ b/_data/tasks.yml @@ -31,6 +31,10 @@ toc: section: - docs/tasks/run-application/rolling-update-replication-controller.md +- title: Running Jobs + section: + - docs/tasks/job/parallel-processing-expansion.md + - title: Accessing Applications in a Cluster section: - docs/tasks/access-application-cluster/port-forward-access-application-cluster.md diff --git a/docs/admin/daemons.md b/docs/admin/daemons.md index 4682a62b71..0e46ab6b21 100644 --- a/docs/admin/daemons.md +++ b/docs/admin/daemons.md @@ -51,7 +51,7 @@ A pod template in a DaemonSet must have a [`RestartPolicy`](/docs/user-guide/pod ### Pod Selector The `.spec.selector` field is a pod selector. It works the same as the `.spec.selector` of -a [Job](/docs/user-guide/jobs/) or other new resources. +a [Job](/docs/concepts/jobs/run-to-completion-finite-workloads/) or other new resources. The `spec.selector` is an object consisting of two fields: diff --git a/docs/concepts/configuration/overview.md b/docs/concepts/configuration/overview.md index 0ef55e7b5e..0d2192b03a 100644 --- a/docs/concepts/configuration/overview.md +++ b/docs/concepts/configuration/overview.md @@ -34,7 +34,7 @@ This document is meant to highlight and consolidate in one place configuration b Replication controllers are almost always preferable to creating pods, except for some explicit [`restartPolicy: Never`](/docs/user-guide/pod-states/#restartpolicy) scenarios. A - [Job](/docs/user-guide/jobs/) object (currently in Beta), may also be appropriate. + [Job](/docs/concepts/jobs/run-to-completion-finite-workloads/) object (currently in Beta), may also be appropriate. ## Services diff --git a/docs/concepts/jobs/run-to-completion-finite-workloads.md b/docs/concepts/jobs/run-to-completion-finite-workloads.md index fe17bb32e6..15121b7171 100644 --- a/docs/concepts/jobs/run-to-completion-finite-workloads.md +++ b/docs/concepts/jobs/run-to-completion-finite-workloads.md @@ -277,7 +277,7 @@ Here, `W` is the number of work items. | Pattern | `.spec.completions` | `.spec.parallelism` | | -------------------------------------------------------------------- |:-------------------:|:--------------------:| -| [Job Template Expansion](/docs/user-guide/jobs/expansions/) | 1 | should be 1 | +| [Job Template Expansion](/docs/tasks/job/parallel-processing-expansion/) | 1 | should be 1 | | [Queue with Pod Per Work Item](/docs/user-guide/jobs/work-queue-1/) | W | any | | [Queue with Variable Pod Count](/docs/user-guide/jobs/work-queue-2/) | 1 | any | | Single Job with Static Work Assignment | W | any | diff --git a/docs/concepts/overview/working-with-objects/labels.md b/docs/concepts/overview/working-with-objects/labels.md index af99f0c364..0b0adc8d0f 100644 --- a/docs/concepts/overview/working-with-objects/labels.md +++ b/docs/concepts/overview/working-with-objects/labels.md @@ -157,7 +157,7 @@ this selector (respectively in `json` or `yaml` format) is equivalent to `compon #### Resources that support set-based requirements -Newer resources, such as [`Job`](/docs/user-guide/jobs), [`Deployment`](/docs/user-guide/deployments/), [`Replica Set`](/docs/user-guide/replicasets/), and [`Daemon Set`](/docs/admin/daemons/), support _set-based_ requirements as well. +Newer resources, such as [`Job`](/docs/concepts/jobs/run-to-completion-finite-workloads/), [`Deployment`](/docs/user-guide/deployments/), [`Replica Set`](/docs/user-guide/replicasets/), and [`Daemon Set`](/docs/admin/daemons/), support _set-based_ requirements as well. ```yaml selector: diff --git a/docs/tasks/job/job.yaml b/docs/tasks/job/job.yaml new file mode 100644 index 0000000000..790025b38b --- /dev/null +++ b/docs/tasks/job/job.yaml @@ -0,0 +1,18 @@ +apiVersion: batch/v1 +kind: Job +metadata: + name: process-item-$ITEM + labels: + jobgroup: jobexample +spec: + template: + metadata: + name: jobexample + labels: + jobgroup: jobexample + spec: + containers: + - name: c + image: busybox + command: ["sh", "-c", "echo Processing item $ITEM && sleep 5"] + restartPolicy: Never diff --git a/docs/tasks/job/parallel-processing-expansion.md b/docs/tasks/job/parallel-processing-expansion.md new file mode 100644 index 0000000000..5b1b33a568 --- /dev/null +++ b/docs/tasks/job/parallel-processing-expansion.md @@ -0,0 +1,195 @@ +--- +title: Parallel Processing using Expansions +--- + +* TOC +{:toc} + +# Example: Multiple Job Objects from Template Expansion + +In this example, we will run multiple Kubernetes Jobs created from +a common template. You may want to be familiar with the basic, +non-parallel, use of [Jobs](/docs/concepts/jobs/run-to-completion-finite-workloads/) first. + +## Basic Template Expansion + +First, download the following template of a job to a file called `job.yaml.txt` + +{% include code.html language="yaml" file="job.yaml" ghlink="/docs/tasks/job/parallel-processing-expansion/job.yaml" %} + +Unlike a *pod template*, our *job template* is not a Kubernetes API type. It is just +a yaml representation of a Job object that has some placeholders that need to be filled +in before it can be used. The `$ITEM` syntax is not meaningful to Kubernetes. + +In this example, the only processing the container does is to `echo` a string and sleep for a bit. +In a real use case, the processing would be some substantial computation, such as rendering a frame +of a movie, or processing a range of rows in a database. The "$ITEM" parameter would specify for +example, the frame number or the row range. + +This Job and its Pod template have a label: `jobgroup=jobexample`. There is nothing special +to the system about this label. This label +makes it convenient to operate on all the jobs in this group at once. +We also put the same label on the pod template so that we can check on all Pods of these Jobs +with a single command. +After the job is created, the system will add more labels that distinguish one Job's pods +from another Job's pods. +Note that the label key `jobgroup` is not special to Kubernetes. You can pick your own label scheme. + +Next, expand the template into multiple files, one for each item to be processed. + +```shell +# Expand files into a temporary directory +mkdir ./jobs +for i in apple banana cherry +do + cat job.yaml.txt | sed "s/\$ITEM/$i/" > ./jobs/job-$i.yaml +done +``` + +Check if it worked: + +```shell +$ ls jobs/ +job-apple.yaml +job-banana.yaml +job-cherry.yaml +``` + +Here, we used `sed` to replace the string `$ITEM` with the loop variable. +You could use any type of template language (jinja2, erb) or write a program +to generate the Job objects. + +Next, create all the jobs with one kubectl command: + +```shell +$ kubectl create -f ./jobs +job "process-item-apple" created +job "process-item-banana" created +job "process-item-cherry" created +``` + +Now, check on the jobs: + +```shell +$ kubectl get jobs -l jobgroup=jobexample +JOB CONTAINER(S) IMAGE(S) SELECTOR SUCCESSFUL +process-item-apple c busybox app in (jobexample),item in (apple) 1 +process-item-banana c busybox app in (jobexample),item in (banana) 1 +process-item-cherry c busybox app in (jobexample),item in (cherry) 1 +``` + +Here we use the `-l` option to select all jobs that are part of this +group of jobs. (There might be other unrelated jobs in the system that we +do not care to see.) + +We can check on the pods as well using the same label selector: + +```shell +$ kubectl get pods -l jobgroup=jobexample --show-all +NAME READY STATUS RESTARTS AGE +process-item-apple-kixwv 0/1 Completed 0 4m +process-item-banana-wrsf7 0/1 Completed 0 4m +process-item-cherry-dnfu9 0/1 Completed 0 4m +``` + +There is not a single command to check on the output of all jobs at once, +but looping over all the pods is pretty easy: + +```shell +$ for p in $(kubectl get pods -l jobgroup=jobexample -o name) +do + kubectl logs $p +done +Processing item apple +Processing item banana +Processing item cherry +``` + +## Multiple Template Parameters + +In the first example, each instance of the template had one parameter, and that parameter was also +used as a label. However label keys are limited in [what characters they can +contain](/docs/user-guide/labels/#syntax-and-character-set). + +This slightly more complex example uses the jinja2 template language to generate our objects. +We will use a one-line python script to convert the template to a file. + +First, copy and paste the following template of a Job object, into a file called `job.yaml.jinja2`: + + +```liquid{% raw %} +{%- set params = [{ "name": "apple", "url": "http://www.orangepippin.com/apples", }, + { "name": "banana", "url": "https://en.wikipedia.org/wiki/Banana", }, + { "name": "raspberry", "url": "https://www.raspberrypi.org/" }] +%} +{%- for p in params %} +{%- set name = p["name"] %} +{%- set url = p["url"] %} +apiVersion: batch/v1 +kind: Job +metadata: + name: jobexample-{{ name }} + labels: + jobgroup: jobexample +spec: + template: + name: jobexample + labels: + jobgroup: jobexample + spec: + containers: + - name: c + image: busybox + command: ["sh", "-c", "echo Processing URL {{ url }} && sleep 5"] + restartPolicy: Never +--- +{%- endfor %} +{% endraw %} +``` + +The above template defines parameters for each job object using a list of +python dicts (lines 1-4). Then a for loop emits one job yaml object +for each set of parameters (remaining lines). +We take advantage of the fact that multiple yaml documents can be concatenated +with the `---` separator (second to last line). +.) We can pipe the output directly to kubectl to +create the objects. + +You will need the jinja2 package if you do not already have it: `pip install --user jinja2`. +Now, use this one-line python program to expand the template: + +```shell +alias render_template='python -c "from jinja2 import Template; import sys; print(Template(sys.stdin.read()).render());"' +``` + + + +The output can be saved to a file, like this: + +```shell +cat job.yaml.jinja2 | render_template > jobs.yaml +``` + +or sent directly to kubectl, like this: + +```shell +cat job.yaml.jinja2 | render_template | kubectl create -f - +``` + +## Alternatives + +If you have a large number of job objects, you may find that: + +- even using labels, managing so many Job objects is cumbersome. +- You exceed resource quota when creating all the Jobs at once, + and do not want to wait to create them incrementally. +- You need a way to easily scale the number of pods running + concurrently. One reason would be to avoid using too many + compute resources. Another would be to limit the number of + concurrent requests to a shared resource, such as a database, + used by all the pods in the job. +- very large numbers of jobs created at once overload the + Kubernetes apiserver, controller, or scheduler. + +In this case, you can consider one of the +other [job patterns](/docs/concepts/jobs/run-to-completion-finite-workloads/#job-patterns). diff --git a/docs/user-guide/cron-jobs.md b/docs/user-guide/cron-jobs.md index db3153688a..936c354cae 100644 --- a/docs/user-guide/cron-jobs.md +++ b/docs/user-guide/cron-jobs.md @@ -11,7 +11,7 @@ title: Cron Jobs ## What is a cron job? -A _Cron Job_ manages time based [Jobs](/docs/user-guide/jobs/), namely: +A _Cron Job_ manages time based [Jobs](/docs/concepts/jobs/run-to-completion-finite-workloads/), namely: * Once at a specified point in time * Repeatedly at a specified point in time @@ -159,8 +159,8 @@ string, e.g. `0 * * * *` or `@hourly`, as schedule time of its jobs to be create ### Job Template The `.spec.jobTemplate` is another required field of the `.spec`. It is a job template. It has exactly the same schema -as a [Job](/docs/user-guide/jobs), except it is nested and does not have an `apiVersion` or `kind`, see -[Writing a Job Spec](/docs/user-guide/jobs/#writing-a-job-spec). +as a [Job](/docs/concepts/jobs/run-to-completion-finite-workloads/), except it is nested and does not have an `apiVersion` or `kind`, see +[Writing a Job Spec](/docs/concepts/jobs/run-to-completion-finite-workloads/#writing-a-job-spec). ### Starting Deadline Seconds diff --git a/docs/user-guide/jobs/expansions/index.md b/docs/user-guide/jobs/expansions/index.md index 8d5cb87bd4..da9981b62d 100644 --- a/docs/user-guide/jobs/expansions/index.md +++ b/docs/user-guide/jobs/expansions/index.md @@ -2,194 +2,7 @@ title: Parallel Processing using Expansions --- -* TOC -{:toc} +{% include user-guide-content-moved.md %} -# Example: Multiple Job Objects from Template Expansion +[Parallel Processing Expansion](/docs/tasks/job/parallel-processing-expansion/) -In this example, we will run multiple Kubernetes Jobs created from -a common template. You may want to be familiar with the basic, -non-parallel, use of [Jobs](/docs/user-guide/jobs) first. - -## Basic Template Expansion - -First, download the following template of a job to a file called `job.yaml.txt` - -{% include code.html language="yaml" file="job.yaml.txt" ghlink="/docs/user-guide/job/expansions/job.yaml.txt" %} - -Unlike a *pod template*, our *job template* is not a Kubernetes API type. It is just -a yaml representation of a Job object that has some placeholders that need to be filled -in before it can be used. The `$ITEM` syntax is not meaningful to Kubernetes. - -In this example, the only processing the container does is to `echo` a string and sleep for a bit. -In a real use case, the processing would be some substantial computation, such as rendering a frame -of a movie, or processing a range of rows in a database. The "$ITEM" parameter would specify for -example, the frame number or the row range. - -This Job and its Pod template have a label: `jobgroup=jobexample`. There is nothing special -to the system about this label. This label -makes it convenient to operate on all the jobs in this group at once. -We also put the same label on the pod template so that we can check on all Pods of these Jobs -with a single command. -After the job is created, the system will add more labels that distinguish one Job's pods -from another Job's pods. -Note that the label key `jobgroup` is not special to Kubernetes. You can pick your own label scheme. - -Next, expand the template into multiple files, one for each item to be processed. - -```shell -# Expand files into a temporary directory -mkdir ./jobs -for i in apple banana cherry -do - cat job.yaml.txt | sed "s/\$ITEM/$i/" > ./jobs/job-$i.yaml -done -``` - -Check if it worked: - -```shell -$ ls jobs/ -job-apple.yaml -job-banana.yaml -job-cherry.yaml -``` - -Here, we used `sed` to replace the string `$ITEM` with the loop variable. -You could use any type of template language (jinja2, erb) or write a program -to generate the Job objects. - -Next, create all the jobs with one kubectl command: - -```shell -$ kubectl create -f ./jobs -job "process-item-apple" created -job "process-item-banana" created -job "process-item-cherry" created -``` - -Now, check on the jobs: - -```shell -$ kubectl get jobs -l jobgroup=jobexample -JOB CONTAINER(S) IMAGE(S) SELECTOR SUCCESSFUL -process-item-apple c busybox app in (jobexample),item in (apple) 1 -process-item-banana c busybox app in (jobexample),item in (banana) 1 -process-item-cherry c busybox app in (jobexample),item in (cherry) 1 -``` - -Here we use the `-l` option to select all jobs that are part of this -group of jobs. (There might be other unrelated jobs in the system that we -do not care to see.) - -We can check on the pods as well using the same label selector: - -```shell -$ kubectl get pods -l jobgroup=jobexample --show-all -NAME READY STATUS RESTARTS AGE -process-item-apple-kixwv 0/1 Completed 0 4m -process-item-banana-wrsf7 0/1 Completed 0 4m -process-item-cherry-dnfu9 0/1 Completed 0 4m -``` - -There is not a single command to check on the output of all jobs at once, -but looping over all the pods is pretty easy: - -```shell -$ for p in $(kubectl get pods -l jobgroup=jobexample -o name) -do - kubectl logs $p -done -Processing item apple -Processing item banana -Processing item cherry -``` - -## Multiple Template Parameters - -In the first example, each instance of the template had one parameter, and that parameter was also -used as a label. However label keys are limited in [what characters they can -contain](/docs/user-guide/labels/#syntax-and-character-set). - -This slightly more complex example uses the jinja2 template language to generate our objects. -We will use a one-line python script to convert the template to a file. - -First, copy and paste the following template of a Job object, into a file called `job.yaml.jinja2`: - - -```liquid{% raw %} -{%- set params = [{ "name": "apple", "url": "http://www.orangepippin.com/apples", }, - { "name": "banana", "url": "https://en.wikipedia.org/wiki/Banana", }, - { "name": "raspberry", "url": "https://www.raspberrypi.org/" }] -%} -{%- for p in params %} -{%- set name = p["name"] %} -{%- set url = p["url"] %} -apiVersion: batch/v1 -kind: Job -metadata: - name: jobexample-{{ name }} - labels: - jobgroup: jobexample -spec: - template: - name: jobexample - labels: - jobgroup: jobexample - spec: - containers: - - name: c - image: busybox - command: ["sh", "-c", "echo Processing URL {{ url }} && sleep 5"] - restartPolicy: Never ---- -{%- endfor %} -{% endraw %} -``` - -The above template defines parameters for each job object using a list of -python dicts (lines 1-4). Then a for loop emits one job yaml object -for each set of parameters (remaining lines). -We take advantage of the fact that multiple yaml documents can be concatenated -with the `---` separator (second to last line). -.) We can pipe the output directly to kubectl to -create the objects. - -You will need the jinja2 package if you do not already have it: `pip install --user jinja2`. -Now, use this one-line python program to expand the template: - -```shell -alias render_template='python -c "from jinja2 import Template; import sys; print(Template(sys.stdin.read()).render());"' -``` - - - -The output can be saved to a file, like this: - -```shell -cat job.yaml.jinja2 | render_template > jobs.yaml -``` - -or sent directly to kubectl, like this: - -```shell -cat job.yaml.jinja2 | render_template | kubectl create -f - -``` - -## Alternatives - -If you have a large number of job objects, you may find that: - -- even using labels, managing so many Job objects is cumbersome. -- You exceed resource quota when creating all the Jobs at once, - and do not want to wait to create them incrementally. -- You need a way to easily scale the number of pods running - concurrently. One reason would be to avoid using too many - compute resources. Another would be to limit the number of - concurrent requests to a shared resource, such as a database, - used by all the pods in the job. -- very large numbers of jobs created at once overload the - Kubernetes apiserver, controller, or scheduler. - -In this case, you can consider one of the -other [job patterns](/docs/user-guide/jobs/#job-patterns). diff --git a/docs/user-guide/jobs/work-queue-1/index.md b/docs/user-guide/jobs/work-queue-1/index.md index f926f4211f..8b76ff1185 100644 --- a/docs/user-guide/jobs/work-queue-1/index.md +++ b/docs/user-guide/jobs/work-queue-1/index.md @@ -9,7 +9,7 @@ title: Coarse Parallel Processing using a Work Queue In this example, we will run a Kubernetes Job with multiple parallel worker processes. You may want to be familiar with the basic, -non-parallel, use of [Job](/docs/user-guide/jobs) first. +non-parallel, use of [Job](/docs/concepts/jobs/run-to-completion-finite-workloads/) first. In this example, as each pod is created, it picks up one unit of work from a task queue, completes it, deletes it from the queue, and exits. @@ -255,7 +255,7 @@ do not need to modify your "worker" program to be aware that there is a work que It does require that you run a message queue service. If running a queue service is inconvenient, you may -want to consider one of the other [job patterns](/docs/user-guide/jobs/#job-patterns). +want to consider one of the other [job patterns](/docs/concepts/jobs/run-to-completion-finite-workloads/#job-patterns). This approach creates a pod for every work item. If your work items only take a few seconds, though, creating a Pod for every work item may add a lot of overhead. Consider another diff --git a/docs/user-guide/jobs/work-queue-2/index.md b/docs/user-guide/jobs/work-queue-2/index.md index 4fd806d392..8da4fc8fd3 100644 --- a/docs/user-guide/jobs/work-queue-2/index.md +++ b/docs/user-guide/jobs/work-queue-2/index.md @@ -9,7 +9,7 @@ title: Fine Parallel Processing using a Work Queue In this example, we will run a Kubernetes Job with multiple parallel worker processes. You may want to be familiar with the basic, -non-parallel, use of [Job](/docs/user-guide/jobs) first. +non-parallel, use of [Job](/docs/concepts/jobs/run-to-completion-finite-workloads/) first. In this example, as each pod is created, it picks up one unit of work from a task queue, completes it, deletes it from the queue, and exits. @@ -205,7 +205,7 @@ As you can see, one of our pods worked on several work units. ## Alternatives If running a queue service or modifying your containers to use a work queue is inconvenient, you may -want to consider one of the other [job patterns](/docs/user-guide/jobs/#job-patterns). +want to consider one of the other [job patterns](/docs/concepts/jobs/run-to-completion-finite-workloads/#job-patterns). If you have a continuous stream of background processing work to run, then consider running your background workers with a `replicationController` instead, diff --git a/docs/user-guide/pod-templates.md b/docs/user-guide/pod-templates.md index a10c60c62b..27e45d82eb 100644 --- a/docs/user-guide/pod-templates.md +++ b/docs/user-guide/pod-templates.md @@ -5,7 +5,7 @@ title: Pod Templates --- Pod templates are [pod](/docs/user-guide/pods/) specifications which are included in other objects, such as -[Replication Controllers](/docs/user-guide/replication-controller/), [Jobs](/docs/user-guide/jobs/), and +[Replication Controllers](/docs/user-guide/replication-controller/), [Jobs](/docs/concepts/jobs/run-to-completion-finite-workloads/), and [DaemonSets](/docs/admin/daemons/). Controllers use Pod Templates to make actual pods. Rather than specifying the current desired state of all replicas, pod templates are like cookie cutters. Once a cookie has been cut, the cookie has no relationship to the cutter. There is no quantum entanglement. Subsequent changes to the template or even switching to a new template has no direct effect on the pods already created. Similarly, pods created by a replication controller may subsequently be updated directly. This is in deliberate contrast to pods, which do specify the current desired state of all containers belonging to the pod. This approach radically simplifies system semantics and increases the flexibility of the primitive. diff --git a/docs/user-guide/replication-controller/index.md b/docs/user-guide/replication-controller/index.md index 824ab21841..632e73f4c7 100644 --- a/docs/user-guide/replication-controller/index.md +++ b/docs/user-guide/replication-controller/index.md @@ -249,7 +249,7 @@ Unlike in the case where a user directly created pods, a ReplicationController r ### Job -Use a [`Job`](/docs/user-guide/jobs/) instead of a ReplicationController for pods that are expected to terminate on their own +Use a [`Job`](/docs/concepts/jobs/run-to-completion-finite-workloads/) instead of a ReplicationController for pods that are expected to terminate on their own (i.e. batch jobs). ### DaemonSet diff --git a/index.html b/index.html index 30cb9264b5..37a426ebdd 100644 --- a/index.html +++ b/index.html @@ -115,7 +115,7 @@ cid: home Gluster, Ceph, Cinder, or Flocker.

-

Batch execution

+

Batch execution

In addition to services, Kubernetes can manage your batch and CI workloads, replacing containers that fail, if desired.