Skip to content

How to respond to GKE upgrade notifications

This guide shows you how to respond to the GitLab issues that are raised when an upgrade is available for one of our Google Kubernetes Engine (GKE) clusters.

GKE clusters on a release channel are upgraded automatically by Google, so no action is required. However, an automatic upgrade can happen at any time within the cluster's maintenance window, which may be in the middle of the night. These issues give the team a chance to test the upgrade and apply it at a time of their choosing, so that any problems can be dealt with there and then.

Understand where the issue comes from

The issues are created by the GCP Notification to GitLab Issue Terraform module. GKE publishes cluster notifications to a Pub/Sub topic and the module's Cloud Function turns UpgradeAvailableEvent notifications into GitLab issues in the configured project.

Some things to be aware of:

  • An issue is only raised when the available version is the one that GKE will automatically upgrade the cluster to on its release channel. Intermediate versions are ignored.
  • There is a single open issue per cluster. By default it's titled chore(<project>/<cluster>): GKE Cluster Upgrade Available.
  • Control plane (master) and node pool upgrades are both reported against the same issue. The first notification forms the issue description, and any further notifications are added as notes while the issue remains open.
  • Once the issue is closed, any subsequent notification will open a new issue.

The issue should be treated like any other issue and worked on using our standard issue lifecycle workflow.

Decide how to handle the upgrade

How the upgrade is handled is down to each team, as the impact of an upgrade varies between services. Our recommended approach is:

  1. If a non-production cluster is available, upgrade it manually first and check that your workloads continue to work as expected.
  2. Upgrade the production cluster manually at a time that suits the team. That way any issues can be resolved immediately, rather than being discovered after GKE has upgraded the cluster automatically outside working hours.

If the team is comfortable with GKE upgrading the cluster automatically, it's fine to note this on the issue and close it without taking any further action.

Prepare for the upgrade

Before upgrading a cluster you should:

  • Check the GKE release notes for the new version, and the upstream Kubernetes changelog if the upgrade moves to a new minor version (for example 1.32 to 1.33).
  • Check the cluster for use of deprecated Kubernetes APIs which are removed in the new version. The GKE console highlights these on the cluster's details page, and they are explained in the feature and API deprecations documentation.
  • Consider what is running on the cluster and how it will be affected. For example, a zonal cluster's control plane is unavailable while it's upgraded, and nodes are drained and replaced when a node pool is upgraded. See what happens during an upgrade below.
  • Let the users of the service know, if the upgrade is likely to cause any disruption.

Note

If your Terraform configuration pins the cluster or node pool version, update the version in Terraform rather than upgrading via the console, otherwise the configuration will no longer reflect the deployed cluster.

Upgrade the cluster

The control plane must be upgraded before its node pools, since nodes can't run a newer version of Kubernetes than the control plane. Follow the steps below, which are covered in more detail in the GKE documentation on manually upgrading a cluster.

  1. Open the cluster in the Google Cloud console. The issue description contains a link to the cluster.
  2. Upgrade the control plane to the version given in the issue:
    1. On the cluster's Details tab, find Version in the Cluster basics section and click Upgrade available.
    2. Select the version given in the issue and click Save changes.
    3. Wait for the upgrade to complete. This can take some time.
  3. Upgrade each node pool, one at a time:
    1. On the cluster's Nodes tab, click the name of the node pool.
    2. Click Upgrade next to the node pool's version.
    3. Select the same version as the control plane and click Change.
    4. Wait for the upgrade to complete before moving on to the next node pool.
  4. Check that your workloads are healthy.
  5. Add a note to the issue with the outcome of the upgrade, and close it.
Upgrading using gcloud

The same upgrades can be performed using the gcloud CLI. Upgrade the control plane first:

gcloud container clusters upgrade CLUSTER_NAME --master \
  --cluster-version=VERSION --location=LOCATION --project=PROJECT_ID

Then upgrade each node pool:

gcloud container clusters upgrade CLUSTER_NAME --node-pool=NODE_POOL_NAME \
  --cluster-version=VERSION --location=LOCATION --project=PROJECT_ID

What happens during an upgrade

  • Control plane: for a zonal cluster the Kubernetes API is unavailable while the control plane is upgraded, so you won't be able to deploy or change workloads until it has finished. Workloads already running on the nodes continue to run. Regional clusters have multiple control plane replicas which are upgraded one at a time, so the API remains available.
  • Node pools: by default GKE uses surge upgrades, creating a new node before cordoning and draining an old one. Pods running on a drained node are evicted and rescheduled, so long-running work (for example CI jobs on a GitLab runner cluster) may be interrupted. Pod disruption budgets are respected for a limited time while nodes are drained.

See also