Understanding the Platforms Team's IaC factory pattern¶
When the Platforms Team refers to a factory, we mean a single infrastructure as code definition that is deployed many times over, with each deployment supplying its own configuration and keeping its own state. Every deployment is independent and can be created, changed or destroyed without touching any of the others.
This lets one shared definition serve dozens or hundreds of deployments while keeping a predictable blast radius for every operation. Isolating state per deployment avoids cross-coupling between products and services, and makes the effect of any change straightforward to reason about.
Our factories are built from two tools working together. Terraform declares the infrastructure, and Terragrunt decides which deployment to run, what inputs it receives, and where its state lives. This page explains that split, the repository layout our factories use today alongside the one we are moving to, and how the pipelines that run our factories are generated.
When to use the factory pattern¶
The factory pattern is not intended for routine product or service deployments. For those use cases, the DevOps division already has a well-established deployment approach using standard templates and patterns.
By contrast, the factory pattern suits platform-level or foundational infrastructure that must be deployed repeatedly at scale, often across many teams or product groups, while remaining technically owned by a central team. In these scenarios the infrastructure has a consistent shape, requires strong guardrails, and benefits from being managed as one reusable definition rather than duplicated across repositories.
The pattern is particularly effective when a single team needs to provide a standardised set of resources with centrally enforced practices such as security controls, auditing, compliance requirements and operational guardrails. The factory allows that team to evolve the implementation over time, while consumers simply provide configuration for their own deployment.
In summary, the factory pattern is most appropriate when:
- A single team owns and maintains the infrastructure definition
- The same infrastructure pattern must be deployed many times
- Each deployment requires an independent lifecycle and isolated state
- Strong central standards and guardrails must be enforced
- Consumers interact primarily through configuration, not implementation
For team-specific, bespoke or one-off infrastructure deployments, the standard product deployment patterns should continue to be used instead.
How Terragrunt and Terraform work together¶
Terraform has the primitives a factory needs. Partial backend
configuration
leaves part of the state location to be supplied at init time, and variable definition files
supply per-deployment inputs at plan time. What Terraform does not give us is anything to
orchestrate those primitives across many deployments. Choosing which deployments to act on, and
passing each one the right arguments, is left to the caller.
We used to do that ourselves, with a run-*.sh script in each factory repository. That worked, but
it did not scale. Each script was copied from the last and then diverged, so every improvement had
to be made several times over.
Terragrunt is built to orchestrate Terraform at scale, which is why we adopted it. The division of labour is worth stating plainly. Terraform declares what the infrastructure is. Terragrunt decides which instance of it to act on, what inputs that instance gets, and which state file it writes to.
Units¶
A unit is Terragrunt's name for one deployable instance. In our factories a unit is a directory
containing a terragrunt.hcl file. The presence of that file is what makes the directory
addressable by Terragrunt commands, so adding a deployment means adding a directory and running
nothing else to register it.
Shared configuration lives in a single root.hcl at the repository root, which every unit includes.
# product-vars/example/terragrunt.hcl
include "root" {
path = find_in_parent_folders("root.hcl")
}
State isolation¶
root.hcl holds a remote_state block that derives each unit's state location from its path.
Terragrunt passes these values to Terraform as -backend-config arguments during init, so no unit
has to know or restate where its own state lives.
# root.hcl
remote_state {
backend = "gcs"
config = {
bucket = local.gcs_bucket
prefix = "terraform/${basename(path_relative_to_include())}"
}
}
Because the prefix is derived from the unit directory, every deployment gets a fully isolated state file by construction. This is the foundational requirement of the factory pattern, and it is now impossible to get wrong by forgetting an argument.
The static backend block still present in each factory's Terraform is deliberately incomplete. It
fixes the values that are the same for the whole factory, such as the bucket, and leaves the
per-deployment prefix to Terragrunt. Terraform does not allow variables in a terraform block, so
this is the only way to express it.
Repository layouts¶
Two layouts matter here. One is what our factories use today, and one is where we are heading. They differ in where the Terraform lives and how a deployment reaches it.
A single root module¶
gcp-product-factory, gitlab-project-factory and ais-api-factory keep one Terraform root module
at the repository root. Every unit instantiates that same root module with a different variables
file.
.
├── modules/
│ └── workspace/ # local child modules, called by relative path
├── product-vars/
│ ├── example/
│ │ ├── example.tfvars # this deployment's inputs
│ │ └── terragrunt.hcl # marks the directory as a unit
│ └── another-product/
│ ├── another-product.tfvars
│ └── terragrunt.hcl
├── main.tf # the root module
├── variables.tf
├── versions.tf
├── root.hcl
├── .pipeline-helper.yml
└── .gitlab-ci.yml
root.hcl points Terragrunt at the repository root and injects the variables file, which it finds
by convention from the unit's directory name.
terraform {
source = "${get_repo_root()}//"
extra_arguments "var_file" {
commands = get_terraform_commands_that_need_vars()
required_var_files = ["${get_terragrunt_dir()}/${basename(get_terragrunt_dir())}.tfvars"]
}
}
Because root.hcl resolves everything, each unit's terragrunt.hcl is the same three-line include.
In gcp-product-factory all 103 of them are byte for byte identical. That uniformity is the
clearest sign of what this layout is, which is Terragrunt retrofitted onto a repository shaped for
Terraform alone.
A catalog of units and modules¶
The layout we are standardising on separates the two concerns into their own directories. Terraform
modules live in catalog/modules/, reusable Terragrunt unit definitions live in catalog/units/,
and each deployment declares which units it needs and supplies their values.
.
├── catalog/
│ ├── units/
│ │ ├── subscription/terragrunt.hcl # maps values onto the module's inputs
│ │ └── meta/terragrunt.hcl
│ └── modules/
│ ├── subscription/ # plain Terraform, no Terragrunt
│ └── meta/
├── live/
│ ├── example/
│ └── another-product/
├── root.hcl
└── .gitlab-ci.yml
This gives three layers rather than two. The deployment says which units it wants and supplies their values, the catalog unit maps those values onto module inputs and carries the defaults, and the module declares resources and nothing else.
# catalog/units/subscription/terragrunt.hcl
terraform {
source = "${get_repo_root()}//catalog/modules/subscription"
}
inputs = {
environment = local.values.environment
product_name = local.values.product_name
location = try(local.values.location, "uksouth")
}
Note where the defaults sit. In the single root module layout they are in variables.tf, mixed in
with the module's own interface. Here they are in the unit, which keeps the Terraform module
reusable and makes the factory's opinions visible in one place.
This is not a local invention, and the directory names above are not arbitrary. They are what Terragrunt's own documentation recommends, and its Terralith to Terragrunt guide builds up to exactly this layout. Following the upstream convention means a reader who knows Terragrunt already knows their way around our repositories, and we can point at upstream documentation rather than writing our own.
Direction of travel¶
The single root module layout is a hangover from the pre-Terragrunt era. It was the right shape when
a shell script drove terraform directly, because there was only ever one root module to point at.
Carrying it forward constrains us in two ways. A repository can hold only one root module, so a
deployment cannot be composed from several different units, and everything that deployment manages
shares a single state file.
The catalog layout removes both constraints, and it keeps the Terraform modules free of factory-specific wiring so they can be tested and reused on their own terms.
New factories should start from the catalog layout, using the catalog/units and catalog/modules
names above. We expect to migrate the existing factories over time, though that is a larger piece of
work than it looks, because moving a deployment between layouts means moving its state too.
Deployment configuration¶
Each deployment supplies its own configuration. In the single root module factories this is a
.tfvars file named after its directory, which Terragrunt locates by convention.
product-vars/example/example.tfvars
In the catalog layout the values are declared alongside the units the deployment uses, and reach
the Terraform module through the catalog unit's inputs block.
HCL versus YAML¶
We standardise on native HCL variable definition files for factory inputs.
Alternative formats have been considered. Terraform supports JSON natively, but its strictness and lack of comments make it poorly suited to configuration written by hand. YAML is more forgiving, but Terraform does not read it natively and it would need additional tooling to convert at runtime.
Native HCL gives us schema validation through variable blocks, consistent tooling, and no
conversion layer to maintain.
Running a factory locally¶
Every factory wraps Terragrunt in a set of poe tasks, so nobody has
to remember the full invocation or the environment variables the factory expects. Some factories,
including gcp-product-factory and gitlab-project-factory, run Terragrunt through a
docker-compose.yaml service as well, which pins the Terragrunt version across the team.
One decision is worth knowing about before you run anything. Terragrunt adds -auto-approve to
destructive commands by default, which removes the confirmation prompt. We set
TG_NO_AUTO_APPROVE in every factory to put that prompt back, so an apply always shows you the
plan and waits. The cost is that Terragrunt cannot then apply units in parallel, which is why
rolling a change out across a whole factory takes a different route to changing one deployment.
Both workflows are documented separately. See Run a factory deployment locally for the routine case, and Apply a factory change to every deployment for the fan-out case.
CI/CD pipelines¶
A factory holds one definition and many deployments, and which deployments a change affects depends on what changed. A pipeline therefore has to be generated rather than written by hand.
That generation is done by
iac-factory-pipeline-helper,
a tool shared between factories. It scans the deployments directory, works out which deployments a
change affects, and writes the child pipeline YAML for the parent pipeline to trigger. Deployments
are batched across several child pipelines, because a stage holding more than about a hundred jobs
makes the GitLab UI unusable.
The helper is configured by a .pipeline-helper.yml file in the factory repository.
iac_tool: terragrunt
deployments_root: product-vars
includes:
- component: $CI_SERVER_FQDN/uis/devops/platform/ci-components/iac/terragrunt-deploy@0.2.0
inputs:
plan-job-name: .terragrunt-plan
enable-apply: false
drift:
exclusions_file: scripts/exclude_products.yaml
The helper only orchestrates. The jobs it generates extend a hidden template job supplied by the
terragrunt-deploy CI
component, which knows
how to run a plan and publish it as an artefact. Neither the helper nor the component knows anything
specific to a given factory, because Terragrunt already resolves the backend, the inputs and the
module. A generated job only has to say which directory to run in.
Where a factory needs job types the helper does not provide, it can register a plugin. The nightly
GitLab access token rotation in gitlab-project-factory works this way.
Merge request pipelines¶
On a merge request the helper scopes plan jobs to what changed. Everything outside the deployments directory counts as shared, so changing the root module plans every deployment and gives reviewers the full blast radius. Changing one deployment's variables file plans only that deployment.
Plans on a merge request are allowed to report pending changes. Nothing is applied.
Drift detection¶
A scheduled pipeline runs nightly with IAC_DRIFT_MODE set. This plans every deployment rather than
only those affected by a change, and fails any job whose plan finds changes. Deployments named in
the factory's exclusions file are skipped.
A failure means the deployed infrastructure no longer matches what is defined in the repository. The Platforms Team is alerted by email and in the Platforms Team notifications channel in Microsoft Teams, and somebody reviews it the following working morning. This is what stops manual changes going unnoticed.
Applies¶
Generated pipelines plan. They do not apply. gcp-product-factory and ais-api-factory both set
enable-apply: false, and applies are run by hand from a workstation after review.
The one exception is the nightly token rotation in gitlab-project-factory, which does apply from
CI. It is deliberately narrow. The rotation plan is restricted with -target to the token resources
alone, a separate job validates that the plan contains only those resources before anything is
applied, and the apply is confined to the default branch. Anything outside that list fails the
pipeline rather than being applied unreviewed.
Summary¶
Our factories deploy one infrastructure definition many times, with isolated state per deployment. Terraform declares the infrastructure and Terragrunt selects the deployment, supplies its inputs and derives its state location, which replaced the per-repository shell scripts we used to maintain.
Two layouts matter. The single root module is a pre-Terragrunt hangover that ties a repository to
one root module, and the catalog layout separates Terragrunt units from Terraform modules under the
catalog/units and catalog/modules names Terragrunt itself recommends. We are moving to the
latter for new factories. Pipelines are generated by iac-factory-pipeline-helper, which
plans what a change affects, detects drift nightly, and leaves applies to a human in all but one
tightly scoped case.
See also¶
- Run a factory deployment locally walks through planning and applying a change to one deployment.
- Apply a factory change to every deployment covers rolling a change out across a whole factory.
- Guidance on testing Terraform modules covers our approach to testing the shared modules a factory consumes.
- Deployment reference describes the standard approach for ordinary product and service deployments.
- Each factory repository has a
docs/terragrunt.mddescribing its local workflow.gitlab-project-factoryadditionally has adocs/ci-pipeline.mdwith a full account of how its pipeline is generated.