Skip to main content

Controlling Costs

Most cloud bills are not expensive because the resources are big. They are expensive because the resources are always on — including nights, weekends, and the eleven hours a day nobody is looking at them.

Infrastream gives you two levers, and they work in this order:

  1. Size it right once. The estimated cost appears on the change before you approve it, so you pick the size with the number in front of you rather than discovering it on an invoice.
  2. Switch it off when nobody is using it. A hibernation schedule puts a whole project to sleep on a recurring timetable and wakes it up again, automatically.

This guide covers the second lever, which is where the large savings are.


In plain words: what hibernation is​

Think of a non-production project — your staging or development copy — as an office. Right now you are paying rent on it twenty-four hours a day, including the hours the lights are off.

A hibernation schedule is the timer on the lights. You declare the hours nobody is working, and Infrastream shuts the expensive parts down for exactly those hours and brings them back before you return.

Three things matter to a non-technical reader, and all three are good news:

  • Your data is not deleted. Hibernation removes the compute — the running machinery that gets billed by the hour. The stored data, and its backups, stay exactly where they were.
  • It is not a button someone has to remember to press. The schedule is part of the project's declared configuration, so it happens whether or not anyone thinks about it.
  • You can suspend it. If you need the environment during a normally-asleep period — a demo, a release weekend — you declare an exclusion for those dates and the schedule stands down.
Non-production only

Hibernation is designed for development, integration, staging, and sandbox projects. Nothing in the platform stops you from putting a production project to sleep, so treat "never on production" as a review rule your team enforces on the pull request.


What actually gets switched off​

This is the part worth reading carefully, because hibernation does not affect everything in a project, and budgeting on the assumption that it does will disappoint you.

ResourceWhat hibernation doesBilling effect
AlloyDB (PostgreSQL)The database instances are removed. The cluster, its storage, and its backups remain.This is the big one. AlloyDB instances bill per vCPU-hour, and that stops. Storage and backups continue to bill.
Cloud SpannerScaled down to the 100-processing-unit floor, autoscaling switched off.Drops to minimum, does not reach zero.
Applications on Compute EngineThe managed instance group is set to zero instances and its autoscaler is suppressed.Compute billing stops.
Applications on KubernetesReplicas set to zero and no horizontal autoscaler is provisioned.Pods stop; the cluster itself keeps running.
Virtual machinesAutoscaling floor and ceiling both set to zero.Compute billing stops.
Public ingress gatewaysThe gateway instance group is set to zero, and supporting sidecars drop to a single instance cap.The environment becomes unreachable while asleep — which is the intent.
Applications on Cloud RunNothing. Cloud Run scaling comes from the application's own scaling block.Set scaling.min to 0 on non-production services instead; Cloud Run then bills only per request.
Redis, GKE clusters, MCP and Agent deploymentsNothing today.These keep running and keep billing through the window.
Why AlloyDB is removed rather than paused

AlloyDB has no pause operation — an instance either exists and bills, or it does not exist. So the engine expresses "asleep" as "the instance is not declared", and expresses "awake" as "it is". The cluster holding your data is declared unconditionally, which is why the data survives untouched.

The practical consequence is that waking up is not instantaneous. Creating an AlloyDB instance takes several minutes, so schedule the wake-up boundary comfortably before the first person needs the database, not at the exact minute.


Declaring a schedule​

Hibernation is a field on the Project manifest. A window is a recurring period during which the project sleeps, expressed as two cron expressions — when sleep begins and when it ends.

# .../environment/development/project/my-project/my-project.yaml
apiVersion: lowops.manifests.v1
kind: Project
metadata:
name: my-project
environment: development
organizational-unit: my-unit
organization: my-org
spec:
displayName: My Project - Dev
hibernation:
# 'hibernate: true' is a manual override that sleeps the project immediately
# and ignores every window below. Leave it false for schedule-driven sleep.
hibernate: false
windows:
# The map key is your own name for the window.
overnight:
start: "0 20 * * *" # asleep from 20:00
end: "0 7 * * *" # awake again at 07:00

Both start and end are required on every window. A window with only one boundary is a hard error at plan time, not a warning.

Time zones: everything is UTC​

The cron expressions are evaluated in UTC, and the scheduler that fires them is pinned to UTC. There is no time-zone field on a window today, so you must convert your working hours to UTC yourself.

If your region observes daylight saving, that conversion changes twice a year and the manifest has to be edited both times. If your region does not observe daylight saving, a UTC cron is stable forever — write it once and leave it.

Suspending the schedule temporarily​

An exclusion is a one-off period during which the schedule stands down, for a release window or a customer demo. Exclusions are timestamps rather than crons, and they do carry a UTC offset.

spec:
hibernation:
windows:
overnight:
start: "0 20 * * *"
end: "0 7 * * *"
exclusions:
release-week:
start: "2026-10-05T00:00:00Z"
end: "2026-10-09T23:59:59Z"
Exclusions need a reconcile to take effect

An exclusion suppresses hibernation whenever the platform next evaluates the project, but it does not currently schedule its own wake-up. If an exclusion begins at a time when no window boundary happens to fire, the environment stays asleep until the next boundary or the next merged change. Plan an exclusion to begin before a window boundary, or merge any change at the start of the excluded period to force the evaluation.


How the schedule runs itself​

You do not need this section to use hibernation, but you do need it to predict what happens on your Git repository.

Each distinct cron boundary in your organization becomes one scheduled trigger. When a boundary fires, it runs a full reconcile of the organization against the main branch — the platform re-reads the manifests, re-evaluates which projects should currently be asleep, and applies the result.

Two consequences follow, and both are worth knowing before you add your first window:

  • main must always be in an applyable state. A hibernation boundary applies whatever is on main at that moment, not just the hibernation change. If someone merged something broken at 17:00, the 20:00 boundary is when it lands.
  • Identical boundaries are shared. If five projects all sleep at 20:00, that is one trigger, not five. Reusing the same boundary times across projects is cheaper and easier to reason about than giving every project its own slightly different schedule.

Inheritance: parents win, and children cannot opt out​

Hibernation can be declared at any level of the hierarchy — organization, organizational unit, environment, or project — and the computed result is a logical OR up the chain.

If a parent is hibernating, every descendant hibernates. A child setting hibernate: false does not exempt it; there is no opt-out. This is deliberate — it means "put the whole sandbox organizational unit to sleep" is a single, reliable change — but it also means declaring a schedule at organization level will sleep production along with everything else.

The safe default is to declare hibernation on individual non-production projects, and to move it up a level only when you genuinely mean every project underneath.


A worked example: an eight-hour daily window​

Suppose your team is active from 15:00 to 07:00 local time, in a region four hours ahead of UTC that does not observe daylight saving. The idle period is 07:00 to 15:00 local, which is 03:00 to 11:00 UTC.

spec:
hibernation:
hibernate: false
windows:
daily-idle:
start: "0 3 * * *" # 07:00 local — everyone has stopped
end: "0 10 * * *" # 14:00 local — an hour early, so the database is ready by 15:00

Note the wake-up is set an hour before anyone needs it. That hour is cheap and it absorbs the time an AlloyDB instance takes to come back.

To add the weekend, declare a second window rather than complicating the first. Windows combine as an OR — the project sleeps if any window says it should.

      weekend:
start: "0 3 * * 6" # Saturday 07:00 local
end: "0 10 * * 1" # Monday 14:00 local

Together these put the project to sleep for roughly half of every week.


The other lever: right-sizing​

Scheduling saves you the hours. Sizing saves you the rest.

  • Databases. spec.alloydb.cpuCount is the single biggest number on a database bill. The floor is 1 vCPU, which requires a machine shape that is not available in every region — if a 1-vCPU database is rejected where you are, 2 is the effective floor. See Provisioning a Database.
  • Cloud Run services. Setting the minimum instance count to zero on non-production services means you pay per request instead of per hour. Leave a warm minimum only on services where a cold start is genuinely unacceptable, such as a login path.
  • Compute target. Moving a service from Kubernetes to Cloud Run is a change to one field. See Migrating a Service to Cloud Run.
  • Read replicas. spec.alloydb.clusterSize above 1 adds read-pool instances, each billed like the primary. Non-production rarely needs them.

Learn more​