Skip to main content

Growing team: stop being the human API

This arc starts on the day your one-person project becomes a team. The work itself has not got harder — what has changed is that every small change now queues behind the one person who knows how the infrastructure fits together, and that person is spending their day answering questions instead of building. The five stories here take that bottleneck apart: the agent drafts the routine changes, staging becomes a real copy instead of a hopeful guess, two projects talk to each other privately, the rules about who may reach what become a line anyone can read, and the whole estate shows up on one picture with the cost attached.


10. Let the agent handle the routine stuff​

The win: Nobody has to interrupt the one person who knows the infrastructure just to get a memory limit raised.

Who this is forAn engineer on a small team who keeps waiting on the one colleague allowed to touch infrastructure
DifficultyIntermediate
TimeMinutes per change, no upfront setup
What you need firstAt least one app already running on Infrastream, and the pvot agent given permission to open pull requests — proposed changes — in your manifests repository
What it costsNothing extra for asking. If the change resizes something, the cost estimate appears on the proposed change before anyone approves it

In plain words​

New engineers keep asking Sam for small things: bump a memory limit, add an environment variable, stand up a staging copy of one endpoint. Each request used to pull Sam out of real work, because Sam was the only person who knew where the settings lived. Now the request goes to pvot, the platform's AI agent, instead. It reads the live state of what you actually have running, writes the change, and opens it as a pull request — a proposed edit that sits there until a human reads it and says yes. Any engineer on the team can be that human. Sam stops being the bottleneck for work that never needed Sam in the first place.

What you actually write​

For this story you do not write YAML at all. You write a sentence. In the portal's chat surface, an engineer types the request the way they would say it out loud:

"Raise the memory limit on checkout-api in staging to 1Gi and keep at least one instance warm."

pvot finds the DeploymentConfig that governs checkout-api in that environment and drafts the edit for review. The change it proposes looks like this:

apiVersion: lowops.manifests.v1
kind: DeploymentConfig
metadata:
name: checkout-api
project: storefront
environment: staging
organizational-unit: retail
organization: fincorp
spec:
version: v1.4.2
container:
resources:
limits:
cpu: "1"
memory: 1Gi # ← the only value the request asked to change
scaling:
min: 1 # ← keeps one instance warm, as asked
max: 5
# ... ports, health probes and env unchanged

Nothing is applied when pvot finishes writing. The pull request waits for a human to review and merge it, exactly as it would for a change a colleague wrote by hand.

What the platform handles for you​

  • The agent reads your real, current state before it writes anything, so it edits the manifest that is actually live rather than one it guessed at.
  • The change arrives scoped to what was asked — one field on one file — which is what makes it reviewable in thirty seconds by someone who is not an infrastructure expert.
  • Review and approval stay with people: every proposal is gated by the repository's CODEOWNERS rules and by the approval stages on your release track.
  • The same request history lives in Git, so six months later "who raised this and why" is a commit, not a memory.

Learn more​


11. Split into two projects without starting over​

The win: A staging environment that is a genuine copy of production, not something rebuilt from memory.

Who this is forA small team that has outgrown testing changes directly on the live system
DifficultyIntermediate
TimeAbout 15 minutes
What you need firstAn app already described in manifests on Infrastream. The second Google Cloud project can be created for you — you do not have to make it yourself
What it costsRoughly a second copy of what you already run, and this is the strongest candidate on the page for an off-hours schedule — see Controlling Costs

In plain words​

"Test it in production" was a fine joke when it was just Sam. With three engineers shipping every day it stops being funny. The team needs a staging environment — a second, separate copy of the whole system where you can break things safely. Because everything you run is already described in manifests (the plain-text files that say what you want), a second copy is a second set of those files rather than a weekend of rebuilding by hand. Same API shape, same database schema, same login wiring, running in its own isolated space where nothing it does can touch real customers.

What you actually write​

A new Environment for staging, and the Project that lives inside it. Both are the same shapes you already used for production, so in practice you are copying your existing files and changing the names.

# .../organizational-unit/retail/environment/staging/staging.yaml
apiVersion: lowops.manifests.v1
kind: Environment
metadata:
name: staging
organizational-unit: retail
organization: fincorp
spec:
displayName: Retail Staging
hibernation:
windows:
out-of-hours: # ← paused outside working hours
start: "0 19 * * 1-5" # cron: stop at 19:00 on weekdays
end: "0 8 * * 1-5" # cron: start again at 08:00
---
# .../environment/staging/project/storefront.yaml
apiVersion: lowops.manifests.v1
kind: Project
metadata:
name: storefront
environment: staging # ← the only line that differs from prod
organizational-unit: retail
organization: fincorp
spec:
displayName: Storefront Staging
defaultUrlRedirect: https://fincorp.io
# ... permissions, region and egress as in production
Schedules inherit downwards, and children cannot opt out

Put the schedule on the Environment and every Project underneath it sleeps on those hours. The inheritance is one-directional: if a parent is hibernating, a child setting hibernate: false does not exempt it. If one project genuinely has to stay up, declare the schedule on the individual projects that should sleep rather than on the environment above them.

To suspend a schedule temporarily — a demo, a release weekend — declare an exclusion rather than removing the window. See Controlling Costs.

What the platform handles for you​

  • A separate Google Cloud project and folder for staging, so a mistake there cannot reach production data or production users.
  • The inherited settings — region, permissions, domain segment — flow down from the environment, so the copy stays consistent instead of drifting into a slightly-different system.
  • A non-production project can be put on an off-hours schedule: the platform scales its workloads down overnight and at weekends and brings them back in the morning, so staging costs money only while people are awake to use it.
  • Promotion between the copies is ordered and gated by your release track, so a version reaches production by passing through staging rather than by someone remembering to deploy it.

Learn more​


12. Let two projects talk to each other privately​

The win: Three different kinds of machine on one private network, declared once instead of discovered during an outage.

Who this is forA team whose services no longer all live in the same place
DifficultyAdvanced
TimeOne to two hours
What you need firstStory 11 done, and workloads already running in both projects — serverless, virtual machines, Kubernetes, or a mix
What it costsOne always-on internal gateway, plus whatever the machines behind it cost; virtual machines and Kubernetes node pools can go on an off-hours schedule — see Controlling Costs

In plain words​

The team's new billing service runs on Kubernetes in one project. The main app is serverless in another. A virtual machine handles a legacy batch job nobody wants to touch. All three need to talk to each other, and none of that conversation should ever travel across the public internet. A private ingress is an internal-only front door: it gives the three a single address they can reach each other on, inside your own network, with nothing published to the outside world. You declare which projects are allowed through that door, and the platform builds the plumbing.

What you actually write​

One PrivateIngress in the project that hosts the shared internal API, naming the projects allowed to reach it.

apiVersion: lowops.manifests.v1
kind: PrivateIngress
metadata:
name: internal-api-gateway
project: billing
environment: production
organizational-unit: retail
organization: fincorp
spec:
description: "Internal gateway for billing APIs"
region: us-central1
authorizedProjects: # ← who is allowed to reach this door at all
- name: storefront
environment: production
organizationalUnit: retail
- name: batch-jobs
environment: production
organizationalUnit: retail

Each path through that door is a separate HttpRoute manifest that links to this ingress by name and points at a deployment and a port — see the private ingress guide below for the routing half.

What the platform handles for you​

  • An internal load balancer with TLS termination, reachable only from inside your network, so the billing API has no public address to attack.
  • The cross-project network connectivity itself: each authorized project is wired to the gateway, instead of three engineers independently working out each other's IP ranges.
  • An internal hostname for the gateway, so services address each other by name and keep working when the underlying addresses change.
  • The same mesh covers all three runtimes — a Kubernetes cluster, a serverless app, and a virtual machine joined via its meshStrategy setting — so one declared topology replaces three half-remembered ones.

Learn more​


13. Decide exactly who can reach what​

The win: "The batch machine cannot call billing" is a line anyone can read in the repository, instead of a sentence in an incident report.

Who this is forA team that now has enough moving parts to want the connections written down
DifficultyAdvanced
Time30 to 45 minutes
What you need firstStory 12 done, and a decision — in plain words is fine — about which service is allowed to call which
What it costsNothing extra for the rules themselves; the virtual machines and node pools they govern are the cost, and those can go on an off-hours schedule — see Controlling Costs

In plain words​

Now that the projects can talk, the team wants limits. The legacy batch machine should reach the database and nothing else — it has no business calling the billing service, and everyone would rather learn that from a written rule than from an outage. Two written rules do this job together. Reaching out of a project to the wider internet is an allow-list on the project itself: anything not on the list is blocked by default. Reaching in to an internal gateway is the authorized- projects list from Story 12. Both live in files in the repository, which means both are reviewed before they take effect and readable by anyone afterwards.

What you actually write​

The outbound allow-list goes on the Project manifest, and applies to everything running inside that project.

apiVersion: lowops.manifests.v1
kind: Project
metadata:
name: batch-jobs
environment: production
organizational-unit: retail
organization: fincorp
spec:
displayName: Batch Jobs
defaultUrlRedirect: https://fincorp.io
region: us-central1
allowedEgress: # ← the complete list of outside destinations
- "*.googleapis.com" # Google Cloud APIs
- "smtp.sendgrid.net" # nightly report email
# ... permissions and maintenance unchanged

To control who may reach an internal service, edit the authorizedProjects list on the PrivateIngress from Story 12 — removing a project from that list removes its route to the door.

note

allowedEgress governs traffic leaving the project for destinations outside your shared network, and it is additive: it grants access, it cannot be used to forbid traffic that is allowed by default, such as traffic inside the network. Internal reachability is controlled by which projects you authorize on the ingress, not by this list.

What the platform handles for you​

  • Firewall rules generated from the allow-list, so a service that was never meant to call the outside world simply cannot, with no code change required to enforce it.
  • A default-deny posture for outbound traffic, which means a forgotten dependency shows up as a connection timeout during testing rather than as a surprise in production.
  • Changes take effect without restarting your applications, so tightening a rule is not an outage.
  • Every edit to either list is a reviewed change in Git, so the question "when did this service get permission to do that?" has an answer with a date and a name on it.

Learn more​


14. See the whole thing on one graph​

The win: "What are we running, and what does it cost?" becomes something you look at, not something you ask three people.

Who this is forAnyone on the team, including people who do not write code
DifficultyBeginner
TimeInstant, no setup
What you need firstAt least one project on Infrastream and an account that can sign in to the portal
What it costsNothing — the view is part of the portal you already have

In plain words​

Three engineers, two projects, a private connection between them, a website and a chat bot and a mobile app on the front. Working out what actually exists used to mean asking everyone on Slack and hoping the answers agreed. The portal shows it as one picture instead: every resource, every connection between them, which project each one lives in, and the cost attached — kept current as changes are merged. You do not have to understand a single manifest field to read it, and you do not have to ask permission to look.

What you actually write​

Nothing. This is the one story on the page with no file to author. The picture is built from the manifests you have already merged, which is exactly why it is trustworthy: it is a drawing of the declared state, not a diagram someone updated by hand in March and forgot about. Open the portal and the graph is already there.

If you want a written answer rather than a picture, ask the agent — questions such as "what changed in the storefront project in the last hour, and by which pull request?" are answered from the same underlying state. And when looking at the graph makes you want to change something on it, you do not have to work out the manifest yourself.

Ask the pvot agent to draft this one — see Using Pvot.

What the platform handles for you​

  • One view across every project and environment, so nothing is invisible just because it belongs to somebody else's corner of the system.
  • Cost attributed to the resources that incur it, which turns a monthly surprise into something you can see while deciding whether to build the thing.
  • The connections between resources, not only the resources — so you can see what would be affected before you change something.
  • Continuous updating as changes merge, because the platform builds the same dependency graph to provision your infrastructure that it draws for you to read.

Learn more​


Where to go next​

At this point the team is no longer routing every change through one person, and everything it runs is visible in one place. The next arc is what happens when the organization gets large enough that isolation, audit, and migration matter more than speed: bringing an existing estate under management without a cutover weekend. Continue with Enterprise (Stories 15–18), or go back to Solo Builder (Stories 1–9) if you skipped the beginning.