KYLE GANIEssays

E-002

Kyle GaniMission Control12 min read

The margin you cannot see

When the product is a system, the change customers love can quietly turn your margin negative. You find out at month close, weeks too late to trace it back.

You ship a change your customers love. Usage climbs that week. The graph goes up and to the right, and everyone reads it as a good month. Then the bill for that usage comes in. The growth everyone loved is losing you money, because the feature it drove runs on the most expensive hardware you rent. You find out at month close. That is five weeks after the decision that caused it, and nothing in the report ties the number to the change.

Most product leaders meet their cost of goods this way. You get it after the fact, in a report nobody can argue with. You can watch usage in real time. What that usage does to the money stays dark for weeks, and the runway leaks away in the gap.

In ordinary software the gap does not matter much. One more customer costs close to nothing. The margin takes care of itself. Watching usage is the whole job. Systems businesses do not get that. When the product is infrastructure, developer tools, or an API, it spends money every time a customer touches it: compute, storage, bandwidth, model inference. The cost of goods is real. Product decisions set it.

I lead a product organization using a loop I call Mission Control: three capabilities, built on purpose, that decide whether a team compounds or just gets heavier. The first is sight, a picture of the future specific enough to rule options out. Most people meet it on a roadmap, where it shows up as the customers you decide not to serve. On a profit and loss statement it is narrower and harder: knowing which change you just shipped moved your margin, by how much, while there is still time to act.

A profit and loss statement, the P&L, is the record of what a business earned against what it spent. Margin is what is left after the cost of serving a customer is taken out of what they pay you. Unit economics is the profit or loss on a single unit of what you sell. Every systems business has one, whether that is an API call, a build minute, or a second of compute. Mine was the second of compute. The stories here come from an infrastructure platform, but the chain from the unit to the profit line is the same whatever yours is.

Every product decision is a financial decision

In a systems business, a product decision is a financial decision whether or not you can see the finance. Raise the storage quota on the free tier. Make retries more generous. Default a heavy job to a faster instance type because it feels sluggish. Every one of those is a kindness to the customer, and every one changes what that customer costs you to serve. The knob is in the product, and the cost shows up in the P&L a month later. Finance reports the number. The decisions that moved it were yours.

The margin does not care whether you were looking.

Half of this is easy. That easy half is the trap. Usage meters in near real time: requests, builds, seconds, which customers are growing. That graph is honest, it is immediate, and it feels like visibility. It is only half the picture. Cost is hidden when you make the decision and visible weeks later, by which time the decision is old and so is the quarter.

This week: take the last change you shipped to make the product feel better. A heavier default, a looser limit, a faster tier under a slow job. Name the one line in your cost of serving that it moved. You do not need the number yet, only the link. Trace a knob to a cost once and you will never turn it blind again, and the next change like it gets the cost question before it ships.

Sight is something you build

The number does not exist until you build it. Your bill shows spend and your product shows usage, with nothing joining them. The chain is what joins them.

It starts with the meter. Usage reaches you as a raw stream of events, late, duplicated, sometimes mislabeled. The first job is turning that stream into one settled unit of consumption: this customer, this app, this request, this container, this instance type, this region. Until that unit is solid, nothing built on it can be.

Then the unit gets a price. Every unit ties back to something you buy, and that bill keeps moving after you get it. It is a slowly settling estimate, revised for days. On our platform each second of compute tied to a specific machine type, an SKU, with its own cost per hour. In another business the line is model calls or bandwidth. Either way, yesterday’s cost is still changing today, so you re-check the past instead of trusting the first answer.

Then the hardest join, the one nobody sees. Your supplier charges you for capacity: machines, clusters, committed volume. What your customers send you is workloads, and many workloads share the same machine. Matching the charge back to the customers who caused it is the central problem of unit economics. The naive move is to divide the bill by usage. It is quietly wrong. Not every unit costs the same, and someone still has to carry the capacity nobody used. An honest system leaves that leftover in plain sight instead of smearing it across customers to make the picture tidy.

Then the unit meets the contract. Tiers, minimums, credits, rates promised months ago and never revisited: contracts decide what a unit is worth. The same second of compute can be profit on one customer and loss on another. This is where the deal that exploded your ARR meets the hardware it runs on, and you learn whether the rate you signed covers the cost. The output is an estimate of revenue. It stays honest only while it matches what customers are actually invoiced, so it gets reconciled against the billing system continuously.

Then the two sides meet. Revenue against cost is margin, which is what the whole chain exists to produce: profit or loss on each customer, each app, each instance type. Every customer can look profitable while the platform loses money. Idle capacity and cross-region traffic belong to no customer, so they never show up in a per-customer view. Utilization, meaning how much of what you paid for actually did work, is the biggest lever on the number and the one nobody sees.

The last link is the P&L: one number for the whole business. It is only worth trusting if you can trace it back, dollar by dollar, to the usage that caused it. Some dollars will not trace cleanly. Report those as their own line instead of hiding them inside someone else’s.

A report
Tells you the number after the month has closed.
Sight
Tells you which change moved the number while you can still act on it.

None of this is real time, whatever people mean by the phrase. The numbers build through the day and then keep settling for days afterward. So label what is still moving, and treat the latest figure as provisional until it stops changing.

This week: take one number your team already reports. A cost, a revenue figure, a margin. Trace it back toward the raw event it came from. Every place the trail goes cold is a link you have been trusting without seeing. The gaps, in the order you hit them, are your build list.

A spike and a trend look identical on the day

The first thing new sight did was make us twitchy.

Early on, a customer’s usage jumped hard. For the first time we could watch it move the numbers in something close to real time, and the instinct that comes with sight is to act. So we acted. We treated the new, higher level as the truth and repriced against it. Then the customer settled back to roughly where they had been, and the decision we rushed gave up revenue we never needed to give up. Our new sight had handed us a faster way to be wrong.

This is the failure the framework guards against: confusing a reflex with an adaptation. Adaptation, another of the three capabilities, is changing the route without moving the destination. A reflex changes the route before you know whether the road changed at all. A loud week from one customer is noise. Real signal is a durable shift in how they use you, and on the day it happens the two look identical. Sight will not separate them. Only time will, and acting early is what costs you the time.

The fix was to take the decision out of the moment entirely. We were never going to get better judgement in the moment. While nothing was moving, we wrote down what has to be true before usage changes a commercial term: how large the move, how long it must hold, who signs off, and what we check while we wait. The cheapest check turned out to be asking the customer. A spike usually has an ordinary explanation, a launch week or a one-off backfill. A durable trend loses nothing in the wait, and the wait is what kills the reflex.

The same volatility runs the other way, and there the cause is invisible to you. A customer’s usage can collapse for reasons that live entirely inside their business: a cost-cutting round, a project losing its sponsor. No amount of sight into your own platform will show you a budget meeting you were not in. That one gets solved in the contract. Commitments and notice periods, sized so a sudden downturn on their side is absorbed by the deal instead of your P&L. Price the swing in while the deal is being signed, because you will not see it coming afterwards.

This week: write your version of those rules while nothing is moving. The smallest change that counts as real, how many days it has to hold, who signs off before it changes a price. Then read your newest agreement and ask what happens to it the quarter that customer cuts costs. Do both in the calm. Once the graph is moving, the graph decides.

The numbers get worse when they get honest

The spike cost us one decision. The harder months came from the sight itself, because building it makes the P&L look worse before it looks better. For the first time you can see the losses you were blind to. The margins did not get worse. They got honest. The corner of the product that was underwater, a class of hardware in our case, does not start losing money the day you can measure it. It starts looking like it is losing money, which is the first step to fixing it. Ours got fixed.

That is the capability the framework calls conviction: holding the trajectory through the quarter where the numbers argue against it, because your picture told you this ugly middle was coming and roughly what it would look like. If you only hold course when the numbers agree, you do not have a strategy. You have a mood.

The fix is a hundred small corrections in the product itself. Packing workloads tighter so the capacity you already pay for does more work. Moving a job onto the instance type that actually fits it. Cutting the data that crosses regions for no reason. You change one thing and wait for the system to react. Read whether it moved the way you expected. Keep what worked.

A lot of those months went into the instrument itself. A cost charged to the wrong source. A figure that double-counted. A month-end projection I changed my mind about four times in one afternoon. You cannot read the platform’s reaction on a gauge you do not trust. People treat that grind as the boring part before the insight. It is the insight, one correction at a time.

Breakeven is a hundred small tweaks, each one checked against the system the next morning.

This week: find the number you have been avoiding because you suspect it is bad. Build just enough to see it honestly, even roughly, even ugly. A loss you can see is a problem you can work on. The one you cannot see is the one that takes the company. Then make a single change against it and look tomorrow to see whether it moved.

The loop runs on the money the way it runs on the roadmap

For a long time the last mile was me. Every morning I pulled the numbers from each supplier’s console, checked them against each other and against our own records, and chased down whatever did not line up. Then I worked out what had moved since the day before, where the month was heading, and what we were paying for but not using. Then I typed it into the team channel by hand. The sight existed. It just cost a person tedious hours every morning to exercise it.

The last step was handing those hours to a system. Everything I did by hand now runs as a scheduled job. It pulls the numbers, reconciles them, writes the morning summary and posts it to the team channel before anyone has logged in. The summary covers what moved in cost and margin overnight and the likely reasons. Where in the product to go and look. Which customers grew or went quiet, ranked by the dollars each one moves. Where the month is trending, and where money is going out without buying anything. It stops there. What happened, in plain language, and the decision left where it belongs.

The first morning it ran without me, the hours I used to spend were simply gone, and the read was sharper than mine. The automation matters less than what it made possible. The loop now runs on the money the same way it runs on the roadmap. You make a move: retire idle capacity, change how a workload scales, renegotiate what you buy. The next morning the system tells you how it answered. You poke it, you read the reaction, and that decides the next poke. It stopped being a report you skim and became a system you play against.

Sooner or later a systems business gives its product leaders a piece of the company they were never trained to run: the pricing, the margins, the cost of the thing they ship. The reflex is to defer to whoever owns the spreadsheet and feel around in the dark. The P&L reads like someone else’s instrument. It is the same loop you already run on the roadmap, pointed at the money. Build the sight. Hold the conviction through the quarter where the numbers get worse before they get better. Adapt when the signal turns out to be real.

Then make one small change and check in the morning to see how the system answered. This is the clearest case I have for the argument that anchors everything I write here: capability beats capacity. We never added a person to fix these numbers. What changed was what the same people could see, and a team that can see what its decisions do to the money will beat a bigger one that finds out at month close.

E-002/Filed JUL 28, 2026

Next in the registry

Mission Control8 min read

Wrong at scale

Hiring more people speeds up the plan you already have. It does nothing for the moment the plan turns out to be wrong.

The framework behind the writing, and the next dispatch when it ships.

Mission ControlSubscribe