LIVE ON AMD INSTINCT · PATENT PENDING

You bought an AI factory. Does it pay?

InferPulse puts one ratio on every workload running on your AMD estate — value produced per hour over cost per hour, recomputed every five seconds. Every inference microservice lands in one of four zones: scale it, maintain it, optimise it, or stop it. Read natively from ROCm, AMD-SMI, Helios, XGMI, KFD and vLLM, not through a generic GPU abstraction.

THE PROBLEM

Nobody commissions a plant without a P&L per line. Somehow we did it with GPUs.

If you commissioned a manufacturing plant, you would know the unit economics of every line in it before the first shift. Which lines are profitable, which are marginal, and which are running because somebody started them and nobody stopped them. That is not sophisticated finance. It is the minimum condition for operating a capital asset.

AI infrastructure was bought differently, and for understandable reasons. Fast, on strategic urgency, in the middle of a capability race, with a business case written at the level of the whole estate. That was defensible in year one. It is not defensible in year three, when the CFO asks which half of the spend is working and the honest answer is that nobody has ever measured it at that resolution.

Finance sees the invoice. Engineering sees utilisation. Neither can tell you the return on a specific workload — which means the estate gets defended as one number or cut as one number.

And when the pressure arrives, the failure mode is predictable. The cuts land on whatever is easiest to switch off rather than whatever is least productive. Which means the workloads that die are the ones with the weakest internal sponsor, not the weakest economics. That is not a failure of will. It is a failure of instrumentation: if you cannot rank workloads by return, you will rank them by politics, because politics is the only ranking available.

InvoiceWhat finance can see today
UtilisationWhat engineering can see today
ReturnWhat neither can see, per workload
5sRecomputation interval
every workload, continuously
0Behavioural drift dimensions
CUSUM statistical process control
0Native AMD telemetry feeds
into the economics, not a side panel
0Board and CFO-ready reports
no placeholder content, real factory data
THE METRIC
WEI

The Workload Economic Index is a ratio, not a dashboard.

WEI is the value a workload produces per hour divided by what it costs per hour. Every microservice gets a WEI score every five seconds, and the number integrates behavioural drift, SLA status, GPU cost profile and an industry value model — so a workload that is technically healthy but producing degraded output does not keep scoring as if it were fine.

The point of the metric is that it produces an action, not a conversation.

ZoneWEIWhat it meansAction
Scale≥ 3.0Value production materially exceeds cost. Your most profitable workloads, and usually your most under-resourced.Increase GPU allocation
Maintain1.5 – 3.0Stable positive economics. Healthy, but watch for drift pushing the ratio downward.Monitor weekly
Optimise1.0 – 1.5Declining. Behavioural drift, SLA breach or fabric saturation is eroding the value side.Root cause investigation
Kill< 1.0Negative economics. Every hour this runs costs more than it produces.Suspend or reconfigure now
inferpulse · workload economicsidle
$ select a workload to score
Illustrative sequence built on real platform behaviour. Not connected to a live cluster.
CORE CAPABILITIES
FOUR PILLARS

Measure the output, price the input, catch the drift, and put it in front of the person who signs.

PILLAR 01

Native AMD telemetry

Built for AMD infrastructure, not adapted for it.

Generic GPU abstractions lose exactly the information that makes the economics real. InferPulse reads the AMD stack directly: Instinct MI300X, MI325X and MI355X cost profiles, ROCm platform counters for compute occupancy and memory bandwidth, AMD-SMI for utilisation, power and VRAM, XGMI Infinity Fabric bandwidth with per-GPU workload attribution, KFD process-level assignment for precise cost attribution, vLLM latency and time-to-first-token, and AITER kernel mode distribution.

  • Every feed plugs into the WEI computation, not a separate monitoring view
  • Architecture-aware thresholds across CDNA3, CDNA4 and CDNA5
  • Fabric-bound vs compute-bound distinguished, so you fix the right thing
PILLAR 02

Workload economics & portfolio

A ranked portfolio position, recomputed every five seconds.

Every workload is classified automatically on every tick, with a plain-language explanation and dollar figures attached. The recommendations table says what to do and what it saves. Click any workload reference anywhere in the product — fabric matrix, drift chart, SLA table, leaderboard — and get the full economics, behavioural history and value attribution for that service.

  • Four zones with an action, not a status colour
  • Universal drill-down from any reference to full workload detail
  • WEI headroom quantifies the recoverable value still in the estate
PILLAR 03

Drift as an economic signal

Six dimensions under statistical process control, continuously.

Drift is usually treated as a quality issue, which routes it to an engineering backlog. Treated as an economic signal it routes to the person who owns the budget, and it moves. InferPulse tracks six behavioural dimensions using CUSUM — the same technique that has run manufacturing quality control for sixty years — and when a dimension drifts, WEI adjusts in real time, before a human notices the degradation.

  • Economics move the moment behaviour moves, not at month end
  • Helios power overlaid on WEI — racks in combined crisis surface first
  • Fault correlation links hardware events to economic impact
PILLAR 04

Reporting for the people who ask

Seven reports, because the CIO, CFO, CISO and compliance need different vocabulary.

Factory economics board pack with grade, portfolio matrix, priority actions with dollar savings, factory P&L and a 30/60/90-day forecast. WEI ROI analysis with payback period and net annual return. Neocloud multi-tenant report with per-tenant economics. EU AI Act compliance certificate per workload. Behavioural audit report with eight-week WEI and drift history. Conflict investigation report tracing an instruction conflict to business impact. Every one prints to PDF in one click with no placeholder content.

  • One underlying truth, expressed in each stakeholder's language
  • Real factory data only — nothing synthetic in a board pack
  • Patent pending on the scoring method
PROOF
ON REAL SILICON

Connected live to AMD Instinct, not modelled against a spec sheet.

InferPulse has run against a live AMD Instinct MI300X instance on the AMD Developer Cloud, reading production telemetry from a real inference workload — Qwen3-8B served through vLLM — and computing WEI from it end to end. Device metrics via a ROCm-backed exporter, inference telemetry from the serving layer, economics computed on top.

Two things matter about that. The first is that the numbers came from silicon rather than a simulation. The second is what the exercise exposed: on a virtualised instance, several telemetry endpoints simply are not reachable, and a platform that assumes a clean bare-metal feed produces nothing. Handling the degraded case is most of the engineering.

MI300XAMD Instinct, live instance on AMD Developer Cloud
126msTime to first token, measured through vLLM
2.25Live WEI on the served workload — SCALE zone
Grade AFactory grade computed from live telemetry
ESTATE VIEW · WEI DISTRIBUTION512 WORKLOADS
Kill — negative economics Optimise — declining Maintain — stable positive Scale — under-resourced and profitable
DEPLOYMENT
& TRUST

Reads the factory. Never sits in the inference path.

Anything that adds latency to inference gets removed the first time a workload misses its SLA. InferPulse is a side-channel observer: it reads telemetry that is already being produced, and never touches serving.

NO LATENCY

Side-channel by design

Zero changes to the inference path and no proxy in front of the serving layer. If InferPulse stops, inference does not notice.

DEGRADED FEEDS

Partial telemetry handled

Virtualised instances and locked-down clusters do not expose every endpoint. The economics degrade gracefully rather than going blank.

TENANCY

Your cluster, your data

Deployable inside your own environment. Multi-tenant mode for neocloud and GPUaaS operators, with per-tenant economics and isolation.

COMPLIANCE

Evidence as a by-product

Per-workload certifiable statements covering the EU AI Act, NIST AI RMF, GDPR and ISO 42001, with hash chain summary and signature block.

HOW A WORKLOAD GETS PRICED
01Native AMD telemetry read from ROCm, AMD-SMI, XGMI, KFD and vLLM
02Six drift dimensions scored under CUSUM control
03Value and cost per hour computed, drift-adjusted
04Zone assigned with a dollar-denominated recommendation
INSIGHTS

Notes from inside the AI factory.

What we keep finding when the economics of a GPU estate get measured at workload resolution, and what it changes about how the estate gets run.

FACTORY ECONOMICS · SEPTEMBER 2026 · 8 MIN

The cuts always land in the wrong place

If you cannot rank workloads by return, you will rank them by politics — because politics is the only ranking available.

Nobody commissions a manufacturing plant without knowing the unit economics of every line in it. You would know which lines are profitable, which are marginal, and which are running because somebody started them and nobody stopped them. That is not sophisticated financial management. It is the minimum condition for operating a capital asset.

AI infrastructure got bought differently, and for reasons that were entirely defensible at the time. It was bought fast, on strategic urgency, in the middle of a capability race, with a business case written at the level of the entire estate.

That holds for a year. It does not hold for three, and year three is where a lot of organisations now are.

The meeting everyone is about to have

It starts with a reasonable question from finance: which half of this is working?

The honest answer in most organisations is that nobody has ever measured it at that resolution. Finance can see the invoice. Engineering can see utilisation. Neither can tell you the return on a specific workload, which means the estate can only be defended as one number or cut as one number.

And when it gets cut as one number, watch where the cuts land. They land on whatever is easiest to switch off. Not the least productive workload — the least defended one. The team that shipped something quietly and moved on. The experiment whose sponsor changed roles. Meanwhile the workload with an executive champion and a slide survives regardless of what it returns.

Why the ranking has to be a ratio

The instinct is to build a dashboard. Utilisation by cluster, spend by team, tokens by application. I have sat in front of a lot of those dashboards and they produce discussion rather than decision, because none of the numbers can be compared to each other.

What changes the meeting is a single ratio per workload: value produced per hour over cost per hour. One number, comparable across everything, with an action attached at each band.

  • Well above one. This is your most profitable compute, and it is usually under-resourced because nobody knew.
  • Comfortably above one. Stable. Watch it for drift pushing it down.
  • Just above one. Declining. Something is eroding the value side and it is worth knowing what.
  • Below one. Every hour this runs costs more than it produces. There is no version of that argument that improves with time.

The point of the ratio is not elegance. It is that it produces an action instead of a conversation, and it moves the decision from whose project survives to what the portfolio returns.

The part people resist

Two objections come up every time, and both are fair.

The first is that value is hard to quantify for some workloads. True — and the answer is not to abandon the measurement but to be explicit about the value model, per industry and per workload type, and let people argue with it. An explicit model somebody disputes is more useful than an implicit one nobody can see.

The second is that killing a workload has switching costs the ratio does not capture. Also true, and worth handling as an override with a reason recorded, rather than as a reason to avoid measuring. The override is the interesting data point. If a workload has been overridden for three quarters running, that tells you something the ratio alone does not.

What good looks like

The test I would apply: can you name, right now, the three least productive workloads in your estate and say what each costs per month?

If the answer takes more than a few seconds, the next budget cut will be made on politics whether or not anyone intends it to be.

That is not a technology gap. It is the same gap a plant manager would have closed in the first month, and it has been solved before in every other capital-intensive industry. Compute is simply the newest one, and it is currently being managed with less rigour than a warehouse.

Vyasa Murthy is the founder of Venture Vertex LLC.
Write to vyasa.murthy@venture-vertex.com — the most useful replies start with why this will not work in your environment.

DRIFT · SEPTEMBER 2026 · 7 MIN

Drift is an economic signal, not a quality ticket

A workload that is technically healthy but producing degraded output should stop scoring as if it were fine. Immediately, and without waiting for a human to notice.

Here is a failure mode that costs more than any outage and generates no alerts at all.

A model-serving workload runs at healthy utilisation. Latency is inside its envelope. No errors, no restarts, no saturation. Every infrastructure dashboard is green. And the output quality has been sliding for eleven days, because a retrieval connector changed a chunk size, or an upstream prompt was edited, or the traffic mix shifted.

Nothing crashed. So nothing fired.

Where drift currently goes to die

When drift is eventually noticed — usually by a customer, or by somebody reading outputs by hand — it gets raised as a quality issue. Which means it enters an engineering backlog, gets prioritised against feature work, and waits.

Meanwhile the compute keeps billing. The workload keeps consuming GPU hours at full cost while producing output that is worth materially less than it was a fortnight ago. Nobody has quantified that gap, so nobody is urgently closing it.

That is a routing problem, not an engineering failure. The signal went to the wrong function.

Borrowing from manufacturing

The technique that solves this is sixty years old and has nothing to do with AI. Statistical process control — CUSUM in particular — was built for exactly this shape of problem: detecting a small, sustained shift in a process mean, as early as possible, without firing on ordinary noise.

A threshold alert asks whether today's value crossed a line. CUSUM asks whether the accumulated deviation from the baseline has become implausible. That difference matters enormously for drift, because drift is rarely a step change. It is a slow slide where no individual day looks wrong.

Track several behavioural dimensions that way — instruction governance, behavioural integrity, inference performance, session integrity and so on — and you get early warning on a class of degradation that no infrastructure metric represents.

The move that makes it matter

Detection alone still routes to the backlog. The change that makes drift actionable is to wire it into the economics.

When a behavioural dimension drifts, the value side of that workload's ratio adjusts in real time. Cost is unchanged — the GPUs cost what they cost. So the return falls immediately, the workload moves down the portfolio ranking, and it surfaces in the same view the CFO and the platform owner are already using to make allocation decisions.

Now it is not a quality ticket competing with feature work. It is a workload losing money, visible to the person who can act on it.

The reframe costs nothing technically and changes the response time by an order of magnitude, because it moves the problem from a queue into a P&L.

One caution

Do not let the economic framing erase the engineering detail. Knowing a workload's return has fallen tells you to look; it does not tell you what broke. The value of the drift dimensions is that they narrow it — a fall concentrated in instruction governance points somewhere very different from one concentrated in inference performance or fabric saturation.

Both halves are needed. The economics get the attention. The dimensional breakdown makes the attention productive.

Vyasa Murthy is the founder of Venture Vertex LLC.
Write to vyasa.murthy@venture-vertex.com — the most useful replies start with why this will not work in your environment.

FIELD NOTES · SEPTEMBER 2026 · 6 MIN

What we learned plugging into a live MI300X

Handling the degraded case turns out to be most of the engineering. A platform that assumes a complete feed produces nothing on a virtualised instance.

There is a comfortable stage in building infrastructure software where everything works against your own test harness. The feeds are complete, the endpoints respond, the metrics arrive in the shape the parser expects. It is a pleasant stage and it teaches you very little.

We took the platform and ran it against a live AMD Instinct MI300X instance, serving a real model through a real inference stack, and computed the economics end to end from what came back.

Two things were worth the exercise, and only one of them was the result.

The result

The numbers came from silicon rather than a model of silicon. Device-level metrics through a ROCm-backed exporter, inference telemetry from the serving layer, time-to-first-token measured rather than estimated, and a workload economic ratio computed on top of all of it.

That matters for a specific commercial reason. Anyone can show a projection. Very few platforms in this category have run against the hardware they claim to understand, and the gap between those two states is usually where the interesting problems live.

The lesson, which was more valuable

On a virtualised instance, several of the telemetry endpoints we designed around simply were not reachable. Not broken — not exposed. Different virtualisation configurations surface different subsets of the hardware interface, and a locked-down enterprise cluster will be stricter still.

A platform that assumes a clean bare-metal feed does one of two things when that happens. It produces nothing, or worse, it produces a number that looks plausible and is quietly wrong because one input silently defaulted.

So we spent the following sprints on the unglamorous half of the problem:

  • Explicit degradation. When an input is unavailable, say so in the output rather than substituting a default and moving on.
  • Confidence that travels with the number. An economic ratio computed from six feeds and one computed from three are not the same claim, and the interface should not present them as though they were.
  • Substitution paths. Where one feed is unreachable, derive what can be derived from what is present and mark what cannot.
  • Architecture awareness. Thresholds that are correct for one GPU generation are wrong for another, so the generation has to be detected rather than configured.

Why this is the real product

A demonstration environment rewards the happy path. A customer environment is almost never the happy path — there is a security policy, a virtualisation layer, a cluster someone else operates, and an endpoint that was closed for a good reason two years ago by a person who has since left.

The interesting engineering in infrastructure software is not the calculation. It is what the calculation does when an input goes missing.

If you are evaluating anything in this category, that is the question I would lead with. Not what does it measure, but what does it do when it cannot measure. The answer separates software that has been run in anger from software that has been demonstrated.

Vyasa Murthy is the founder of Venture Vertex LLC.
Write to vyasa.murthy@venture-vertex.com — the most useful replies start with why this will not work in your environment.

Bring your GPU count and your workload mix.

Thirty minutes is enough to show you the economics of your own factory. If the answer turns out to be that everything is fine, that is a useful answer too.

AMD factory owners, neocloud operators and GPUaaS providers.