Perspective
How We Decide What to Publish as a Number
Jayesh Verma
August 2026
9 min read
Our evidence policy, in public, including the part where we broke it last week. A number you cannot attribute to one deployment is worth less than no number at all.
Start with the uncomfortable part, because it is what prompted this page.
Last week we wrote a line that read: our measured Tier-1 carrier figures are 27,000 or more devices under one pane, 2,000 or more nodes under governed execution, around 60 percent of recurring faults resolved end to end, around 50 percent faster to root cause, around 70 percent faster patch rollout and around 35 percent fewer repeat incidents.
Every one of those numbers is real, is measured, and appears on this site with a case study behind it. The sentence is still wrong, because those six figures come from four different deployments and the sentence presents them as one. Two are Tier-1 telecom operators on different continents. One is a Tier-1 enterprise in North America that is not a carrier at all. Stacking them behind the words our Tier-1 carrier figures produces a composite customer who does not exist.
Nobody fabricated anything. That is exactly why it is worth writing down: the failure mode in this category is almost never invention. It is aggregation, and aggregation feels like summarising right up until somebody asks which customer.
The rule we should have followed
A headline set describes one deployment. If a figure comes from a different customer, it does not join the set. It lives on its own case study page with its own context, and it never stands next to figures it did not earn its place beside.
Applied honestly, that reduces our headline to four numbers from a single Tier-1 telecom operator in India.
| Figure | Basis |
|---|---|
| 27,000+ devices in one governed pane | Estate size at go-live. Countable, not estimated. |
| 43% MTTR reduction | Measured in production within six months of go-live, against the documented pre-deployment baseline. |
| 85% of routine manual operations automated | Measured in production, as governed MOPs with pre-check, post-check and audit trail. |
| ~60% of incidents auto-created with cause and context | Proportion of incidents raised by the platform carrying a probable cause, not a bare alarm. |
Those four are smaller and less impressive than the composite we wrote, and they are the ones that survive the only question that matters, which is: which customer, and can I see it. The case study is here, and it carries its own caveat that we should not have dropped: the MTTR and automation figures are measured production results from that deployment, and the remaining improvement figures in it are indicative.
The other deployments keep their own numbers on their own pages. The North America telecom closed loop reports around 60 percent of recurring network faults resolved end to end and around 50 percent faster to root cause. The day two operations engagement reports more than 2,000 nodes under governed execution and around 70 percent faster mass patch rollout, at a Tier-1 enterprise, not a carrier. Those are good results. They are not the carrier story and they will not be told as though they were.
Two tiers, marked on every page
The second rule predates the mistake and is already visible across the site, though we have never explained it in one place. Every use case carries one of two labels, and the label controls what the page is allowed to claim.
- Production patternDeployed, measured, published. The page links to the case study containing the measurement. Four of our 38 use cases carry this label, and all four are telecom.
- Platform capabilityThe platform does this, modelled onto a scenario we have not measured with a named customer. Every timeline on the page is marked illustrative. Thirty four use cases carry this label.
- What is not allowedA modelled timeline written to look measured. An industry benchmark reported as ours. A figure from an analogous public deployment quoted without saying whose it is.
The ratio is deliberately unflattering. Four measured out of thirty eight is not a strong number for a company that would like to look established, and we publish it as a count on the use cases page rather than burying it. The alternative is to relabel thirty four pages as production evidence, which takes an afternoon and destroys the only thing that makes the four mean anything.
The rules, stated so you can hold us to them
- One deployment per headline set. Figures from different customers do not appear in the same sentence, slide, or frame of a film.
- Every figure resolves to a case study. If a number on this site does not link to the deployment that produced it, that is a defect and we want to hear about it.
- Measured and modelled are labelled, never blended. A modelled timeline says so on the page it appears on, not in a footnote on a different page.
- Borrowed numbers are attributed or dropped. Industry research is useful and we cite it as somebody else research. An analogous vendor deployment is never presented as our result, however comparable the estate.
- The baseline is stated or the percentage is not published. A percentage improvement without a denominator is a decoration.
- Forecasts are marked as forecasts. We have an engagement with roughly 100,000 nodes forecast at target scale. That is a forecast, it is labelled as one on its page, and it does not appear in any headline.
Why this is worth the smaller numbers
The commercial argument is simple and it is not altruism. In a market where every vendor page claims a large percentage improvement and almost none of them states the deployment, the numbers have stopped carrying information. A buyer reading their eleventh AIOps site does not compare 93 percent against 43 percent. They discount the entire class of claim and go looking for something they can verify.
At that point the only asset with any value is a smaller number they can trace. Which is the same argument, in a different domain, as the one we make about execution: an Action Ticket is trusted because its record can be reconstructed a year later, not because the model was confident at the time. Evidence works the same way. Provenance beats magnitude.
Use this on us, and on everyone else
If you take one thing from this page, take the four questions rather than our numbers. Put them to any vendor in this category, including us.
- Which single deployment produced this figure? If the answer names more than one, you are looking at a composite.
- What was the baseline, and who measured it? A percentage improvement with no stated denominator cannot be checked and should not be weighed.
- Over what window, and how long after go-live? Six months into a deployment and six weeks into one are different claims wearing the same number.
- Which figures on this page are modelled? Every vendor has some. The informative thing is whether they will tell you which without being pushed.
Frequently asked questions
Why publish an evidence policy at all?
Because in this category the numbers have stopped carrying information. When every vendor page claims a large percentage improvement and none of them says from which deployment, over what period, against what baseline, a buyer learns to discount all of them equally. Publishing the standard is the only way to make our numbers mean more than the average, and it costs us nothing except the numbers we were not entitled to use.
What is the difference between a production pattern and a platform capability?
A production pattern has been deployed, measured and published, and the use case carries a link to the case study containing the measurement. A platform capability is something the platform does, modelled onto a scenario we have not yet measured with a named customer. Of the 38 use cases on this site, 4 are labelled production pattern. The other 34 are labelled platform capability, and every timeline on them is marked illustrative rather than measured.
Why are your headline numbers smaller than your competitors?
Partly because they come from one deployment rather than the best figure from each of four, and partly because a 43 percent MTTR reduction measured against a documented baseline is a different kind of object from a percentage with no stated denominator. We would rather publish the smaller defensible number. If a competitor number is both larger and traceable to a single named deployment with a stated baseline, that is a fair comparison and we will lose it honestly.
How can I check any of this?
Every headline figure on this site links to the case study it came from, and every case study states the industry, the geography, the estate size, the measurement window and which of its figures are measured rather than indicative. If you find a number anywhere on this site that does not resolve to a single deployment, tell us and we will either source it or remove it. That offer is the policy.