Two dashboards can show “NRR: 110%” and still be describing different things.
That is the useful starting point for KPI selection. Amplitude says in its 2025 Form 10-K that there are no uniform standards for calculating many of its key business metrics, so similarly titled measures reported by different companies may not be comparable. Amplitude 2025 Form 10-K
The implication reaches beyond SaaS. A finance team, a subscription business, and a product team can all have perfectly sensible metrics while answering different questions. Before copying any of them, I want to know what decision the number is supposed to change.
For Finance / FP&A, that might be cash, forecasting, profitability or resource allocation. A subscription team may be diagnosing recurring revenue, retention and expansion. Product teams are often closer to activation, engagement and repeated value delivery. One company can need all three views at once.
A metric is ready only when someone else can reproduce its meaning
The useful object is not the dashboard label. It is the operating definition behind it.
Call that a metric contract. It should tell a second analyst what the metric is for, how it is calculated, who or what is eligible, which cohort is included, the unit of analysis, the time window, the underlying system, any exclusions or adjustments, the source of truth, and who owns changes to the definition.
That sounds administrative until two systems disagree. At that point the contract is the difference between investigating a business movement and arguing about whose query is “right”.
There is also a strategic benefit. Starting with the decision makes it easier to decide whether a number is an outcome, a driver, a guardrail or merely a diagnostic signal. A metric that cannot change a decision may still be interesting, but it has a different job.

SaaS makes the definition problem visible very quickly
Amplitude’s ARR is a point-in-time operating measure based on subscription agreements. The company explicitly says that ARR is neither annualised U.S. GAAP revenue nor a forecast of revenue. Amplitude 2025 Form 10-K
Both ARR and revenue can therefore be useful without being interchangeable. One describes recurring subscription run-rate under the company’s definition; the other follows accounting recognition.
NRR needs even more context. Amplitude calculates it from an existing-customer cohort, includes expansion, nets contraction and attrition, excludes new customers and overage charges, and presents trailing-twelve-month NRR using period-end measures in an averaging process. Amplitude 2025 Form 10-K
A bare 110% does not tell you any of that.
Before comparing it with another company, I would want the cohort definition, treatment of new customers and overages, contraction/churn rules, and whether the reported number is a point-in-time observation or an average. Box makes the same broader comparability caution around its company-defined operating and non-GAAP measures. Box FY2026 Form 10-K
This is why SaaS benchmarking gets shaky long before the arithmetic gets difficult.
Product analytics adds a second source of drift: instrumentation
In Google Analytics 4, Total Users and Active Users are distinct metrics. Active User status depends on specified engagement or event conditions rather than simply counting everyone who appears in the dataset. Google Analytics: Understand user metrics
The engagement layer has its own rules too. GA4’s Engagement Rate is engaged sessions divided by total sessions, with an engaged session defined through duration, a key event, or page/screen-view criteria. Google Analytics: Engagement rate and bounce rate
Amplitude uses a different default convention: a user who logs at least one event in the chosen interval is generally active, and teams can mark events as inactive so that those events no longer make a user active. Amplitude Docs: Helpful definitions
Now imagine Active Users fall by 8% after an instrumentation clean-up. Treating the movement as a product-engagement decline before checking event taxonomy would be a category error. The data pipeline may have changed what “active” means.
The North Star idea has a related boundary. Amplitude ties a good North Star to customer value, product strategy and the product’s core engagement model. Amplitude: Every Product Needs a North Star Metric That makes it a useful design heuristic, not a licence to transplant one company’s preferred metric into another product.
Finance has the same problem in a different costume
Free cash flow looks familiar enough that people often skip the definition.
Box’s non-GAAP free cash flow begins with operating cash flow and deducts a company-defined set of items including net capital expenditure, finance lease principal payments and capitalised software costs. It reconciles the measure back to operating cash flow. Box FY2026 Form 10-K
The label “FCF” therefore does not reveal the adjustment set.
SEC guidance on non-GAAP financial measures reinforces the need for that visibility. Where applicable, the measure should be quantitatively reconciled to the most directly comparable GAAP measure, with enough detail to understand the reconciling items. SEC: Non-GAAP Financial Measures C&DIs
IFRS 18 operates in a different regime, but it also creates disclosure and reconciliation requirements for qualifying management-defined performance measures. IFRS 18 Presentation and Disclosure in Financial Statements
The two regimes should not be conflated. For an analyst, however, they point to a common discipline: custom performance measures need visible definitions and adjustment logic if they are going to travel across teams, periods or companies.
Benchmarks need an equivalence check before they become targets
I would not reject an external benchmark simply because it is imperfect. I would classify it.
A benchmark is much stronger when seven things line up: the formula, the population or cohort, the time window, the business model and segment, the measurement or accounting system, exclusions and adjustments, and the benchmark’s date and provenance.
When several of those are unknown, the number can still be context. It should not quietly become a target.
This distinction matters because “What is a good NRR?” sounds like a request for one number. In practice it is a request for a comparable population and a compatible metric definition first.
The meeting version
If someone proposes a new KPI tomorrow, I would ask these questions before adding another tile to the dashboard:
- What decision will this number influence?
- Which economic or behavioural system does the question belong to?
- Is the metric an outcome, driver, guardrail or diagnostic signal?
- Can we state its formula, eligibility rules, cohort, time window and source?
- Can another analyst reproduce it from the same source of truth?
- Is the external comparison genuinely equivalent, or only directionally useful?
- Who owns the action when the number moves?
The last question is especially revealing. A metric can be carefully defined and still be operationally useless if nobody can act on it.
There is no permanent “best KPI” set for Finance, SaaS or Product. What can be made durable is the method for deciding whether a particular metric is fit for a particular decision, under a definition everyone can inspect.