The enterprise sales team has done what management asked. New enterprise ARR is up 25%, the dashboard is green and the reward logic has fired. The figure is illustrative, not a recommended target.

Elsewhere, the operating picture has become less comfortable. Discounts are deeper. A few deals have been pulled into the quarter. Some customers are a weak fit. Implementation has inherited commitments that are difficult to deliver with the available capacity. The sales result is real, yet it may coexist with weaker contribution economics, retention quality or delivery performance.

Calling this cheating would skip the more useful diagnosis. A performance system should survive contact with a capable employee who understands the rules and wants to succeed under them. If such an employee can maximise the measured result while making the underlying strategy worse, the organisation has designed a behavioural contract it probably did not intend.

The first test is therefore simple: can the headline target reach 100% while the strategic outcome deteriorates?

Place the OKR inside the performance system

An Objective and a set of Key Results can create focus, but they sit inside a longer chain from strategy to behaviour. The distinction matters because each element has a different job.

The Objective describes the strategic change that matters. An outcome KPI or Key Result observes whether that result is moving. A driver metric represents a mechanism management currently believes may influence the outcome. An initiative is work the team chooses to do. A target attaches an expected level or direction to a measure. A health or guardrail metric watches a dimension that should not be consumed while the focal result improves.

Locke and Latham’s review of goal-setting research supports a bounded claim: goals can affect motivation and performance, with effects shaped by factors such as task complexity, ability, commitment and feedback. That evidence does not establish that a branded OKR implementation will improve every organisation.

Kaplan and Norton, writing in the Balanced Scorecard tradition, argue for measures derived from strategy and for combining outcomes with performance drivers. A driver, however, remains a hypothesis until the relationship has been tested. More enterprise demos may indicate stronger demand, or simply weaker targeting. A larger pipeline may reflect better market coverage, or looser qualification.

For the SaaS company, the working model might be: move up-market while protecting durable economics and delivery quality; accelerate enterprise growth; observe new enterprise ARR; watch qualified pipeline, win rate, sales-cycle quality and implementation readiness; invest in account-based marketing, enterprise sales enablement, security readiness and solution engineering; keep retention quality, contribution economics and implementation capacity visible.

That map is still relatively safe. The behavioural problem appears when one element in it starts carrying material consequences for pay, status or resources.

Consequences change what the metric is used for

Suppose a large share of the sales reward now depends on enterprise ARR. The team can improve the result through better selling, but the design also changes the relative attractiveness of other actions. A deeper discount can remove friction. A deal expected next quarter can be pulled forward. Qualification can loosen. Implementation commitments can be made before delivery capacity has caught up. None of those moves has to be fraudulent for the current-period number to improve faster than the economics behind it.

Performance-governance loop showing Measures flowing through Target and Incentive to Behaviour, which can lead to Strategic Outcome or Wrong Victory; Guardrails and Review oversee the mechanism, while Feedback and Reset returns a wrong victory to Measures.

This is where four nearby numerical concepts stop being interchangeable. A baseline anchors the current position. A target states an intended result. A trigger tells management when review or escalation should begin. A guardrail protects another dimension while the focal target is pursued. Calling all four “KPIs” hides the decisions they are meant to support.

Research provides several bounded reasons to take the behavioural mechanism seriously. Bevan and Hood document gaming and measurement distortion in an English public health-care target regime. That setting cannot establish how often the same behaviour occurs in technology companies; it does show a concrete target regime reshaping behaviour around its measurement rules.

In budgeting and compensation, Jensen describes incentives for sandbagging and information distortion when negotiated targets are mechanically tied to rewards. The claim does not extend to “metric-linked pay is bad”. It gives management a reason to inspect what a consequential proxy omits before allowing the formula to settle the decision.

A broader review by Franco-Santos, Lucianetti and Bourne, covering 76 empirical studies, finds effects on behaviour, organisational capabilities and performance that vary with the design and context of the measurement system.

A support team paid heavily for tickets closed within 24 hours illustrates the same mechanism at smaller scale. Faster closure can coexist with more repeat contacts or unresolved issues. The useful question is not whether the SLA should disappear. It is what a team can sacrifice while still satisfying it.

For the SaaS company, that is enough to move the question away from intent. The operating issue is whether the reward system has made short-term ARR improvement rational even when contribution economics, retention quality or delivery capacity absorb the cost.

Stress the target before somebody else does

Before a consequential target goes live, give it to an imaginary operator who is excellent at the job and excellent at reading rules. Ask how that person would maximise the metric without breaking the stated rules.

For enterprise ARR, several paths become obvious:

  • discount harder to reduce purchase friction, leaving contribution economics weaker;
  • pull future deals into the current quarter, improving timing in one period at the expense of another;
  • accept marginal customer fit, raising booked ARR while increasing future retention or support risk;
  • promise implementation scope ahead of delivery capacity;
  • concentrate on the booked number because renewal and expansion quality arrive later.

The exercise is diagnostic. It does not predict that every sales team will take these routes. It identifies what the target fails to protect.

That distinction changes the next design step. A discounting failure points towards contribution economics. Poor-fit wins make retention or qualification quality relevant. Repeated over-selling exposes implementation capacity as a constraint that the sales target is currently allowed to ignore.

Design governance around the failure surface

Keep enterprise ARR as an important outcome measure. Then prevent it from acting as the sole certificate of success.

The first layer is small and specific. Contribution economics can guard against growth purchased with excessive discounting. Retention quality can expose weak-fit acquisition. Implementation backlog or capacity utilisation can surface delivery debt created by selling ahead of the organisation. There is no universal guardrail value for any of these; the boundary depends on the business model, risk appetite and strategic stage.

The metric also needs an owner and a method. The SEC’s guidance on KPIs and metrics in MD&A concerns disclosure, not internal OKR practice, so it should be used only as a governance analogy. Its emphasis on calculation, usefulness, management use and changes in methodology is a helpful reminder that a significant metric needs context. Internally, that means being explicit about calculation, population, period, role in the strategy chain, data ownership and the conditions under which the metric itself should be reviewed.

Control then extends beyond the dashboard. Robert Simons’ Levers of Control distinguishes belief, boundary, diagnostic and interactive control systems. It is one influential framework rather than a universal taxonomy. In this example, a rule against selling implementation commitments beyond agreed delivery capability can operate as a boundary. ARR and its guardrails can be diagnostic signals. The question of whether the move-up-market strategy still makes sense belongs in recurring management attention rather than in a formula.

Governance also needs to know what happens when the signals conflict. A review trigger can specify when a green ARR result must be challenged. Named owners can determine which evidence enters that review and who has authority to hold, adjust or override a formulaic reward outcome.

The UK Corporate Governance Code 2024 provides one bounded example from UK listed-company remuneration governance. It requires remuneration arrangements to allow discretion over formulaic outcomes, while related FRC guidance discusses the remuneration committee exercising judgement. The context is specific, but the design point travels: if ARR shows 110% attainment while delivery capacity has been materially overdrawn, a governance system should have a legitimate way to examine that conflict rather than pretending the formula settled it.

Adding more measures is an easy way to obscure the problem again. The UK public-sector FABRIC framework describes performance information as focused, appropriate, balanced, robust, integrated and cost-effective. It is not a mandatory corporate KPI framework. Here it is useful as a constraint on metric sprawl: information should improve decisions, not accumulate because every failure mode has acquired its own dashboard tile.

For this SaaS company, the redesigned arrangement is therefore modest: ARR remains focal; a few guardrails protect the dimensions exposed by the stress test; metric ownership and review conditions are explicit; boundaries cover unacceptable ways of winning; and formulaic attainment can be challenged through a defined evidence path.

Know when to reopen the design

The redesigned system can still age badly.

Bourne and colleagues treat updating as part of the design and use of performance-measurement systems. That matters because the strategy may change, the customer mix may shift, the data may be redefined, or people may simply learn how to operate more effectively within the current rules.

A reset should be evidence-led rather than calendar-led. For the enterprise-growth system, management should reopen the design if enterprise growth stops being the central strategy, if the existing guardrails cease to represent the customer mix, if a supposed driver repeatedly fails to predict or explain the outcome, if the same optimisation failure returns despite controls, or if a change in metric definition or data source breaks comparability.

None of those conditions implies a universal monthly, quarterly or annual review cycle. They specify what would invalidate parts of the current model.

Put the next target through the seven-field test

The enterprise ARR case can now be reduced to a working decision interface. Before a consequential target enters a scorecard, bonus plan or performance review, fill in all seven fields.

FieldQuestionEnterprise ARR example
Strategic outcomeWhat real result is supposed to improve?Sustainable enterprise growth rather than only this quarter’s booked ARR
Focal measure and roleIs the number an outcome, driver, activity proxy or guardrail?New enterprise ARR is an outcome proxy; it does not represent the whole strategy
Behaviour made rationalWhat becomes attractive when the number is consequential?Discounting, pull-forward, relaxed qualification, oversold delivery
Sacrificial dimensionWhat can deteriorate while the focal number improves?Contribution economics, retention, delivery capacity, timing quality
Guardrail / counter-metricWhich dimensions must remain protected?Retention quality, contribution economics, implementation capacity
Review / escalation / override ownerWho can challenge a green result, and when?A defined cross-functional path across sales, finance and customer/delivery owners
Reset conditionWhat evidence would force the system itself to change?Strategy shift, failed driver hypothesis, repeated gaming signal, structural customer-mix change

The illustrative 25% target may survive that test, or it may not. Nothing here is sufficient to set a target for a real company. The practical output is the governance around it: management can see which behaviours the target makes rational, what can be sacrificed without turning the headline red, who is allowed to challenge the apparent win, and what evidence would force the design to be reopened.

References