HaveStack Request a meeting
Systems under management

What a response time actually promises.

A service agreement is only as good as the definitions underneath it. This is what HaveStack means by a response time, what happens after an incident, how releases reach production, and what arrives in the client’s inbox at the end of each month.

Applies to
Every system under a maintenance agreement
Cadence
Report each month, access review each quarter
Evidence
Incident record, release log, monthly report

Three words that are not interchangeable.

Most disagreements about a service agreement are really disagreements about definitions, and three terms carry almost all of the weight.

A service level indicator is a measurement: the proportion of requests served under a given latency, the proportion of successful writes, the availability of an endpoint. A service level objective is the target for that measurement, set internally. A service level agreement is the external commitment, with a stated consequence if it is missed.

The order matters. An agreement written before anyone has decided what is being measured commits the practice to a number it cannot verify and the client to a promise they cannot check.

The error budget, and what it is for.

If availability is targeted at anything below one hundred per cent, the difference is a budget. It is the amount of unreliability the objective permits over a period, and it converts an argument into arithmetic.

The use of it is a decision rule. While the budget is healthy, change is cheap and features ship. When it is burning, the work moves to stability until it recovers. That removes the recurring standoff between shipping and reliability by making it a measured condition rather than a matter of who is more persuasive in the room.

Incidents are written up without indicting anyone.

A blameless write up identifies the contributing causes of an incident without attributing it to an individual, on the working assumption that everybody involved acted reasonably given the information they had at the time.

This is not politeness. A process that assigns fault reliably produces incomplete accounts, because the people who know most about what happened have the strongest reason to say least. The organisation then loses the one thing an incident was good for.

Every incident on a HaveStack system produces a written record: what the client experienced, the timeline, the contributing causes, the correction made, and the change that reduces the chance of a repeat. The client gets it whether or not they asked.

  • What the user experienced, before what the system did
  • A timeline with detection, response and resolution times
  • Contributing causes, in the plural, without naming an individual
  • The immediate correction, and the durable change
  • Action items with an owner and a date
  • Sent to the client, not filed internally

Scheduled, never applied unannounced.

The change failure rate is one of the four delivery metrics popularised by the DORA research programme: the proportion of releases that cause a degradation requiring remediation. It is a useful figure because it resists the intuition that shipping less often is safer. In practice the teams with the lowest failure rates are usually the ones releasing in small, frequent, reversible increments.

HaveStack releases into an agreed window, announces the window in advance, and keeps each release small enough to be reversed. A client should never discover a deployment by noticing that something moved.

A monthly report the client can file.

The report exists so that somebody who was not in any of the conversations can understand the state of the system. It is written for a director or an auditor, not for an engineer, and it is a document rather than a dashboard link that will be dead in a year.

It covers availability against the objective, incidents and their status, changes applied, capacity against provisioning, outstanding risks, and what is scheduled next. Where a target was missed, it says so and says why.

Access reviewed each quarter.

Access accumulates. People join a project, get what they need, and keep it after they move on. The risk is not usually malice, it is arithmetic: the set of credentials that can reach a production system grows monotonically unless somebody prunes it.

Every account and integration with access to a maintained system is reviewed quarterly against whether it is still required. Access is withdrawn on the working day someone leaves, rather than at the next review, and the review confirms it happened.

Written down, or it does not happen.

Maintenance that is not written down is a favour, and a favour stops the moment somebody is busy. These are the clauses that make the practice above enforceable.

What goes into the agreement

  • The indicators being measured, before any target is written
  • Objectives per system, and the response times attached to each severity
  • A written incident record for every incident, sent to the client
  • Named release windows, announced in advance
  • A monthly report covering availability, incidents, changes and risks
  • Quarterly access review, and withdrawal on the day someone leaves

Terms used here.

SLI
Service level indicator. A quantitative measurement of one aspect of the service, such as the proportion of requests served under a latency threshold.
SLO
Service level objective. The internal target for an indicator, set before any external commitment is made.
SLA
Service level agreement. The external commitment, with a stated consequence when it is missed.
Error budget
The unreliability an objective permits over a period. It decides when to ship and when to stabilise.
Change failure rate
The proportion of releases that cause a degradation requiring remediation. One of the four DORA delivery metrics.

References.

Published guidance this page draws on. HaveStack does not claim authorship of the practice described; it claims to follow it.

  1. 01
  2. 02
  3. 03
  4. 04

The rest of what stays under management.

Engagement

Start with a meeting.

Maintenance is available on systems HaveStack built and, after a written technical review, on systems it did not. A first meeting runs about forty minutes and establishes what is already in place.

  1. Submit the brief
  2. Reply within two working days
  3. Meeting scheduled
Request a meeting