
SLAs you can actually meet
Most service level agreements are written to win a tender rather than to run a service. The tell is a resolution time promised for a problem whose cause is not yet known.
Response time is when somebody starts work and can be promised honestly. Resolution time depends on the cause and often on a third party, so it should be a target with an escalation path rather than a guarantee.
Set priorities on business impact rather than on technical severity, agree them in advance, and keep the list to four levels. More than four and nobody uses them consistently.
A monthly report is worth reading if it shows the top five recurring causes rather than only ticket counts. Counting tickets measures activity; recurring causes measure whether anything is improving.
Response time and resolution time are different promises
A supplier can honestly promise when work starts. Promising when it finishes requires knowing what is wrong, which by definition nobody does at the moment a ticket is raised.
An SLA that guarantees a four-hour resolution for a priority one incident either has a definition of resolution that includes a workaround, or it has a penalty clause the supplier has priced into the rate and expects to pay occasionally. Ask which. Neither is dishonest, but they are very different products.
Four priorities, defined by impact
| Priority | Definition | Example |
|---|---|---|
| One | Business stops or money is being lost now | Tills down in a shop, a ward system unavailable |
| Two | A department is blocked, or a workaround exists but costs real time | A site offline with 4G failover holding |
| Three | An individual is blocked | A laptop that will not start |
| Four | Request or question, no work is stopped | New starter setup, software request |
Write the examples into the agreement. Definitions get argued about at two in the morning; examples do not.
Agree who declares a priority one
The most common friction in a support relationship is not response times. It is a customer calling something a priority one that the supplier logged as a three, and both being convinced they are right.
Fix it by naming people rather than criteria. A short list on your side who can declare a priority one, a matching list on the supplier side who can accept it, and a rule that disputes are resolved after the incident and not during it.
Say which hours the clock runs
An SLA of four hours means one thing at ten in the morning and something else entirely at five on a Friday afternoon. Every response time needs the cover window written next to it, and the window should match when your organisation actually loses money rather than when offices are typically open.
For a retailer that includes Saturdays. For a distribution centre it includes nights. For an accountancy firm it very likely does not, and paying for cover you do not need is as wasteful as lacking cover you do.
The part nobody writes down: other suppliers
A large share of incidents end up waiting on somebody else: an internet provider, a software vendor, a payment processor. If the SLA is silent on that, every one of those becomes an argument about whose clock is running.
Two clauses fix it. First, define that the clock pauses while a named third party holds the ticket, and that the pause is visible in the report. Second, agree who chases them, because the answer should be your supplier rather than you. That second point is most of the value of service governance: one party owning the whole chain instead of four suppliers each owning their own piece.
What a monthly report should contain
- Volume and trend. Tickets by priority, this month against the last six.
- Performance against the agreement. Response and resolution, with misses explained rather than hidden.
- Top five recurring causes. The most valuable section and the one most often missing.
- What changed. Changes made, and the incidents that followed them.
- Backup and monitoring status. Success rate, oldest restore point, tests performed.
- What we recommend next. Two or three items with a reason, not a wish list.
A report that only counts tickets tells you the desk was busy. A report that names recurring causes tells you whether the underlying problem is being fixed, which is the only question that matters over a year.
Questions we get about this
What comes up when an SLA is being negotiated.
What is the difference between response time and resolution time?
Response time is when somebody starts working on the ticket, which a supplier can promise honestly. Resolution time depends on what is wrong and often on a third party, so it belongs in an agreement as a target with an escalation path rather than as a guarantee.
How many priority levels should an SLA have?
Four. Fewer and everything important collapses into one level; more and nobody applies them consistently. What matters more than the number is that each level has a written example from your own organisation next to it.
What if the delay is caused by another supplier?
Agree in advance that the clock pauses while a named third party holds the ticket, that the pause is visible in the monthly report, and that your supplier chases them rather than you. Without that clause, multi-vendor incidents turn into arguments about whose responsibility it was.
Should an SLA have penalty clauses?
They are less useful than they look. A penalty compensates you slightly for a bad month but does nothing to prevent the next one. A monthly review that names recurring causes and tracks whether they are being fixed changes behaviour far more than a credit note does.
Where this lands in our work
Where the agreements live.
Want this looked at for your own sites?
Half an hour on a call is usually enough to tell you whether we are the right party for it, and we will say so if we are not.