Quality & Support
Support that reduces
the tickets it receives
A support contract that measures itself on tickets closed has the wrong incentive. Ours is measured on incident recurrence and on the size of the backlog of known defects, both of which should fall every quarter.
Overview
24/7 cover with named engineers, published response times and a shrinking incident count.
Application support divides into three kinds of work, and confusing them is why most contracts feel expensive. There is keeping the lights on: monitoring, incident response, patching. There is corrective work: fixing what the last release broke. And there is small change: the twenty-hour requests that never justify a project.
We price and report those three separately. A monthly report shows incidents by root cause, mean time to restore, patch currency, and how much of the retainer went on preventable failure. When the same cause appears twice, the fix goes into the next sprint as engineering work rather than into the runbook as a workaround.
Cover is provided by named engineers who learn your system, rather than by a rota of whoever is on shift that night. That costs more per hour and less per incident, and it is the difference between a first response that acknowledges a ticket and one that has already read the logs and warned the downstream teams.
- 14 min
- median P1 acknowledgement across 2,700 tickets last year
- -58%
- recurring incidents at Pilbara Freight over four quarters
- 99.94%
- availability on the supported estate at Cordage Bank
Capabilities
What this covers
Six areas we staff properly. If your problem sits outside them, the honest note at the foot of this page says so.
Tiered incident response with published SLAs
P1 acknowledged in 15 minutes and worked continuously until service is restored, with severity defined by business impact rather than by which system the alert came from.
Proactive monitoring and alert tuning
Alerts tied to user-visible symptoms rather than to CPU graphs, tuned until every page is actionable. A noisy alert that nobody investigates is worse than no alert at all.
Patch, dependency and certificate management
A tracked upgrade cadence for runtimes, libraries and base images, with certificate expiry monitored months ahead. Most emergency weekends we inherit began as a deferred minor upgrade.
Small change and enhancement delivery
A ring-fenced share of the retainer for changes under about forty hours, delivered on a two-week cadence, so minor requests stop queueing behind the next funded project.
Problem management and defect reduction
Every repeat incident gets a root cause investigation and a scheduled permanent fix. The known-error backlog is published, aged and burned down rather than quietly growing.
Disaster recovery testing
Restores are exercised on a schedule against the documented RTO and RPO, in a clean environment, because a backup that has never been restored is an untested assumption.
Deliverables
What you get
- Service agreement with severity definitions and published response times
- Named engineer roster with documented system knowledge handover
- Tuned monitoring and alerting with actionable-page criteria
- Monthly service report covering root cause, MTTR and patch currency
- Known-error backlog with ageing and scheduled permanent fixes
- Tested recovery runbooks with dated restore evidence
Stack
What we build it with
- PagerDuty
- Prometheus
- Grafana
- Loki
- OpenTelemetry
- Datadog
- Jira Service Management
- Terraform
- Ansible
- Renovate
- Kubernetes
- Velero
Process
How the engagement runs
Two-week increments against a written definition of done. You can stop at any increment boundary and keep everything built so far.
Takeover assessment
Two to three weeks documenting the architecture, access paths, deployment process and current incident history before we accept the pager.
Stabilise
The first quarter targets alert noise, expired dependencies and the incidents that recur monthly, which is usually where the fastest reduction is available.
Shadow and assume
Your team stays on call alongside ours for an agreed period, then we take primary and your team becomes escalation rather than disappearing overnight.
Steady state
Monitoring, patching, incident response and small change run on a published cadence, with a monthly review of causes rather than of ticket volumes.
Reduce and hand back
Recurring causes are engineered out, and the contract is designed to shrink. Where a system should be retired instead of supported, we say so.
When this is the wrong engagement
If the application is due for replacement within twelve months, a full support contract is money better spent accelerating the replacement.
FAQ
Questions we get asked
- Can you support an application you did not build?
Yes, and most of what we support we inherited. The takeover period is two to three weeks of documenting architecture, access and deployment before we accept the pager, and if that assessment finds the system cannot be safely supported as it stands, we say so before signing.
- What counts as an incident versus a change?
Severity is defined by business impact and agreed in writing before the contract starts. Anything that stops users doing their work is an incident; anything asking the system to do something new is a change, and changes under roughly forty hours come from the small-change allocation.
- Is 24/7 cover worth it?
Only if the cost of downtime overnight exceeds the premium, which is true for payments, logistics and clinical systems and false for a great many internal tools. We often recommend business-hours cover with a defined out-of-hours escalation path, and it costs about a third as much.
- What happens if we want to bring support back in-house?
The contract includes a 60-day exit with a documented handover: runbooks, alert configurations, the known-error backlog and a shadow period where your engineers take primary while we stay on escalation. Nothing useful is held in a private wiki that we keep.

