Skip to main content
Image coming soon

The MSP Network Operations Runbook for Multi-Client SLAs

$199.00
Adding to cart… The item has been added

A focused course, tailored for you

The MSP Network Operations Runbook for Multi-Client SLAs

A written course and implementation playbook for running a multi-tenant NOC where every client thinks they are the only one and the ticket queue still closes on time.

An after-hours alert fires, the on-call engineer cannot tell which client is affected without three terminal windows, and the SLA clock is already running. The cost of that confusion is measured in the next quarterly business review, not the ticket.

$199 one-time
Tailored to your situation. Access within 24 hours. 30-day money-back.

Includes a hand-built implementation playbook delivered alongside course access, generated for your specific situation.

Why this course

Running network operations for a portfolio of clients is not the same job as running a single enterprise network. The same alert engine has to route to the right rotation, label the right tenant, escalate to the right account manager, and produce a report the client recognises as theirs. The same change window has to respect the client who runs retail and freezes for the holiday, the client who runs healthcare and freezes for open enrolment, and the client who runs logistics and freezes for absolutely nothing. The same four-person team has to staff that without the senior engineer becoming the single point of recovery. Most MSP NOCs grow into this shape one client at a time and the runbook lags the reality by quarters. The result is engineer burnout, missed SLA credits that get written off quietly, and clients who start shopping during their renewal cycle because the QBR deck never quite explains why their network is fine when their dashboard says otherwise. The fix is operational, not technological. The monitoring stack is already there. What is missing is the discipline layer on top of it.

What you walk away with

  • A tenant-aware alert and escalation map you can apply to your current monitoring stack within the first week of work.
  • A four-to-six person on-call rotation design that does not collapse when one person is sick or takes leave.
  • A change-window policy that survives clients with conflicting freeze calendars and emergency cutover requests.
  • A quarterly business review pack the client engages with instead of skipping past the first three slides.
  • A documented MTTA and MTTR baseline per client that gives the next contract negotiation a defensible position.

The 12 modules

Module 1. The Multi-Tenant Alert Map
Walk through tenant tagging at the monitoring layer so an alert at 03:42 names the client, the circuit, the responsible engineer, and the escalation path in the first line. Includes the tag taxonomy that does not break when a client gets acquired and the alert-suppression rules that prevent the same circuit from paging four times for one incident. Worked example covers a portfolio of fifteen mixed clients.
Module 2. On-Call Rotation Math for Small Teams
The rotation design problem for a four-to-six person NOC where one engineer is also the architect and another is part-time on a client project. Covers shift length, handover format, weekend coverage that does not eat Monday productivity, and the secondary escalation that keeps the senior engineer off the primary phone three weeks out of four. Includes the rotation calculator template.
Module 3. Tiered Response Without Tier Inflation
How to structure tier one, tier two, and tier three response so the right engineer touches the ticket at the right time and the tier-one queue does not silently become a routing layer. Covers the runbook depth that lets a tier-one engineer close eighty percent of common incidents and the explicit hand-up criteria that prevent endless escalation. Includes worked runbooks for the seven most common MSP network incidents.
Module 4. Change Windows Across Conflicting Freeze Calendars
The policy and the negotiation language for a change-window calendar that serves a retail client who freezes November to January, a healthcare client who freezes October to December, and a logistics client who never freezes. Covers the standing-window approach, the emergency-cutover policy, and the change-advisory-board format that closes a decision in twenty minutes instead of two meetings. Includes the policy template.
Module 5. SLA Definition That Survives an Audit
Define availability, mean time to acknowledge, mean time to repair, and credit triggers in language that the client legal team accepts and the NOC can actually measure. Covers the difference between measured availability and reported availability, the exclusion clauses that are defensible versus the ones that erode trust, and the credit cap structure that does not bankrupt the contract when something goes wrong. Includes the SLA schedule template.
Module 6. The Per-Client Operational Baseline
Build a per-client baseline document that captures the circuit inventory, the vendor contacts, the change history, the recurring incident patterns, and the accountable parties on the client side. Covers the keep-it-current discipline that prevents the baseline from going stale within two quarters and the handover format that lets a new engineer pick up a client in under a day. Includes the baseline template and the quarterly refresh checklist.
Module 7. Vendor Escalation Without Burned Bridges
Carrier escalation, hardware vendor RMA, and ISP back-channel relationships are the difference between a four-hour outage and a fourteen-hour outage. Covers the contact map you maintain per vendor per region, the escalation script that gets a real engineer on the call inside thirty minutes, and the post-incident vendor scorecard that informs the next procurement cycle. Includes the vendor contact template.
Module 8. The Quarterly Business Review That Lands
QBR decks that open with availability percentages lose the client attention by slide three. Covers the QBR structure that leads with the incidents the client actually felt, the trend lines that justify the next investment, and the recommendation section that gives the client CIO something to walk into their own board with. Includes the QBR deck template and the talking-track for the two most awkward client questions.
Module 9. Incident Postmortems Without Blame
Post-incident reviews that the team actually engages with and the client actually accepts. Covers the timeline-reconstruction discipline, the root-cause taxonomy that distinguishes process failure from vendor failure from human error, and the action-item tracking that closes loops instead of accumulating a backlog of unread postmortem documents. Includes the postmortem template and the client-facing summary format.
Module 10. Monitoring Stack Hygiene
The monitoring stack accumulates dead checks, false positives, and orphaned dashboards faster than anyone has time to clean up. Covers the quarterly hygiene sweep that retires the dead checks, the false-positive triage process that reaches a defensible alarm rate, and the dashboard rationalisation that gives each client one dashboard their team can read without a tour. Includes the hygiene checklist.
Module 11. Capacity and Renewal Conversations
The conversation about a client outgrowing their current circuit, their current contract, or their current support tier is the conversation that decides whether the next renewal is upsell or churn. Covers the data the NOC owes the account team a quarter before the renewal lands, the framing that makes the upsell sound like a recommendation rather than a sales push, and the timing that catches the budget cycle on the client side. Includes the renewal-brief template.
Module 12. Building the Operational Layer Without a Project
Implementation in a live operating NOC that cannot stop running for a transformation programme. Covers the sequencing that delivers the runbook in two-week increments, the engineer time budget that does not destroy the ticket queue, and the client communication that lets the improvements show up in their QBR without a meeting about the meeting about the meeting. Includes the four-quarter rollout plan.

How this addresses your situation

Specific modules that map to what you said you are dealing with.

Module 1 and 2 apply when the after-hours page lands on the wrong engineer or names the wrong client.
Module 4 and 5 apply when a client books a cutover during another client's freeze window or disputes an SLA credit.
Module 8 and 11 apply when a QBR ends with the client asking why their dashboard says one thing and your report says another.
Module 6, 7 and 10 apply when a new engineer joins and needs six weeks to be useful on a client they have never seen.

What you get with this course

  • Twelve written modules with downloadable templates and worked examples for every module.
  • The per-buyer implementation playbook hand-built against your actual client mix, ticket volume, and current monitoring stack.
  • Editable templates for the alert tagging schema, on-call rotation calculator, change-window policy, SLA schedule, per-client baseline, vendor contact map, QBR deck, and incident postmortem.
  • A four-quarter rollout plan that fits a live operating NOC without a transformation project.
  • Thirty-day satisfaction guarantee.

What you will have in hand by Day 1, Week 1, Month 1

Within 24 hours your account in the learning environment is provisioned and the tailored implementation playbook is delivered alongside it.

Weeks one and two: complete modules one through four, deploy the tenant-aware alert map and the on-call rotation.

Weeks three through six: complete modules five through eight, refresh SLA schedules and run the next QBR with the new pack.

Quarter two: complete modules nine through twelve, embed the postmortem and renewal disciplines into the operating rhythm.

Before and after

Before

Multi-tenant alerts that confuse the on-call engineer for the first ten minutes of every after-hours incident. A rotation that depends on one senior engineer always being available. QBRs that the client politely skips through. SLA credits written off because nobody wants to argue the measurement methodology.

After

Alerts that name the client and the circuit on the first line. A rotation where the senior engineer is off the primary phone three weeks out of four. QBRs the client engages with. SLA measurement that holds up to audit and protects the contract value.

What happens if you do not address this

The cost shows up in the next renewal cycle. Clients who do not feel served during incidents shop their network contract when the term comes up, and the procurement team uses the QBR pack as evidence in the conversation. Internally, the cost shows up as senior-engineer attrition. The engineer who carries the rotation longer than anyone else is the engineer who eventually leaves for a vendor role with predictable hours.

Who it is for

A network operations lead, senior network engineer, or NOC manager inside a managed services provider supporting between five and forty client networks. Comfortable with the underlying technology, accountable for SLA performance, on the hook for the client conversations when something slips, and short the time to design the operational layer from scratch. Often the person who built the current rotation and is now the person who cannot take a clean week off because the rotation depends on them being available as backup.

Who this is NOT for. Not for enterprise network engineers running a single corporate network with no external SLA. Not for resellers who do not operate the network after handover. Not for executives looking for a strategy deck without operational depth.

How it arrives

Text-based course in the Art of Service learning environment, plus downloadable templates and worked examples for every module, plus the hand-built implementation playbook delivered alongside course access.

Time investment. Around six to eight hours total reading across the twelve modules, plus the implementation work the playbook sequences across your operating quarters. Most operators complete modules one through four in the first two weeks and apply the rest as the operating cadence allows.

Why $199 is the right number

Generic NOC operations courses on the open web cover monitoring tools and incident response but assume a single enterprise network. Vendor-sponsored MSP training covers the vendor product but skips the operational layer entirely. A senior consultant engagement gets the runbook designed but costs five figures and stops the day the consultant leaves. This course gives the operator the runbook design, the templates, and an implementation playbook tuned to the actual client portfolio, for the price of a monitoring stack add-on.

FAQ

We already use a monitoring platform. Does this replace it?
No. The course is the operational layer that sits on top of whatever monitoring stack you run. The templates are tool-agnostic and assume you keep your current platform.
How does the playbook get tuned to our client mix?
After purchase a short intake captures your client count, sector mix, current SLA structures, team size, and monitoring platform. The playbook is hand-built against those inputs and delivered alongside course access.
Is there a money-back guarantee?
Thirty days. If the materials do not fit the operation, the purchase is refunded in full.
Do we need to take the team through this?
The course is designed for the operations lead or NOC manager. The templates and rollout plan are what get rolled out to the team. The course itself is read by the one or two people accountable for the operational design.

30-day money-back guarantee. If after a week of working through the materials this is not what you needed, reply to the receipt email and a full refund is processed. No questions, no forms.

Within 24 hours your account in the learning environment is provisioned and the tailored implementation playbook is delivered alongside it.