Skip to content
CAMPUX Cloud Bootcamp Phase Four · Class Thirty-Six
Phase Four — Operate, Secure & AI
Reading 18 min · Drills 4 · A six-part track
Aligned to AZ-400 and AZ-104
Class Thirty-Six

When it breaks, and how you get it back

Something will fail on a Saturday in November; this track is the two disciplines that decide whether that costs nine hours or twenty-three minutes — responding to the incident in front of you, and recovering the data and service behind it, both rehearsed before you need them.

§1

Two disciplines that share one bad Saturday

Everything you have built in this bootcamp assumed the good day: the deploy that worked, the cluster that scaled, the bill that stayed sane. This track is about the other day — the one where checkout returns a blank page at eleven-fifteen on the busiest Saturday of the year, and the room turns to look at you. There are two jobs to do when that happens, and they are not the same job. One is to respond: stop the bleeding, keep people informed, and learn from it without a witch-hunt. The other is to recover: get the data and the service back from wherever they went, on a timetable the business agreed to in advance. This class teaches both, because on the day they arrive together.

Standalone bootcamps split these across two courses and leave learners to discover, in production, that they are one skill. The engineer running the incident is the same engineer who decides whether to fail over to the paired region; the postmortem that closes the response is the same document that resets the recovery targets for next time. So this track keeps them under one roof and connects them at the exact seam where they meet — the two numbers, RPO and RTO, that a postmortem produces and a recovery plan spends. Learn the response half and you can run a crisis. Learn the recovery half and you can end one. Learn them together and you become the person a company will not operate without.

One half stops the bleeding. The other brings it back.

§2

The shape of the whole thing, drawn once

Before the six parts, hold the picture they all hang off. A disaster is a single fixed moment. Two lanes run away from it.1 The respond lane runs forward in the next minutes and hours: you detect the fault, triage what changed, mitigate to restore service, communicate on a cadence, and — days later — hold a blameless postmortem. The recover lane answers a different question that the same moment forces: is the data intact, and is the service still somewhere? If a disk or a database is gone, you restore it from a vault; if a whole region is gone, you fail over to the copy waiting in another one. How much data you may lose and how long you may be down are the two numbers — RPO and RTO — that size every choice in that lane.

Fig. 1 · Two lanes out of one disaster — the six parts, placed
Two lanes run from one disaster: an upper respond lane (parts 36a-c) and a lower recover lane (parts 36d-f), sized by RPO and RTO. the disaster · 11:15 Respond — parts 36a·b·c detect triage mitigate communicate,then postmortem Recover — parts 36d·e·f restore from a vaultAzure Backup fail over a regionSite Recovery the BCDR plantiers, pairs, zones RPO & RTO size everything in this lane
Copy this once by hand. The top lane is parts 36a–36c and it is a procedure you run while frightened; the bottom lane is parts 36d–36f and it is a plan you priced while calm. The red line at the left is the disaster itself — one moment, two lanes running away from it. The seam is written at the end: the postmortem that closes a response sets the RPO and RTO the recovery plan spends. Interviewers ask you to place these, not recite them.
§3

The track — six parts, alert to all-clear

The rest of Class Thirty-Six is six pages, each a full class in its own right with its own drills, situations, and lab notes. Read them in order the first time; the recovery half leans on vocabulary the response half sets, and the two numbers introduced at the seam run through everything after. Later they work as reference — the page you reopen the week you are actually on call, or the afternoon you finally write the BCDR plan.

36a · Fundamentals & detection
What counts as an incident, how severity is set, and the loop you run under pressure — plus the signals that tell you something broke before a customer does, and what being on call actually asks of you.
36b · Triage & mitigation
The one question that finds most outages — what changed? — read against the Activity Log and the estate, and the discipline that saves the day: mitigate first, root-cause second. Rollbacks, feature flags, and the safe fast fix.
36c · Communication & postmortems
The update cadence that keeps leadership out of the incident channel, and the blameless postmortem that turns an outage into a system fix instead of a name. Where the response half hands the recovery half its two numbers.
36d · Azure Backup
Where recovery points live: the Recovery Services vault and the Backup vault, backup policies and retention, restoring a VM or a file — including across regions — and the soft delete that survives a hostile administrator.
36e · Azure Site Recovery
The standby in another region: continuous replication, the full drill of test failover, failover, reprotect, and failback, and why a failover you have never rehearsed is a hope, not a plan.
36f · BCDR strategy
The judgement no service configures for you: region pairs and zone redundancy, sorting workloads into honest tiers, and writing the BCDR plan — including, in writing, what you chose not to protect. Carries the lab.
§4

The one habit that carries the whole track

Beginners meet an incident by improvising and meet recovery by hoping. The engineers who are calm in both have internalised a single habit: decide the hard things while nothing is on fire, and rehearse them until they are boring. The severity levels, the comms cadence, the rollback command, the failover runbook, the RPO and RTO each tier is owed — all of it is cheap to settle on a Tuesday afternoon and impossibly expensive to invent at three on Sunday morning. The whole track is, in a sense, one long argument for doing the thinking in advance and leaving only the execution for the bad day.

This is why the response half gives you a loop you can run frightened, and why the recovery half makes you test the restore and the failover before you trust them. A backup you have never restored is a rumour; a failover you have never run is a slide. The senior signal in this whole domain is not heroics during the outage — it is the visible, unglamorous evidence that you removed the need for heroics beforehand. Learn that disposition here and the six parts stop being a checklist and become a temperament.

Case File · Campux Retail

The nine-hour outage that started all of this

the founding wound, finally addressed at both ends

Before this bootcamp existed, Campux was dark for nine hours when a single disk failed in the server closet and the only backup had never been tested. That outage is the reason the reader — Campux's first cloud engineer — has a job at all. This class is where it stops being possible from both directions. The response half gives Campux a loop so the next fault costs minutes of confusion, not hours; the recovery half gives it recovery points, a paired-region standby, and a plan that prices each workload's protection against what losing it actually costs.

Across the six parts, Campux earns the sentence that renews an engineer's budget every year: the last unplanned outage cost us nine hours; the last incident cost us twenty-three minutes and no data.2 The disk in the closet took whatever it wanted. Nothing since has been allowed to — and the drills below rehearse the judgement that keeps it that way.

§5

Why this track pays for the year

Incident response and recovery are the least glamorous work a cloud engineer does and the work that most reliably keeps them employed. Nobody is promoted for the outage that photographs well; people are trusted with more when the outage that could have been a headline was a footnote instead. The hiring signal is specific and rare: you can describe the loop without notes, you name mitigation before root cause, you run a postmortem that produces a fix and not a culprit, and you can price resilience in the two numbers a business actually signs. That combination — calm under load plus planning before it — is what a hiring manager is straining to detect and what the drills in this hub rehearse.

So treat the six parts as an investment with a legible return. Each one is a paragraph you can say in an interview and a task you can do on the job, and together they move you from "I hope nothing breaks" to "I know exactly what we do when it does, and I have tested it." The parts build the hands; the hub, below, rehearses the judgement they serve.

On the job

The engineer the outage does not rattle

You · Cloud Engineer · the pager has just gone off

Checkout is down and the channel is filling with questions. You are calm, not because you are brave, but because you decided all of this on a quiet Tuesday: the severity is set, the loop is running, the update goes out on the clock, and if the data is hurt you know which vault to open and which region to fail into — because you tested both last quarter. When it is over you write the postmortem that fixes the system, and reset the two numbers for next time. That temperament, more than any single service, is what this track is for.

Class Thirty-Six · Hub

Examination

Four drills, then two situations. These test the judgement the track is built on; the hands are built in the six parts. The situations have no marking scheme — write your answer before you reveal the reasoning, or the exercise is worthless. Nothing is stored.

Drill 01Recall
This track keeps two disciplines under one roof. Which pair, and what joins them at the seam?
Marked

B. The response half stops the bleeding — detect, triage, mitigate, communicate, postmortem; the recovery half brings the data and service back — backup, failover, a priced plan. They meet at the two numbers: a postmortem sets the RPO and RTO the business will fund, and the recovery plan spends them. D is the trap: backup and Site Recovery are both in the recovery half — they do share a vault, but naming them as the two disciplines drops the entire response side, which is where a real outage begins. Miss that and you are the engineer who can restore a database but cannot run the room while it burns.

Drill 02Recall
The track's central habit, in one line, is:
Marked

B. Severity levels, comms cadence, the rollback command, the failover runbook, each tier's RPO and RTO — all cheap to settle on a Tuesday and ruinous to invent at 3am. A sounds streetwise and gets people hurt: improvisation is what the loop and the runbook exist to replace. D is the specific expensive mistake the strategy part dismantles — protecting analytics scratch data to checkout grade is a finding, not prudence, because resilience is bought in real money. The whole track argues for doing the thinking in advance and leaving only execution for the bad day.

Drill 03Select three
Which three are honest markers of a senior in this domain, as opposed to a hero?
Marked

The loop without notes, the tested restore and failover, and the blameless postmortem. Each is evidence that the thinking happened before the outage and is shared, not hoarded. The two rejects describe the same anti-pattern wearing a cape: the indispensable hero who fixes everything personally and keeps the runbook in their head is a single point of failure the company cannot promote and cannot let take a holiday. A backup you have never restored is a rumour; a failover you have never run is a slide; a runbook only you can read is a liability. Seniority here is the removal of drama, not the supply of it.

Drill 04Spot the error
A new engineer sketches the plan for the six-part track below. One line inverts the discipline the whole track teaches. Which?
# how I'll handle the next Saturday outage
1.  Detect from alerts, set a severity, start the loop.
2.  Find the root cause first; do not touch anything
    until I fully understand why it broke.
3.  Send a leadership update on a fixed cadence.
4.  After it's over, hold a blameless postmortem and
    reset the RPO and RTO for the affected tier.
Marked

Line two. The single most expensive instinct in incident response is the engineer's urge to understand before acting. It inverts the rule the response half is built on: mitigate first, root-cause second. If a rollback or a feature flag restores checkout in four minutes, you do that first and diagnose the failed deploy afterward, at leisure, with the revenue flowing. Insisting on full understanding first means customers stay down for the length of your investigation — which on a November Saturday is priced in thousands of pounds a minute. The other lines are healthy: a fixed cadence keeps leadership out of the channel (C is the anti-pattern, not line three), and resetting the two numbers after a postmortem is exactly the seam that joins the track's two halves (D misreads it).

Situation 01Write before you reveal
Your manager, trimming the roadmap, says: "We already have Azure Backup turned on. Do we really need to spend a quarter on all this incident and DR process? If something breaks we'll just restore." How do you respond?
Do not argue that backup is useless — it isn't. Examine the word "just".
Reasoning

Concede the tool before you question the plan. Backup being on is genuinely good; open by saying so, or you sound like you are defending your own headcount. Then put the weight on the word doing the damage: "just restore." A restore is only the recovery half, and only its slowest tool — it brings back last night's data after however long the restore takes, which is fine for a catalogue and useless for checkout in a regional outage. "Just restore" quietly assumes the data is the only thing that breaks, the backup has been tested, someone knows which recovery point to pick, and no customer or executive needs managing while it runs. None of those are free.

Price the gap in the currency that moves a roadmap. The nine-hour outage in the case file was a restore that had never been tested, run by people improvising under pressure with no comms and no severity call. Every hour of that was revenue and trust. The quarter buys the difference between that and a rehearsed twenty-three minutes: a loop, an update cadence, a tested failover for the tier that cannot wait for a restore, and a plan that says in writing what each workload is owed. That is not process for its own sake; it is the conversion of a random disaster into a decision the business already made.

Offer the smaller yes. You do not need the whole quarter before any of it pays off. Propose sequencing: settle severity and comms and one tested restore first — days of work that would have saved most of the nine hours — then the failover and the full BCDR plan for the revenue tier. That reframes the ask from "a quarter of overhead" to "the cheapest insurance we will ever buy, staged so it earns its keep as we go." Managers fund that; they defund "process".

Situation 02Write before you reveal
An interviewer says: "Walk me through what happens, end to end, when a service you own goes down hard — say the database is corrupted during your busiest hour." You have two minutes. What do you say?
They are testing whether you have a structure or will improvise. Walk the two lanes of Figure 1.
Reasoning

Name the structure before the steps. Open with the frame, so they know you are not improvising: "Two things happen at once — I respond to the incident and I recover the data, and I keep them separate in my head." Then walk the respond lane: detect and confirm the symptom and scope, set a severity, and start communicating on a cadence so leadership stays out of my way. Crucially, mitigate before I diagnose — if I can fail reads over to a replica or put the site in a safe degraded mode, I buy time before I care why the database corrupted.

Then walk the recover lane against the two numbers. "In parallel I ask what the tier's RPO and RTO are, because that decides the tool. If it's minutes and minutes, I fail over to the zone-redundant or replicated copy — I'm not waiting on a restore. If a restore is the right call, I pick the last clean recovery point before the corruption, not the newest, and I know which one because retention was set for exactly this." Naming that you would test-failover quarterly, so this is rehearsed, is the sentence that lands.

Close on the postmortem, because it proves you learn. "When it's back, I record the actual recovery time against the objective, and if the gap is uncomfortable I tighten the plan before the next one — blamelessly, because I want the honest timeline, not a name." Two minutes, two lanes, the two numbers, and a loop that improves itself: that is the answer that ends the reliability portion of the interview and tells them you have done this, not read about it.

Examination record · first attempt
0/4
Class Thirty-Six · Complete
Retain this much

Five things worth carrying out of this hub

  1. An outage demands two separate jobs: respond (stop the bleeding, keep people informed, learn without blame) and recover (bring the data and service back on an agreed timetable). This track teaches both because they arrive together.
  2. The two halves meet at two numbers. RPO is how much data you may lose; RTO is how long you may be down. A postmortem sets them; a recovery plan spends them.
  3. The habit that makes it all cheap: decide the hard things while nothing is on fire, and rehearse them until they are boring. A backup you have never restored is a rumour; a failover you have never run is a slide.
  4. The track is six parts, alert to all-clear: fundamentals & detection, triage & mitigation, communication & postmortems, Azure Backup, Azure Site Recovery, and BCDR strategy. Read them in order once, then keep them as reference.
  5. Seniority here is the removal of drama, not the supply of it. The reward is trust — the outage that could have been a headline becomes a footnote, and that is what closes the pay gap.
Notes
  1. The clean split into "respond" and "recover" is a teaching device; in a real severe incident the lanes interleave, and the same engineer often runs both while a second manages comms. Do not read the two lanes as two people or two phases in strict sequence — read them as two questions ("is it stopping?" and "is it coming back?") you hold at once. The six parts separate them only so you can learn each without the other's noise.
  2. "Twenty-three minutes, not nine hours" is a deliberately concrete figure carried from the case file, not a benchmark you should quote to anyone. Recovery times depend entirely on the tier, the tool, and how recently you rehearsed; treat the specific minutes with suspicion and the direction — rehearsed recovery is dramatically faster than improvised — as settled. Every part in the track hedges its own numbers where they drift.