Who this is for
You want the reliability-and-scale side of infrastructure work — less initial building, more keeping complex systems healthy under real load.
How this path works
Each cluster below links to its real lesson page. Go in order the first time through — later clusters assume the ones before them. Once you've done a full pass, use it as reference.
The path — in order
Stack Floor
01
What's actually running underneath the systems you're keeping alive.
Domain 0 · Stack Floor
02
Failover, the on-premises version, before cloud-scale HA.
Domain 0 · Stack Floor
Compute
03
The compute layer SRE work centres on.
Domain 1 · Compute
04
Autoscaling — the difference between an outage and a blip.
Domain 1 · Compute
Networking
05
Networking issues are frequently the real root cause.
Domain 3 · Networking
06
Distributing load so no single point buckles.
Domain 3 · Networking
Operations
07
You can't fix what you can't see.
Domain 5 · Operations
08
Alerts that mean something, not noise.
Domain 5 · Operations
IaC & DevOps
09
Fixing it once, in code, not by hand every time.
Domain 6 · IaC & DevOps
10
Safe, repeatable deployment as a reliability practice.
Domain 6 · IaC & DevOps
Resilience
11
Planning for the failure you haven't seen yet.
Domain 7 · Resilience
12
Availability zones and region pairs at cloud scale.
Domain 7 · Resilience
Certification this path maps to
Being honest: no single Microsoft certification is built specifically for SRE. AZ-104 (Azure Administrator) is the practical baseline most SREs hold, paired with real hands-on Azure Monitor experience.
Where this leads
Platform Engineer and DevOps Engineer share most of this skill set with different emphasis — platform-building versus pipeline-building versus keeping-it-running.