Karka

Free course

Observability and Monitoring

The operating practice of monitoring digital services: service level objectives, symptom-based alerting, humane on-call, incident response and blameless review.

What this course does not do

This course covers the monitoring and incident practice of digital services only. It confers no competence and no authorisation for physical workplace emergency response, first aid, evacuation duties or any statutory safety appointment; those are proven by external credentials and practical sign-off, never by an online course.

Professional electiveCorePlatform Engineering
7 modules 35 lessons 3 enrolled ~12h of material

First three lessons free. Full access with Founding Annual Access, ₹3,999 for your first year.

Observability and Monitoring

By the end

What you'll build

  • Write a service level indicator and objective derived from a real user journey
  • Design symptom-based alerts with defined severity, routing and ownership, and retire alerts that no longer earn their place
  • Combine black-box and white-box monitoring, including synthetic journeys and saturation signals
  • Run an on-call rota with handover, escalation and a measurable view of load
  • Coordinate an incident with clear roles, communications and a reconstructable timeline
  • Run a blameless review that produces a small number of actions that are completed

Curriculum

What's inside

7 modules · 35 lessons

  1. 01

    Service level thinking

    5 lessons
    • Writing a service level objective from a real user journey
    • Indicators, objectives and agreements distinguished
    • Choosing measurement windows
    • Error budgets and what they actually permit
    • Indicator, objective, agreementmatch pairs
  2. 02

    Alerting

    5 lessons
    • Alerting on symptoms: designing a page always worth waking for
    • Severity, routing and ownership
    • Burn-rate alerting in outline
    • Reducing noise and retiring dead alerts
    • What deserves a pagefill blank
  3. 03

    Health monitoring and synthetic checks

    5 lessons
    • Black-box and white-box monitoring used together
    • Designing synthetic journeys and probes
    • Monitoring dependencies you do not control
    • Capacity and saturation signals
    • Standing up a synthetic journeysequence order
  4. 04

    On-call practice

    5 lessons
    • Running an on-call rota people can sustain
    • Handover, escalation and shadowing
    • Runbooks a tired engineer can follow
    • Measuring on-call load and acting on it
    • The 02:40 pagescenario
  5. 05

    Incident response

    5 lessons
    • Coordinating an incident: roles, communications and timeline
    • Detection, triage and mitigating before diagnosing
    • Keeping stakeholders informed during an incident
    • Severity classification and when to escalate
    • Who does what while it is on firematch pairs
  6. 06

    Learning from incidents

    5 lessons
    • Running a blameless review that produces real change
    • Timelines and contributing factors
    • Writing actions that are small enough to finish
    • Looking for patterns across many incidents
    • Wording a review that changes somethingfill blank
  7. 07

    Practice and check

    5 lessons
    • Signal, role, decisionmatch pairs
    • The first hour of an incidentsequence order
    • Say it preciselyfill blank
    • Checkout is failing at 02:40scenario
    • Course quizquiz

The shape of it

How this course works

Short lessons

35 lessons across 7 modules, each small enough to finish in one sitting.

Practice as you go

Every lesson ends with a small space for what you noticed — the doing is the learning.

Progress you can see

Your progress is saved lesson by lesson, ready whenever you come back.

Ready when you are.

Make an account and this course opens up — your progress is saved from the very first lesson.