By the end
What you'll build
- Choose between counters, gauges and histograms for a given measurement and explain the consequences
- Predict and control label cardinality before it becomes a cost or reliability problem
- Write structured logs with correlation identifiers and a defensible retention and sampling policy
- Instrument a service boundary so trace context propagates and latency can be attributed
- Correlate a trace with the metrics and logs that explain it
- Reduce telemetry cost while keeping the questions you need to answer answerable
Curriculum
What's inside
7 modules · 35 lessons
- 01
Observability compared with monitoring
5 lessons- Known unknowns: answering questions you did not plan for
- Signals, telemetry and the cardinality problem
- Instrumenting for questions rather than for dashboards
- Telemetry cost as a design constraint from day one
- Signals, questions and what each one costsmatch pairs
- 02
Metrics
5 lessons- Choosing counters, gauges and histograms correctly
- Labels, cardinality and the explosion problem
- Aggregation, quantiles and their common pitfalls
- Naming and unit conventions that survive reuse
- Pick the right metric, then read it rightfill blank
- 03
Logs
5 lessons- Structured logging that is worth the storage it costs
- Levels, sampling and retention decisions
- Correlation identifiers across service boundaries
- Keeping sensitive and personal data out of logs
- Making one request findablesequence order
- 04
Traces
5 lessons- Reading a distributed trace to find where the latency lives
- Spans, context propagation and sampling strategies
- Instrumenting a service boundary
- Linking traces to logs and metrics
- Where the latency livesscenario
- 05
Vendor-neutral instrumentation
5 lessons- Instrumenting once against an open, vendor-neutral telemetry specification
- Automatic and manual instrumentation compared
- Collectors, pipelines and in-flight processing
- Migrating an existing estate incrementally
- Pipeline pieces and where they runmatch pairs
- 06
Making telemetry usable
5 lessons- Designing a service dashboard someone can read under pressure
- Drilldown paths from summary to example
- Telemetry ownership and lifecycle
- Cutting telemetry cost without losing visibility
- Dashboards someone else can readfill blank
- 07
Practice and check
5 lessons- Signals and what they meanmatch pairs
- Working a latency alert at 03:00sequence order
- Getting the terms exactly rightfill blank
- High latency, unfamiliar servicescenario
- Course quizquiz
The shape of it
How this course works
Short lessons
35 lessons across 7 modules, each small enough to finish in one sitting.
Practice as you go
Every lesson ends with a small space for what you noticed — the doing is the learning.
Progress you can see
Your progress is saved lesson by lesson, ready whenever you come back.
Ready when you are.
Make an account and this course opens up — your progress is saved from the very first lesson.
