Signal Path: From metric to post-mortem
A practical path through SRE and observability: you instrument a real service, measure reliability with SLOs and error budgets, correlate metrics, traces and logs, then close the loop with alerts, a runbook, an incident and a post-mortem. All on Docker Compose, so it runs on your own VPS (Hetzner, Oracle, Hostinger) with Coolify or Portainer.
- modules
- 10
- phases
- 5
- final project
- 1
Prometheus Grafana Loki Tempo OpenTelemetry Alertmanager Docker Compose
Module 1 · Phase 1 · Foundations
SRE and observability fundamentals
Before installing anything: what problem we are solving, and with which signals.
Monitoring is not observability
Monitoring answers questions you already knew you would ask: is the container up, did CPU cross 80%, does the endpoint respond? Observability lets you answer questions you never anticipated: why did only some /checkout requests take 8 seconds between 14:02 and 14:07?
The raw material is the data your system emits — telemetry — and the three classic signals are metrics, logs and traces. Each answers a different question, and the real value shows up when you connect them.
The metric tells you THAT there is a problem. The trace tells you WHERE it is. The log explains WHY. The SLO and burn rate tell you HOW URGENT it is.
The four golden signals
Latency (how long requests take), traffic (how much work arrives), errors (how many fail) and saturation (how close to the limit you are). Those four numbers answer in ten seconds whether a service is healthy.
Two derived methods organise the work: RED (Rate, Errors, Duration) looks at the service from the user's side; USE (Utilization, Saturation, Errors) looks at the resource — CPU, memory, disk, database connections. A good dashboard uses both, in that order: impact first, cause second.
Thirty panels and no answers. A pretty dashboard is not observability: if nobody can go from "something is wrong" to "this request spent its time in this external call" in under five minutes, you do not have it yet.
Percentiles, not averages
If 99 requests take 100 ms and one takes 10 s, the average lies and that user is gone. So you measure p50 (typical experience), p95 (the one used in agreements) and p99 (the long tail, where timeouts and retries live).
Map the golden signals of one of your services
Pick a real service you already run. For each signal, write down which number represents it today and where it would come from: latency (p95 of which endpoint), traffic (req/s), errors (what counts as an error: only 5xx, or business 4xx too?) and saturation (CPU, memory, connection pool). If you do not know where a number comes from, mark it as a gap.
Draw your current telemetry flow
Diagram the path each signal takes today from your app to wherever someone looks at it: app → which collector? → which backend? → which UI? Mark the missing legs in red. By the end of module 8 this same diagram must be complete.
Module deliverable
Goes into the final project: That definition of a failed request is literally the SLI you will measure in module 4 and the one that will fire your alerts in module 5. Choose it carefully.
Module 2 · Phase 2 · Metrics
Metrics with Prometheus
Instrument your app and ask your first questions in PromQL.
Three metric types cover almost everything
A counter only goes up and counts events: requests, errors, messages processed. A gauge goes up and down and measures instantaneous state: open connections, queue depth, memory. A histogram spreads observations across buckets and is what lets you compute latency percentiles server-side.
The minimum instrumentation for any HTTP service is two metrics: a request counter with route, method and status labels, and a duration histogram with the same labels.
Every label combination is a time series. Putting user_id, order_id or a URL containing IDs into a label creates millions of series and takes the instance down. Identifiers belong in logs and traces, never in metric labels. Use the route (/orders/:id), not the URL.
Prometheus on Docker Compose
services:
prometheus:
image: prom/prometheus:latest
command:
- --config.file=/etc/prometheus/prometheus.yml
- --storage.tsdb.retention.time=15d
volumes:
- ./prometheus.yml:/etc/prometheus/prometheus.yml:ro
- prom-data:/prometheus
ports: ["9090:9090"]
The scrape config points at your service by internal network name: in Compose (and in a Coolify or Portainer stack) containers reach each other by service name, so targets: ['api:3000'] is enough. Never expose /metrics to the internet unprotected.
The three queries you will always use
# traffic (req/s)
sum(rate(http_requests_total[5m]))
# error rate (share of 5xx)
sum(rate(http_requests_total{status=~"5.."}[5m]))
/ sum(rate(http_requests_total[5m]))
# p95 latency
histogram_quantile(0.95,
sum by (le) (rate(http_request_duration_seconds_bucket[5m])))
rate() over a counter gives the per-second slope, not the total. Always rate() before sum(), never the other way round: summing raw counters across replicas that restart produces phantom numbers.
Instrument a real service
Add prom-client (Node) or prometheus_client (Python) to one of your services. Expose /metrics with a http_requests_total{route,method,status} counter and a http_request_duration_seconds histogram. Check that the route is normalised (no IDs in the label) and that the histogram buckets cover your real latency range.
Bring up Prometheus and write the three queries
Build the docker-compose.yml with Prometheus scraping your service every 15s. Drive load at it (hey, k6 or a curl loop) and answer in the Prometheus UI: what is your RPS, your error rate, your p95? Save the three queries with a comment on what each answers.
Module deliverable
Goes into the final project: That histogram is what will feed the exemplars in module 8. Name it well and pick sensible buckets now and the correlation later comes for free.
Module 3 · Phase 2 · Metrics
Grafana: dashboards that answer questions
A dashboard is not a pile of charts: it is a reading order.
Order matters more than the panels
An operational dashboard reads top to bottom like a triage. Layer 1: impact — availability, error rate, latency, error budget. Layer 2: RED — traffic, errors and duration per route. Layer 3: saturation — CPU, memory, replicas, restarts, connection pool. Layers 4 and 5 (from module 6 on): logs and traces for the window you are looking at.
If someone opens the board at 3 AM, the first screen — without scrolling — must answer: are users affected, and since when? Everything else goes below.
Grafana goes in the compose too
grafana:
image: grafana/grafana:latest
environment:
- GF_SECURITY_ADMIN_PASSWORD=${GRAFANA_PASSWORD}
- GF_USERS_ALLOW_SIGN_UP=false
volumes:
- grafana-data:/var/lib/grafana
- ./grafana/provisioning:/etc/grafana/provisioning:ro
ports: ["3001:3000"]
The provisioning volume is what separates a toy from something operable: datasources and dashboards defined as files, versioned in git, reproducible on another VPS. A dashboard hand-built in the UI and never exported is lost the day the volume breaks.
Variables: one board, many services
Instead of cloning the dashboard per service, define a $service variable with label_values(http_requests_total, service) and use it in every query. The same board serves your API, your worker and the WhatsApp bot.
Panels with automatic axes and no unit. If the axis does not say seconds or percent, and the SLO threshold is not drawn as a line, nobody knows whether 0.42 is good or bad.
Dashboard v1 in four panels
Build a board with: availability (1 − error ratio) as a large stat, RPS per route, error rate with a painted threshold, and p95/p99 on the same chart. Set correct units on every axis and a threshold line on latency. Export the JSON to grafana/dashboards/red-v1.json.
Make it reproducible
Move the Prometheus datasource and the dashboard into provisioning files. Delete the Grafana volume, bring the stack back up and confirm everything returns on its own. If it does not, you still have configuration living only in the UI.
Module deliverable
Goes into the final project: This board is the base of the final dashboard. In module 4 you add the SLO and error-budget row, and in module 8 the correlated panels.
Module 4 · Phase 3 · Measured reliability
SLIs, SLOs, error budget and burn rate
Turning "it's fine" into a number you can argue about with product.
Three acronyms that are not synonyms
The SLI is what you measure (availability = 99.92%). The SLO is the internal target you set (99.9% over 30 days). The SLA is a contractual commitment with money or credits attached. Almost no service needs an SLA; almost every one should have an SLO.
The error budget
If the SLO is 99.9%, you are allowed to fail 0.1%. Over 30 days that is 43 minutes and 12 seconds of unavailability. That budget is a decision tool: while budget remains you can ship fast and take risks; once you have burned it, the priority becomes stabilising.
SLO 99.9% → 43m 12s / 30 days
SLO 99.95% → 21m 36s / 30 days
SLO 99.99% → 4m 19s / 30 days
Each extra nine multiplies cost: redundancy, on-call, slower deploys. You pick the SLO by looking at what users tolerate and what the business will pay for, not at what sounds impressive in a meeting.
Burn rate: how fast you are spending
Burn rate is how much faster than allowed you are consuming the budget: error_ratio / (1 − SLO). A burn rate of 1× consumes the budget in exactly 30 days. 14.4× consumes it in a little over two days — that is worth waking someone. A sustained 3× is not urgent tonight, but it destroys the month.
Recording rules: compute once
groups:
- name: slo
interval: 30s
rules:
- record: job:http_error_ratio:rate5m
expr: |
sum(rate(http_requests_total{status=~"5.."}[5m]))
/ sum(rate(http_requests_total[5m]))
- record: job:http_availability:ratio30d
expr: 1 - (
sum(rate(http_requests_total{status=~"5.."}[30d]))
/ sum(rate(http_requests_total[30d])))
With recorded rules, dashboards and alerts query a short name instead of recomputing a heavy query every 10 seconds. It also unifies the definition: one single truth about what "availability" means across your whole stack.
Define and justify your SLO
Write two SLOs for your service in docs/slo.md: one for availability and one for latency (for example, 99% of requests under 500 ms). For each, note the error budget in minutes per month and one sentence on why that number and not a higher one. If you cannot justify it, it is not an SLO yet.
Recording rules + budget panel
Create rules/slo.yml with error ratio at 5m, 30m, 1h, 6h and availability at 30d. Load them into Prometheus, check in /rules that they evaluate without error, and add a remaining error budget panel (1 − (error_ratio_30d / (1 − SLO))) as a percentage to the dashboard.
Module deliverable
Goes into the final project: The rules from this module are exactly what will fire the alerts in module 5. Without an SLO there are no good alerts, only invented thresholds.
Module 5 · Phase 3 · Measured reliability
Alerts that don't wake people for nothing
Alert on user symptoms, not machine symptoms.
The 3 AM test
Before creating an alert, answer this: if it fires at 3 AM, does someone have to get up and act now? If the answer is no, it is not a page: it is a ticket, or plain noise. Alert fatigue is the fastest way to get a team to stop looking at their phone.
CPU > 70%. High CPU alone does not mean impact: it may be the normal midday peak. Alert on availability, latency and burn rate — that is, on what the user suffers.
Multi-window burn-rate alerts
Google's standard recipe uses two windows per severity, to detect fast without firing on a thirty-second spike: the long window confirms the problem is real, the short one confirms it is still happening.
groups:
- name: slo-alerts
rules:
- alert: ErrorBudgetFastBurn
expr: |
job:http_error_ratio:rate1h > (14.4 * 0.001)
and job:http_error_ratio:rate5m > (14.4 * 0.001)
for: 2m
labels: { severity: page }
annotations:
summary: "Fast error-budget burn"
runbook: "https://git.your-domain/runbooks/availability.md"
- alert: ErrorBudgetSlowBurn
expr: |
job:http_error_ratio:rate6h > (6 * 0.001)
and job:http_error_ratio:rate30m > (6 * 0.001)
for: 15m
labels: { severity: ticket }
Alertmanager decides who hears about it
Prometheus detects the condition; Alertmanager does the rest: it groups related alerts so you do not send twenty messages, deduplicates, silences during maintenance and inhibits child alerts once the parent has fired. Routing is by label: severity: page goes to the channel that wakes people, severity: ticket to the one read during business hours.
Instead of configuring every integration inside Alertmanager, send a single webhook_configs to an n8n flow. There you format the message, enrich it with dashboard and runbook links, and route to WhatsApp, Slack or email by severity and time of day. One integration to maintain.
Every alert needs a runbook
The runbook annotation is not decoration: it is the link the on-call person opens half asleep. If the alert has no runbook, it is not finished.
Write two burn-rate alerts
Create rules/alerts.yml with a fast-burn alert (14.4× over 1h and 5m, severity page) and a slow-burn one (6× over 6h and 30m, severity ticket), both with a summary annotation and a runbook link. Validate with promtool check rules before reloading.
Routing and a real test
Configure alertmanager.yml with two routes by severity and a webhook to an n8n flow that formats and delivers the message. Then break something on purpose (kill a dependency, return 500 on a fraction of requests) and time it: how long from the first error to the message on your phone. That number is your MTTD.
Module deliverable
Goes into the final project: The MTTD you measure here is the first metric of your incident process in module 9. Keep it: at the end of the course you will compare it against the GameDay one.
Module 6 · Phase 4 · Logs and traces
Structured logs and Loki
Stop reading loose text, start querying events.
A log is an event, not a sentence
Free-text log: Error processing user order. Structured log:
{"ts":"2026-09-08T14:03:11Z","level":"error","service":"api",
"route":"/orders/:id","status":500,"duration_ms":831,
"trace_id":"b2fe643bc4d341a1f7076f265910e649","order_id":"o_8812",
"msg":"payment gateway timeout"}
The second can be filtered, grouped and counted. The minimum fields that must be there: timestamp, level, service, route, status, duration_ms and trace_id. That last field is what will connect this log to its trace in module 8.
Loki is not Elasticsearch
Loki does not index log content: it indexes only the stream labels and stores the rest compressed. That is why it is cheap to run on a VPS, and why labels must be few and low-cardinality: service, env, level. Nothing more.
Using trace_id, user_id or order_id as a stream label. Every distinct value creates a new stream and the index becomes unmanageable. Those fields belong inside the log JSON and are filtered with | json | trace_id="...", which is just as fast for what you need.
Collection on Docker
loki:
image: grafana/loki:latest
command: -config.file=/etc/loki/local-config.yaml
volumes: [ "loki-data:/loki" ]
alloy:
image: grafana/alloy:latest
volumes:
- ./alloy/config.alloy:/etc/alloy/config.alloy:ro
- /var/run/docker.sock:/var/run/docker.sock:ro
command: run /etc/alloy/config.alloy
Alloy discovers containers through the Docker socket, applies labels from container names and labels, and pushes to Loki. A simpler alternative if your VPS is already tidy: the Docker logging driver pointing straight at Loki.
LogQL in three moves
# every error from the service
{service="api", level="error"}
# only 5xx, parsing the JSON
{service="api"} | json | status >= 500
# error rate per route, last 5 minutes
sum by (route) (
rate({service="api"} | json | status >= 500 [5m]))
On a VPS, disk is finite. Set retention up front (say 14 days of logs and 30 of metrics) in the Loki config, not when you run out of space on a Sunday.
Move your app to structured logging
Replace console.log / print with a JSON logger (pino, winston, structlog). Make sure every request logs a line with route, status, duration_ms and trace_id, and that errors carry the stack in its own field rather than concatenated into the message.
Loki + three useful queries
Add Loki and Alloy to the compose, add the datasource in Grafana and save three LogQL queries: all 5xx in the last hour, errors for one specific route, and error rate per route. Add a logs panel to the dashboard, below the metrics panels.
Module deliverable
Goes into the final project: Right now the trace_id in your logs points nowhere: module 7 creates the trace, and module 8 turns that field into a clickable link.
Module 7 · Phase 4 · Logs and traces
Distributed tracing with OpenTelemetry and Tempo
Knowing exactly where the time went, span by span.
Trace, span and context
A trace is the full path of a request through your system. Each measurable step is a span: the inbound HTTP request, the Postgres query, the GET to the payments API, the job that got queued. Spans form a tree, and that tree shows immediately where the time went:
POST /orders 2.41s
├── auth.verify 38ms
├── db.query orders_insert 142ms
├── http POST payments.api 2.14s ←
└── queue.publish order.created 21ms
The conclusion is not "the app is slow": it is "the payment gateway is eating 89% of the time". Those are two completely different conversations.
OpenTelemetry: instrument once
OTel is the open standard that keeps you from being locked to a vendor: you instrument with its SDK and then decide where the data goes — Tempo today, Datadog or Grafana Cloud tomorrow, without touching app code.
Auto-instrumentation covers the obvious parts (HTTP server, HTTP client, database driver, Redis) with very little code. Manual spans go where your expensive logic lives: the heavy computation, the LLM call, the PDF render.
Nobody instruments what they do not suspect. Put a manual span around every third-party call and every operation that can exceed 100 ms. The spans you never created are exactly the ones you will miss during the incident.
The Collector as a buffer
otel-collector:
image: otel/opentelemetry-collector-contrib:latest
command: ["--config=/etc/otel/config.yaml"]
volumes: [ "./otel/config.yaml:/etc/otel/config.yaml:ro" ]
ports: ["4318:4318"] # OTLP/HTTP
tempo:
image: grafana/tempo:latest
command: ["-config.file=/etc/tempo.yaml"]
volumes:
- ./tempo/tempo.yaml:/etc/tempo.yaml:ro
- tempo-data:/var/tempo
The app speaks OTLP to the Collector, and the Collector decides what to do: drop health-check spans, sample, add environment attributes, and export to Tempo. Switching backend later means changing an exporter, not redeploying the application.
TraceQL to find the needle
{ resource.service.name = "api" && duration > 1s }
{ span.http.status_code >= 500 }
{ name = "http POST payments.api" && duration > 2s }
Keeping 100% of traces on a small VPS fills the disk fast. Sample (10% is a reasonable starting point) but always keep 100% of error traces: those are the ones you will look at.
Instrument with OTel and ship to Tempo
Add the OpenTelemetry SDK with HTTP and database auto-instrumentation, exporting OTLP/HTTP to the Collector. Add Collector and Tempo to the compose and confirm in Grafana Explore that you see complete traces with their child spans.
Find your slowest operation
Add a manual span around the most expensive operation in your service, with useful attributes (payload size, external provider, cache hit or miss). Then use TraceQL to find traces over 1s and note which span eats the time. That is your first data-driven optimisation hypothesis.
Module deliverable
Goes into the final project: You now have all three signals, but still separate: three tabs and a lot of copy-paste. Module 8 joins them.
Module 8 · Phase 4 · Logs and traces
Correlation: metric → trace → log
The three-click jump that cuts MTTR from hours to minutes.
One common key: trace_id
The three signals you have built so far live in different stores. What connects them is an identifier that travels with the request: the trace_id. If your histogram attaches it as an exemplar, your log writes it as a field and your trace has it by definition, you can walk the whole path without writing a single query by hand.
Adding more panels does not shorten an incident. What shortens it is being able to go from the spike on the chart to the concrete request, and from there to the error log line, in three clicks. This entire module exists for that.
Exemplars: from metric to trace
An exemplar is a concrete point attached to a histogram bucket, carrying the trace_id of a real request that landed there. On the latency chart it shows up as a small dot on the curve: you click it and the trace that produced that value opens.
# Prometheus needs the feature enabled
command:
- --enable-feature=exemplar-storage
# and the scrape in OpenMetrics format
scrape_configs:
- job_name: api
static_configs: [{ targets: ["api:3000"] }]
On the app side, the Prometheus library must record the observation with the active trace's exemplar. In Node with prom-client it goes as the third argument to observe(); in Python, via exemplar=.
Derived fields: from log to trace
In the Loki datasource config you define a derived field that detects the trace_id in the JSON and turns it into a button that opens Tempo:
derivedFields:
- name: TraceID
matcherRegex: '"trace_id":"(\w+)"'
url: '${__value.raw}'
datasourceUid: tempo
And in the Tempo datasource you enable the reverse path (Trace to logs): from a span, jump to the logs for that same trace_id within the span's time window. The round trip is closed.
If containers have drifting clocks, the trace → logs jump returns nothing even when the config is correct: Grafana searches a time window that does not line up. Sync NTP on the VPS before you lose a day debugging config.
The full path
Burn-rate alert
↓ (runbook link)
Dashboard: p95 at 12s since 14:02
↓ (click the exemplar)
Trace: POST /orders — 11.8s in payments.api
↓ (trace to logs)
Log: "payment gateway timeout" order_id=o_8812
↓
Root cause found — 4 minutes since the alertEnable exemplars and make the jump
Turn on exemplar-storage in Prometheus, attach the trace_id when observing latency, and switch on Show exemplars in the p95 panel. Generate slow load and confirm you can click a point on the chart and land in the right trace.
Close the loop between logs and traces
Configure the derived field in the Loki datasource and Trace to logs in the Tempo one. Then walk the full path with a stopwatch: from the latency panel to the log line showing the cause. Note how many clicks and how many seconds it took — that number is what you will defend in the final project.
Module deliverable
Goes into the final project: This path is literally the body of your module 9 runbook: instead of writing "check the logs", you will write the exact three clicks.
Module 9 · Phase 5 · Operations
Incident response, severity and runbooks
The process that turns panic into ordered steps.
Four stages and three clocks
An incident is detected, acknowledged, mitigated and resolved — and each leg has its metric. MTTD (detect): first error to alert. MTTA (acknowledge): alert to someone picking it up. MTTR (restore): start to normal service. If you do not measure the three separately, you do not know which to improve: a 20-minute MTTD is not fixed by more on-call people, it is fixed by better alerts.
During the incident the goal is restoring service, not finding root cause. Roll back, kill the feature flag, scale replicas, cut traffic to the failing provider. Investigation belongs in the post-mortem, with the service already healthy.
Severity: define it before you need it
SEV1 Service down or data loss.
All users. People get woken. Comms immediately.
SEV2 Serious degradation: SLO at risk, important
functionality broken. Handled now, in or out of hours.
SEV3 Limited impact or a workaround exists.
Ticket, resolved during business hours.
The criterion is always user impact, never how much code needs touching. A CSS typo in checkout that blocks purchases is SEV1; a downed reporting worker may be SEV3.
Anatomy of a runbook that works
A runbook is for someone opening their phone at 3 AM, not for whoever wrote the system. Five sections:
- Symptom — which alert fired and what the user is seeing.
- Verification — the exact, copyable queries, with dashboard links.
- Mitigation — concrete commands, in order, least to most destructive.
- Escalation — who to call, and from which minute.
- Post-incident — what to save before it is lost: trace IDs, screenshots, the time window.
A runbook that says "check the logs and verify the service is healthy" is not a runbook, it is a wish. If a step cannot be copy-pasted, it is not finished.
Write a real runbook
Pick the most likely failure mode of your service (external dependency down, database saturated, disk full) and write the complete runbook with all five sections. The verification queries must be the ones you built in modules 2, 6 and 8, copyable as-is.
Timed drill
Break the service on purpose and run the whole incident as if it were real: wait for the alert, acknowledge it, follow your own runbook, mitigate. Record the three times and honestly note which runbook step did not work. Fix it the same day.
Module deliverable
Goes into the final project: The drill in this module is the rehearsal. In module 10 you do it for real, with failures you did not pick yourself, and write the post-mortem.
Module 10 · Phase 5 · Operations
Post-mortem and GameDay
Making the incident leave learning behind, not just exhaustion.
Blameless post-mortems, for real
"Blameless" is not politeness: it is the only thing that makes people tell you what actually happened. If whoever pressed the button knows they will be named, next time the timeline will have holes exactly where the useful information was. The right question is never who made the mistake, but what made that action look reasonable at the time.
The six sections
1. Summary — 3 lines: what happened, who was affected, how long.
2. Impact — users, failed requests, error budget consumed,
money if applicable.
3. Timeline — with exact times: first error, alert, ack,
mitigation, resolution. MTTD/MTTA/MTTR come from here.
4. Cause — the technical one and the contributing one (why the
system allowed it to happen).
5. What worked — yes, this section stays. The alert that fired correctly
and the runbook that helped are results too.
6. Action items — each with an owner, a date and a ticket. Without an
owner it is not an action item, it is a comment.
The percentage of action items closed. A beautiful document with twelve tasks nobody touched in three months means the incident will happen again in exactly the same way.
GameDay: breaking things on purpose, in office hours
A GameDay is a planned incident: you pick a day, tell the team, inject real failures and check whether your observability catches them. It is the only honest way to know whether the stack works — because on the day of the real incident you do not want to discover the alert was never configured properly.
Hypothesis: "If the payment provider starts returning 500s,
the fast-burn alert fires within 5 minutes and
the runbook leads us to the cause within 10."
Injection: a proxy returning 500 on 30% of calls
Observation: did it fire? how fast? was the runbook enough?
Learning: what was missing — a span, a panel, a runbook step
Always announce it (it is not a surprise test on people), have the rollback plan written before you start, set a hard stop time, and silence the alert channels so you do not confuse anyone not taking part.
Three failures to start with
A slow or broken external dependency, resource saturation (CPU at the limit or connection pool exhausted), and losing a container (kill it and see whether the system recovers on its own). Those three alone will find at least one hole in your instrumentation.
Run a 60-minute GameDay
Write the plan first: three hypotheses, three injections, a success criterion for each and a rollback plan. Execute it with a stopwatch. Note which failure your stack did NOT detect — that is the most valuable finding of the day.
Write the GameDay post-mortem
Use the six sections on what actually happened, with the measured times. Action items must be concrete: "add a span to the payments call", "lower the alert's for: from 5m to 2m", with owner and date. Close at least one that same week.
Module deliverable
Goes into the final project: With this you close the full loop: instrument, measure, alert, correlate, respond and learn. What remains is packaging it as a stack someone else can bring up.
Final project
Final project: the full stack on a real service
The final deliverable is not a toy lab: it is the observability of a service you actually run, packaged so that someone else can bring it up with docker compose up -d and understand it in fifteen minutes.
What has to be running
- An instrumented service: RED metrics, JSON logs with
trace_id, OTel traces - Prometheus with SLO recording rules and exemplars enabled
- Grafana provisioned from files: datasources, five-layer dashboard, service variable
- Loki + Alloy with defined retention, Tempo with sampling
- Alertmanager with two severities routing into an n8n flow
- The whole stack behind authentication, no ports open to the world
What has to be written
docs/slo.md— two justified SLOs with their error budgetrunbooks/— at least two failure modes, with copyable queriesdocs/postmortem-template.mdand the GameDay post-mortemREADME.md— how to bring everything up from scratch on a clean VPS
Success criteria
| Criterion |
|---|
Your progress is saved in this browser.