Sam Okonkwo
Site Reliability Engineer
Sam has carried the pager for systems on both AWS and GCP and writes the runbook they wish had existed at 3 a.m.
Metrics, Traces, and Logs Without the Cargo Cult
By Sam Okonkwo
Most observability stacks are expensive and unhelpful in the same incident. They collect enormous volume and still cannot answer 'which change caused this'.
Sam Okonkwo works backwards from the questions you ask under pressure and derives the instrumentation that answers them: cardinality budgets that hold, trace sampling that keeps the interesting traces, and alerting tied to user-visible symptoms rather than machine states.
Vendor-neutral, with concrete guidance for CloudWatch and Google Cloud Observability.
6 chapters · 322 pages total
Site Reliability Engineer
Sam has carried the pager for systems on both AWS and GCP and writes the runbook they wish had existed at 3 a.m.