Quick thanks to Jim and XO Ruby for giving me this speaking opportunity. Before we begin, let's do some quick tests.
Everyone's seen this sign before, right? Pretty self explanatory. No parking
What about now? Which side can you park on? Left or right? These were taking right on the street I live on. One on the north side, the other on the south side. There's even a rule here that's not on the sign. You can't park here longer than 3 hours.
Here's another example of signs. These signs are doing A LOT of work. Streetcars, cars, bikes, people, all being held together by a few signs and some lights. I'm no expert, but I feel like this isn't real infrastructure.
Here's one more. If you were walking down this hallway, these are two sets of double doors. Which side do you walk through? The right, right?
But if you chose the right side, you're wrong, you're better off walking through the left. It's not very intuitive is it?
And that brings me to my title slide. I mean, who wants to think, right? Thinking is slow. Thinking leaves room for errors.
My journey into observability began with my tenure at Honeybadger. Building observability products, tools and infrastructure. Then moving on to Phoenix and becoming and end user of those products gave me a fresh perspective on how to best leverage o11y.
Satisficing from Chapter 2 about how wrong our assumptions are about user behavior
Slips vs Mistakes from Chapter 5 about errors slips: plan was right, execution was wrong mistakes: plan was wrong, execution was right
These questions all require knowing how your system is performing at any given moment.
For the long time these were the basic pieces of data when people talked about observability.
Metrics: cheap aggregated numbers that tell you something changed, but not why.
Logs: a free-form story of what happened, easy to write and hard to query.
Traces: a nested group of data, which provides really good insights into debugging slow code.
Known unknowns: recording response time so you can know when requests are slow Unknown unknowns: recording more data so you can know why requests are slow High cardinality: fields with many unique values like user_id or order_id, so you can find the exact request. Multiple dimensions: many fields on one event like plan, region, and app version, so you can slice by any combination.
Same method, same single log line, but now it's a wide event. High cardinality: user_id, order_id, request_id. Every value is unique, so you can find the exact request. Dimensions: plan, coupon, country, app_version, region, flags. Slice by any combination after the fact.
This is similar to: https://github.com/t6d/active_operation https://github.com/collectiveidea/interactor
This only rescues CardError, but the same call can also raise: * Stripe::RateLimitError: too many requests to the API too quickly * Stripe::InvalidRequestError: bad parameters, like a missing customer * Stripe::AuthenticationError: invalid or revoked API key * Stripe::APIConnectionError: network failure talking to Stripe * Stripe::APIError: something went wrong on Stripe's side Every developer has to remember to rescue and report each one.
queue_latency_ms: duration only counts run time. "Took 120ms to send, waited 38 minutes to start." attempt: failed once and retried fine vs failing on attempt 25 for three days. queue: slice latency by queue, e.g. mailers starved because default is flooded. host / ecs_task / availability_zone: fetched once at boot, cheap to add to every event. "Errors only on one task" or "slow jobs only in ca-central-1b" become one GROUP BY.
Reduce the chances of your devs making slips and mistakes.