Part one of this series told the story of November 3rd: one destination started rate-limiting us, our retries synchronized into sixty-second waves, and for 41 minutes every customer paid for it. The…
On November 3rd, between 09:12 and 09:53 UTC, webhook delivery at Halyard degraded badly. Most customers saw end-to-end latencies climb from under two seconds to several minutes, and for twelve of…
incident-review
queues
reliability
About this blog
Halyard delivers webhooks for commerce and billing platforms — about nine million a day, from checkout events to subscription renewals. This blog is the paper trail: when something breaks we publish the postmortem, and when we build something worth copying we publish the design.
We write the way we run incidents — plain language, real numbers, and code you can check. Nothing here is ghostwritten, and nothing goes through marketing.
Behind the posts
Mara IversenInfrastructure lead
Tomás ReyPlatform engineer
Questions we get
How this blog works, in four answers.
Who writes the posts?
The engineers who did the work. Every draft gets a technical review from someone else on the incident or project, and one edit for clarity. Nothing is ghostwritten and nothing goes through marketing.
Can I republish or translate a post?
Yes. Republish or translate any post with attribution and a link to the original. The code snippets are MIT-licensed — use them without asking.
Do you publish a postmortem for every incident?
We publish when customers felt the impact or when the failure taught us something transferable. Internal-only incidents get the same written review; it just stays internal.
Do you take guest posts or sponsorships?
No. Everything here was written by someone at Halyard, about systems we run in production. That is the whole editorial policy.