The engineering blog of Halyard

Notes from the team that keeps nine million webhooks a day moving.

Postmortems with real numbers. Design writeups with real code. This is how we build — and occasionally break — the delivery pipeline.

  • Written by the engineers on call
  • Real incidents, real code
  • New posts when there is something to say

Writing

Everything we have published, newest first. Series posts are threaded — follow the crimson line.

  1. Rebuilding delivery · Part 2 of 25 min read

    The rebuild: a delivery pipeline that says no early

    Part one of this series told the story of November 3rd: one destination started rate-limiting us, our retries synchronized into sixty-second waves, and for 41 minutes every customer paid for it. The…

    • architecture
    • backpressure
    • rate-limiting
  2. Rebuilding delivery · Part 1 of 25 min read

    Postmortem: the retry storm that slowed webhook delivery for 41 minutes

    On November 3rd, between 09:12 and 09:53 UTC, webhook delivery at Halyard degraded badly. Most customers saw end-to-end latencies climb from under two seconds to several minutes, and for twelve of…

    • incident-review
    • queues
    • reliability

About this blog

Halyard delivers webhooks for commerce and billing platforms — about nine million a day, from checkout events to subscription renewals. This blog is the paper trail: when something breaks we publish the postmortem, and when we build something worth copying we publish the design.

We write the way we run incidents — plain language, real numbers, and code you can check. Nothing here is ghostwritten, and nothing goes through marketing.

Behind the posts

  • Mara IversenInfrastructure lead
  • Tomás ReyPlatform engineer

Questions we get

How this blog works, in four answers.

Who writes the posts?
The engineers who did the work. Every draft gets a technical review from someone else on the incident or project, and one edit for clarity. Nothing is ghostwritten and nothing goes through marketing.
Can I republish or translate a post?
Yes. Republish or translate any post with attribution and a link to the original. The code snippets are MIT-licensed — use them without asking.
Do you publish a postmortem for every incident?
We publish when customers felt the impact or when the failure taught us something transferable. Internal-only incidents get the same written review; it just stays internal.
Do you take guest posts or sponsorships?
No. Everything here was written by someone at Halyard, about systems we run in production. That is the whole editorial policy.