{
  "rsas_version": "0.1",
  "page": 0,
  "frozen": true,
  "items": [
    {
      "kind": "entry",
      "id": "itl:10b1c729c8d94057c35336025f24ce68",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2026/07/29/lorenz-and-little",
      "url": "http://brooker.co.za/blog/2026/07/29/lorenz-and-little.html",
      "title": "Lorenz and Little: How Much Does Your Tail Cost?",
      "summary_text": "Lorenz and Little: How Much Does Your Tail Cost? Lorenz and Little sounds like hipster burger bar from 2015. It’s time for Marc’s Amateur Statistics Corner! Today: why I pay a lot of attention to tail latency when optimizing cost. I’ve written before on the importance of tail latency for customer experience (e.g. in 2026, 2021, and 2021, and 2017). Today, I want to talk about tail latency from the perspective of cost and capacity. Like many system operators, I think about tail latency using…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2026-07-29T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:v-X8zp3FdPAajFl0nCbHffUG5ImRCmRr0gmgfuhfmYf3PlN9-1YVFuKC71PqFrLlKVypLrPGomoqcICYItCyCA"
    },
    {
      "kind": "entry",
      "id": "itl:753f4f04a13b5c32e399c450d06e1e68",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2026/07/19/dsql-paper",
      "url": "http://brooker.co.za/blog/2026/07/19/dsql-paper.html",
      "title": "Aurora DSQL: Scalable, Multi-Region OLTP",
      "summary_text": "Aurora DSQL: Scalable, Multi-Region OLTP A paper! Our new paper, Aurora DSQL: Scalable, Multi-Region OLTP, is now available on Arxiv. I’m excited about this one: it’s a fully end-to-end look at how Aurora DSQL works, from query processing, to transactions, to replication, to the control plane. We’ve shared most of this content before in other forms, on this blog, on Marc Bowes’ Blog, Werner’s Blog, in talks, and on the AWS blog. But this version covers all the ground, all in one place. You…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2026-07-19T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:jmSVmnljaJ0zEwseN5n6Kk1oDWjhduE8fUfryKJYoPACyKQP2iZSyzJKdY1t1DuFqKEdH4Quwnq6Bsff5AzrAg"
    },
    {
      "kind": "entry",
      "id": "itl:da4d856623cc6fa4042f6a648f5b8041",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2026/06/19/waiting",
      "url": "http://brooker.co.za/blog/2026/06/19/waiting.html",
      "title": "Meet Alice. Alice is impatient.",
      "summary_text": "Meet Alice. Alice is impatient. What do you mean? Meet Alice. Alice uses your web service. Alice, like most humans, measures her time in seconds and minutes. Alice says your service is slow. You tell Alice that the mean request to your service completes in 100ms, but Alice says that her mean wait time is 1s. You’re both right. Meet Alex. Alex uses your web service. Alex, like most humans, measures his time in seconds and minutes. Alex says that when you have outages, they last a long time and…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2026-06-19T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:8OiVS6HXEEI5ZXlEampOxH_F18sBkHRKl2NP_PwJyaZySgMj-84I2e__qkUJk9qlSShc-BwMPMR8B41phPpYCQ"
    },
    {
      "kind": "entry",
      "id": "itl:c6299d5fac8bccee7da9ae92f0e743da",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2026/06/18/my-blog-and-ai",
      "url": "http://brooker.co.za/blog/2026/06/18/my-blog-and-ai.html",
      "title": "Is this blog written by AI?",
      "summary_text": "Is this blog written by AI? No. None of the human-readable text on this blog is written by AI, and I have no plans to change that. The weird grammar, incorrect assumptions, spelling errors, and annoying tics are all mine. Including the em dashes. I don’t use LLMs for writing. On this blog, or in my professional life. I use agents extensively for brainstorming, research, summarizing, checking facts, handling markup, finding references, analyzing data, and so on. But I think that asking people…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2026-06-18T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:xr7Tg4Wld8rYjbfGteDVBGtpUFSxk1FbEUa4WGsKeNcm8rXPwGEgC0BclZ3eVrgHFZYL1zJgiku0GRjlpyGaDQ"
    },
    {
      "kind": "entry",
      "id": "itl:708da6d372a105604acfcfc10b484dda",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2026/05/20/hypothesis",
      "url": "http://brooker.co.za/blog/2026/05/20/hypothesis.html",
      "title": "Agentic software development hypothesis",
      "summary_text": "Agentic software development hypothesis This is the quality content you come here for, right? Agentic Software Development Hypothesis: Weak form: Any coding task for which a complete specification is available will become trivial. Strong form: Any coding task for which a deterministic oracle is available will become trivial. First objection: Few meaningful tasks have a complete specification. Second objection: Most oracles aren’t deterministic. Strongest form: Any coding task for which a…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2026-05-20T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:OpwFyHwht3He5BdF3FbV8kSD22hoRyw8vGKASGVfzHSNe9VhLU53gOXQoinfXPgEVoHR3oTWNdep4mvM-rvgBA"
    },
    {
      "kind": "entry",
      "id": "itl:8cf40cb497ac45234b95925a7fb95575",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2026/05/18/whats-easy-whats-hard",
      "url": "http://brooker.co.za/blog/2026/05/18/whats-easy-whats-hard.html",
      "title": "What's Easy Now? What's Hard Now?",
      "summary_text": "What’s Easy Now? What’s Hard Now? Take it easy. This is the fourth in a series about how AI is changing software development, after It’s time to be right., What about juniors?, and My heuristics are wrong. What now?. It stands alone, but if you found this interesting you may also find those interesting. I’ve been spending a lot of time thinking about the shape of the capabilities of coding agents. What they’re good at now, what they’re going to be good at. What they’re bad at now, how much of…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2026-05-18T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:tDp6MoFUXQUX-tVdQU9Nt0yN_VOPsAOq1QRWjUSjwnszRWUtVE3OZ0OBQ05ho9MSfleCA6bxArcrRYa2g6tCCg"
    },
    {
      "kind": "entry",
      "id": "itl:65d220b652d1d21480827e2ebb1d332e",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2026/04/30/be-right",
      "url": "http://brooker.co.za/blog/2026/04/30/be-right.html",
      "title": "It's time to be right.",
      "summary_text": "It’s time to be right. Outcomes continue to matter. Earlier this week, I spoke at AI Dev 26. This is what I spoke about there. I’ve been making money, in some form, building software for nearly 30 years. The last five months have been the most exciting of that entire time. I’m extremely optimistic about the future of software, and the future of software engineering as a field. But I have a hypothesis about agentic AI for development, and for knowledge work broadly: in future, the size of the…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2026-04-30T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:ARmi6WGFJg-gPX8aivcfv7han6ZlLLKWtWRL7rgTUe4VQOK9i4Tig2yZ0f559JNttHx9oIxbG9L-L8hUCq_xBg"
    },
    {
      "kind": "entry",
      "id": "itl:ad34bfe61edcdf0d763b0bc8152734e1",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2026/04/09/waterfall-vs-spec",
      "url": "http://brooker.co.za/blog/2026/04/09/waterfall-vs-spec.html",
      "title": "Spec Driven Development isn't Waterfall",
      "summary_text": "Spec Driven Development isn’t Waterfall Write down what you mean. After spending a few months writing (e.g. on the Kiro Blog), and speaking (e.g. Real Python Podcast, SE Radio) about spec-driven development, I’ve noticed a common misconception: spec driven development is a return to a waterfall style of software development. Specification driven development (in Kiro, for example) isn’t about pulling designs up-front, it’s about pulling designs up. Making specifications explicit, versioned,…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2026-04-09T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:7XRyL3IccWi_7AjScUWF1txevNlCAADijFJPOP71f9ocaoWRXryx98RGjbCyGIruEfRxRVLw8EDxVunQEOEDBA"
    },
    {
      "kind": "entry",
      "id": "itl:b9b72f6af79c761afa926ae8301de304",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2026/03/25/ic-junior",
      "url": "http://brooker.co.za/blog/2026/03/25/ic-junior.html",
      "title": "What about juniors?",
      "summary_text": "What about juniors? Start at the beginning. Last week I wrote about how the role of the most senior tech ICs has changed. Today, I wanted to share some thoughts on a more difficult topic: how the role of junior software engineers, folks just starting out on their career, has changed or will change. First, the good news. In last week’s post, I wrote this about senior folks: It’s hard to admit where you’re wrong. It’s hard to go back to being a beginner. Junior engineers don’t have this problem.…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2026-03-25T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:pX5ssZw55uvK1WpuUyv9r6oabWCyuuYfL1sT8F-NlxpdLVdmSzkCz_JrnPdk2wjlEowIphCX-d93fa4gkt6tCg"
    },
    {
      "kind": "entry",
      "id": "itl:d07a72077d6b2a6352c9f6d2716f6f4c",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2026/03/20/ic-leadership",
      "url": "http://brooker.co.za/blog/2026/03/20/ic-leadership.html",
      "title": "My heuristics are wrong. What now?",
      "summary_text": "My heuristics are wrong. What now? More words. More meaning? Some people who ask me for advice get a lot of words in reply. Sometimes, those responses aren’t specific to my particular workplace, and so I share them here. In the past, I’ve written about echo chambers, writing, writing for an audience, time management, and getting big things done. Do you remember Cool Runnings? In the movie, John Candy is a retired bobsled champion, who uses his experience, connections, and lovable curmudgeon…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2026-03-20T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:RmjG-Fknnk2K2vTFKg3M_mL-QOpGiK1WQbHjnBr2Nz9OK8WRg3MHAinkc-cgkAn8Fx-6btH5XFVTh2atNA9pAQ"
    },
    {
      "kind": "entry",
      "id": "itl:52aa4efa5a3b07d94a4640050ac43207",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2026/03/18/apprentice",
      "url": "http://brooker.co.za/blog/2026/03/18/apprentice.html",
      "title": "Music To Build Agents By",
      "summary_text": "Music To Build Agents By I don't have this problem, because I don't use a mouse. Press play, then start reading: Want to learn how to think about agent policy? Start with Goethe’s Der Zauberlehrling. So come along, you old broomstick! Dress yourself in rotten rags! You’ve long been a servant; Obey my orders now! When I talk to customers and teams around me about agents and agent policy, and the work we’re doing on AgentCore Policy (now GA) and Strands Steering, I hear a lot of folks worried…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2026-03-18T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:LI7ezJl8Xo7jyRTNsSq704t1ELL0Lr4rgpZK2GQVdI65CzsUoKdQPVT8OmpeKUa-8tYiGNWOJKJKbHU2qIcNAQ"
    },
    {
      "kind": "entry",
      "id": "itl:97d74b4e967571ff0ccea5980cc4bd30",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2026/02/25/sfq",
      "url": "http://brooker.co.za/blog/2026/02/25/sfq.html",
      "title": "SFQ: Simple, Stateless, Stochastic Fairness",
      "summary_text": "SFQ: Simple, Stateless, Stochastic Fairness Roll the dice. Paul E. McKenney’s 1990 paper Stochastic Fairness Queuing contains one of my favorite little algorithms for distributed systems. Stochastic Fairness Queuing is a way to stochastically isolate workloads from different customers in a way that significantly mitigates the effects of noisy neighbors, with O(1) queues and O(1) time. McKenney starts by describing Fairness Queuing (or queue per client): This fairness-queuing algorithm operates…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2026-02-25T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:aHly2mLyBd9cIVEfIV5rOfX0ethNJQoBJlvU7FbWLDu2NPbJBCNdRVwh5j0D8QQIdo0a57p4E9ovR2COlaDAAg"
    },
    {
      "kind": "entry",
      "id": "itl:c96e3515a4e9c5f7b3d4e8169e20300c",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2026/02/07/you-are-here",
      "url": "http://brooker.co.za/blog/2026/02/07/you-are-here.html",
      "title": "You Are Here",
      "summary_text": "You Are Here Where to next? The cost of turning written business logic into code has dropped to zero. Or, at best, near-zero. The cost of integrating services and libraries, the plumbing of the code world, has dropped to zero. Or, at best, near-zero. The cost of building efficient, reliable, secure, end-to-end systems is starting to drop, but slowly. Where does that leave those of us who have built careers in technology? Our road diverges. Not into the undergrowth of a wood, but into a dense…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2026-02-07T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:q9XoFDWRIXVjEPrX95LvKFeE05lPZK01YzxAvP1JZaODWB4n5XOS389GsHX1KEY8OMlhfFGiRh89JBxe16-rBA"
    },
    {
      "kind": "entry",
      "id": "itl:42634882e6f14e3fcba5bf408d586e4d",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2026/01/21/pass-k",
      "url": "http://brooker.co.za/blog/2026/01/21/pass-k.html",
      "title": "Pass@k is Mostly Bunk",
      "summary_text": "Pass@k is Mostly Bunk Exponentially better results? I'll take three! Measuring the success of AI agents isn’t easy. It’s very sensitive to what success means, it can require a lot of samples, its highly context sensitive. Generally hard. So it doesn’t help that one of the most common metrics used for agents is (mostly) bunk. I’m talking about pass@k. What is pass@k? It’s the probability that at least one of k different attempts will succeed. A six-sided die, where pass means rolling a 6, has a…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2026-01-21T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:YVi0gTcPpFLtjVcejEx73VcWxmC6LHNbEkLXycp0WE4gcmNaheB4KVOWoFyb5uphhu7w2xSnMnYE35PTglgUAw"
    },
    {
      "kind": "entry",
      "id": "itl:128b5d017e33daf54bae61da6c3fb17d",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2026/01/12/agent-box",
      "url": "http://brooker.co.za/blog/2026/01/12/agent-box.html",
      "title": "Agent Safety is a Box",
      "summary_text": "Agent Safety is a Box Keep a lid on it. Before we start, let’s cover some terms so we’re thinking about the same thing. This is a post about AI agents, which I’ll define (riffing off Simon Willison1) as: An AI agent runs models and tools in a loop to achieve a goal. Here, goals can include coding, customer service, proving theorems, cloud operations, or many other things. These agents can be interactive or one-shot; called by humans, other agents, or traditional computer systems; local or…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2026-01-12T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:TXRZk_YsY-chv3ubKaiFM8zW--o--JHMT1fXV2mfrfkacfET8FJSnv3IhnLh6NMQTdF158ZJH7W33OWLAgcvBw"
    },
    {
      "kind": "entry",
      "id": "itl:73917a2ef19ff69197ac571395b00e4f",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2025/12/16/natural-language",
      "url": "http://brooker.co.za/blog/2025/12/16/natural-language.html",
      "title": "On the success of 'natural language programming'",
      "summary_text": "On the success of ‘natural language programming’ Specifications, in plain speech. I believe that specification is the future of programming. Over the last four decades, we’ve seen the practice of building programs, and software systems grow closer and closer to the practice of specification. Details of the implementation, from layout in memory and disk, to layout in entire data centers, to algorithm and data structure choice, have become more and more abstract. Most application builders aren’t…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2025-12-16T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:Kf-imZ48oCYHU30pWYvnBQKNpnwcKq9jedyq0Qy2w0oKwo0lSIFeNSjU80JFrojNAn2P63G0OJHptqlAmKWKCA"
    },
    {
      "kind": "entry",
      "id": "itl:2c65f43bfbd5b694488058e6e3756f94",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2025/12/15/database-for-ssd",
      "url": "http://brooker.co.za/blog/2025/12/15/database-for-ssd.html",
      "title": "What Does a Database for SSDs Look Like?",
      "summary_text": "What Does a Database for SSDs Look Like? Maybe not what you think. Over on X, Ben Dicken asked: What does a relational database designed specifically for local SSDs look like? Postgres, MySQL, SQLite and many others were invented in the 90s and 00s, the era of spinning disks. A local NVMe SSD has ~1000x improvement in both throughput and latency. Design decisions like write-ahead logs, large page sizes, and buffering table writes in bulk were built around disks where I/O was SLOW, and where…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2025-12-15T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:n1igdRT-cfGuMKlveSh77szPGKvoprrfby63Cs0gU6zLFulGqM_hzkDxlScUQuZg_bZ_Z6tuXkX4SoUCFecnDQ"
    },
    {
      "kind": "entry",
      "id": "itl:5c1d8053743687cb69628575336cb5b3",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2025/11/20/what-now",
      "url": "http://brooker.co.za/blog/2025/11/20/what-now.html",
      "title": "What Now? Handling Errors in Large Systems",
      "summary_text": "What Now? Handling Errors in Large Systems More options means more choices. Cloudflare’s deep postmortem for their November 18 outage triggered a ton of online chatter about error handling, caused by a single line in the postmortem: .unwrap() If you’re not familiar with Rust, you need to know about Result, a kind of struct that can contain either a successful result, or an error. unwrap says basically “return the successful results if there is one, otherwise crash the program”1. You can think…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2025-11-20T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:LMkbKzdOi30Zbvy34I11oob-VmDB5uIINUIoe0GGn4p-Gsqcd2FzlVNuovOwo4XWdUKPHBRcRidVogmoai6HBg"
    },
    {
      "kind": "entry",
      "id": "itl:85a396b1ccf53687782e6d8a4ddc10c5",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2025/11/18/consistency",
      "url": "http://brooker.co.za/blog/2025/11/18/consistency.html",
      "title": "Why Strong Consistency?",
      "summary_text": "Why Strong Consistency? Eventual consistency makes your life harder. When I started at AWS in 2008, we ran the EC2 control plane on a tree of MySQL databases: a primary to handle writes, a secondary to take over from the primary, a handful of read replicas to scale reads, and some extra replicas for doing latency-insensitive reporting stuff. All of thing was linked together with MySQL’s statement-based replication. It worked pretty well day to day, but two major areas of pain have stuck with…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2025-11-18T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:wHN-YzMEbe0X23yRGfUEpl-6psUrBkkZWxd1FrPWjWnilFpKZYaEDl1bVdDPDcODZ2RlXPNug4VnBXpL8-0MBQ"
    },
    {
      "kind": "entry",
      "id": "itl:c757ff8140c1d0e4e796d75138d23346",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2025/11/02/thinking-dsql",
      "url": "http://brooker.co.za/blog/2025/11/02/thinking-dsql.html",
      "title": "DSQL: Simplifying Architectures",
      "summary_text": "DSQL: Simplifying Architectures Complexity is a choice. While we were designing and building Aurora DSQL, we spent a lot of time thinking about our experience building and running database-backed systems. We saw that building great, fast, cost-effective, highly-available, systems was harder than it needed to be. We wanted to make it easier. Today, I want to discuss some of Aurora DSQL’s features, and how I think they come together to make your life, as an application, service, or website…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2025-11-02T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:wQotrCy-i7ZwC-pJ5KJ3C-RXU9lAATUq6Z-KKTYcxgAUMzVpgh4ACtk0OTJt71km7nz3sMCrSsd3AIGwndIYDA"
    },
    {
      "kind": "entry",
      "id": "itl:008983049a77fcdba89bb0bb0560721f",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2025/10/22/uuidv7",
      "url": "http://brooker.co.za/blog/2025/10/22/uuidv7.html",
      "title": "Fixing UUIDv7 (for database use-cases)",
      "summary_text": "Fixing UUIDv7 (for database use-cases) How do I even balance a V7? RFC9562 defines UUID Version 7. This has made a lot of people very angry and been widely regarded as a bad move1. More seriously, UUIDv7 has received a lot of criticism, despite seemingly achieving what it set out to do. The legitimate criticism seems to be on a few points. V7 UUIDs: Leak information (namely the server timestamp). Are a bad choice for cases where security or operational requirements require UUIDs that are hard…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2025-10-22T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:0WdWKluoinU4XWwvO9MT6sZd7RrKbbHuvtHe43tsr3V2a3SXzAhXJzSyhjNQinO2MxWfAtDK_ZDPpHglqaghCQ"
    },
    {
      "kind": "entry",
      "id": "itl:f5df6375eb6d0e981858ef167738c875",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2025/10/12/barbarians",
      "url": "http://brooker.co.za/blog/2025/10/12/barbarians.html",
      "title": "Is Systems Research Really Just About Making Numbers Bigger?",
      "summary_text": "Is Systems Research Really Just About Making Numbers Bigger? The Barbarian F.C. of systems research would be pretty cool. Lots of folks online have been talking about Barbarians at the Gate: How AI is Upending Systems Research by Cheng, Liu, Pan, et al this week. Maybe unsurprisingly, given the fact that I work in AI for my day job, and both consume and produce systems research, I found it super interesting. Perhaps the most interesting discussion, however, isn’t about AI at all. It’s about…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2025-10-12T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:Z1ZxboY_HK1tEitsdv4V1Uwu4SgfHEh4qb8jDGJsjqj4UbX0ak2apFhrYuG0Llluul7JTjjVI4jydCmxJUJzBg"
    },
    {
      "kind": "entry",
      "id": "itl:d8a6c08b63c338002ec3e04c122bbc58",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2025/10/05/locality",
      "url": "http://brooker.co.za/blog/2025/10/05/locality.html",
      "title": "Locality, and Temporal-Spatial Hypothesis",
      "summary_text": "Locality, and Temporal-Spatial Hypothesis Good fences make good neighbors? Last week at PGConf NYC, I had the pleasure of hearing Andres Freund talking about the great work he’s been doing to bring async IO to Postgres 18. One particular result caught my eye: a large difference in performance between forward and reverse scans, seemingly driven by read ahead1. The short version is that IO layers (like Linux’s) optimize performance by proactively pre-fetching data ahead of the current read point…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2025-10-05T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:6zsO5U_lu9LzCh9khaP44Y_P11G4ADi3fNLtm2l8m9QB1g0ECfO54xOabEtxP_ivOTDEqCmUoROCigqtHz5WCA"
    },
    {
      "kind": "entry",
      "id": "itl:8a2c60b220ea717c331bec4223dc08ce",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2025/09/18/firecracker",
      "url": "http://brooker.co.za/blog/2025/09/18/firecracker.html",
      "title": "Seven Years of Firecracker",
      "summary_text": "Seven Years of Firecracker Time flies like an arrow. Fruit flies like a banana. Back at re:Invent 2018, we shared Firecracker with the world. Firecracker is open source software that makes it easy to create and manage small virtual machines. At the time, we talked about Firecracker as one of the key technologies behind AWS Lambda, including how it’d allowed us to make Lambda faster, more efficient, and more secure. A couple years later, we published Firecracker: Lightweight Virtualization for…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2025-09-18T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:vxkNGMclrEOh3IwJmR7_nyMr-6UeEdci02lsPW5o36sR1ubejbRU55BPAv8UjmAJ31PRTqq5hbrIaO26WW5fDw"
    },
    {
      "kind": "entry",
      "id": "itl:0a820d4b208a53bb69bc5c190e9e512c",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2025/08/15/dynamo-dynamodb-dsql",
      "url": "http://brooker.co.za/blog/2025/08/15/dynamo-dynamodb-dsql.html",
      "title": "Dynamo, DynamoDB, and Aurora DSQL",
      "summary_text": "Dynamo, DynamoDB, and Aurora DSQL Names are hard, ok? People often ask me about the architectural relationship between Amazon Dynamo (as described in the classic 2007 SOSP paper), Amazon DynamoDB (the serverless distributed NoSQL database from AWS), and Aurora DSQL (the serverless distributed SQL database from AWS). There’s a ton to say on the topic, but I’ll start off on comparing how the systems achieve a few key properties. The key references for this post are: For Dynamo, Dynamo: Amazon’s…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2025-08-15T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:hBSy7oCK7V58XESLWyTs_z5ouegCTeYe3Je8-23CXaz2GbVUirhbK3Og4Yd0-f9qls8cRdJtKzBNfeDT4fOtAA"
    },
    {
      "kind": "entry",
      "id": "itl:82a173bfc17da005d0019343f01c0072",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2025/08/12/llms-as-components",
      "url": "http://brooker.co.za/blog/2025/08/12/llms-as-components.html",
      "title": "LLMs as Parts of Systems",
      "summary_text": "LLMs as Parts of Systems Towers of Hanoi is a boring game, anyway. Over on the Kiro blog, I wrote a post about Kiro and the future of AI spec-driven software development, looking at where I think the space of AI-agent-powered development tools is going. In that post, I made a bit of cheeky oblique reference to a topic I think is super important. I asked Kiro to build a Towers of Hanoi game. It’s an oblique reference to Apple’s The Illusion of Thinking paper, and the discourse that followed it.…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2025-08-12T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:Q4xCaB68i4B5tTn_SjIDG7JUStoNis4PubsES4OrNkJY7nXmM1DhUJVA6voFxlhY9EymZSTB42axyAjgdNZeDg"
    },
    {
      "kind": "entry",
      "id": "itl:f78ef676d518b82a79d38b34a20edec4",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2025/06/20/career",
      "url": "http://brooker.co.za/blog/2025/06/20/career.html",
      "title": "Career advice, or something like it",
      "summary_text": "Career advice, or something like it Cynicism is bad. If I could offer you a single piece of career advice, it’s this: avoid negativity echo chambers. Every organization and industry has watering holes where the whiners hang out. The cynical. The jaded. These spots feel attractive. Everybody has something they can complain about, and complaining is fun. These places are inviting and inclusive: as long as you’re whining, or complaining, or cynical, you’re in. If you’re positive, optimistic, or…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2025-06-20T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:GWz-JJzFyT-8hROcOzI-3MJiOjVjsnLSqloL6mxKRfnChGVTeJ81FrHSFkNrR8_RJr737wnYN34hxuBFhHm6DQ"
    },
    {
      "kind": "entry",
      "id": "itl:0b6d0fa1719cdf75c07e45b58045196c",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2025/06/02/hotos",
      "url": "http://brooker.co.za/blog/2025/06/02/hotos.html",
      "title": "Systems Fun at HotOS",
      "summary_text": "Systems Fun at HotOS One day somebody will tell me what systems means. Last week I attended HotOS1 for the first time. It was super fun. Just the kind of conference I like: single-track, a mix of academic and industry, a mix of normal practical ideas and less-normal less-practical big thinking. I went partially because a colleague twisted my arm, and partially because of this line in the CFP: The program committee will explicitly favor papers likely to stimulate reflection and discussion. That…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2025-06-02T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:ZopvFFhtHJvtQpa_iqtjPBdnpnWGrcJSlJAqIz5SDwBV219swghbSKDCgPYc1x2--zYrck1G2PN5N9rAVT3HDw"
    },
    {
      "kind": "entry",
      "id": "itl:a2cd33fd75f702f3bd5ddd9edc5b2923",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2025/05/20/icpe",
      "url": "http://brooker.co.za/blog/2025/05/20/icpe.html",
      "title": "Good Performance for Bad Days",
      "summary_text": "Good Performance for Bad Days Good things are good, one finds. Two weeks ago, I flew to Toronto to give one of the keynotes at the International Conference on Performance Evaluation. It was fun. Smart people. Cool dark squirrels. Interesting conversations. The core of what I tried to communicate is that, in my view, a lot of the performance evaluation community is overly focused on happy case performance (throughput, latency, scalability), and not focusing as much as we need to on performance…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2025-05-20T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:DEsPpSFoLY5Z3Vl2qXsKT-KGoMlJNkkGBIMTWsKEGJouERa97h0AcBcA2k8_TEdRaTadzOESC0w8YDGWMD9XAw"
    },
    {
      "kind": "entry",
      "id": "itl:6c1b08dad1e97021e92b6cedacddd05e",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2025/04/17/decomposing",
      "url": "http://brooker.co.za/blog/2025/04/17/decomposing.html",
      "title": "Decomposing Aurora DSQL",
      "summary_text": "Decomposing Aurora DSQL Riffing, I guess. Earlier today, Alex Miller wrote an excellent blog post titled Decomposing Transaction Systems. It’s one of the best things I’ve read about transactions this year, maybe the best. You should read it now. In the post, Alex breaks transactions down like this: Every transactional system does four things: It executes transactions. It orders transactions. It validates transactions. It persists transactions. then describes how these steps map to traditional…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2025-04-17T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:wfUSwPMoDxWgJkSZ0841sKFrdR-eKODxa_-M4mDV0IKxK9smSINodUJfti8nosiH2z6-crtrzG4MAhfz3105Bg"
    },
    {
      "kind": "entry",
      "id": "itl:9f658453fd2139dd4a78bdd6783d4ac1",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2025/03/25/two-queues",
      "url": "http://brooker.co.za/blog/2025/03/25/two-queues.html",
      "title": "One or Two? How Many Queues?",
      "summary_text": "One or Two? How Many Queues? Very applied queue theory. There’s a well-known rule of thumb that one queue is better than two. When you’ve got people waiting to check out at the supermarket, having a single shared queue improves utilization and reduces wait times. The reason for this is pretty simple: it avoids the case where somebody is waiting in a queue while there’s a checker available to do the work. It also saves the sanity of the person standing behind a cheque writer or expired coupon…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2025-03-25T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:qIF-e8uoKFVxTcjYNRGqcN8ruIiUbLsteIOa5Wi1UQg3S-AzpDOIvJ6uPJu1ILWRch8RRBWx9NulFp0va3J7Cg"
    },
    {
      "kind": "entry",
      "id": "itl:138a02335f635a7d1736caf621caa4b3",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2025/02/05/feketes",
      "url": "http://brooker.co.za/blog/2025/02/05/feketes.html",
      "title": "What Fekete's Anomaly Can Teach Us About Isolation",
      "summary_text": "What Fekete’s Anomaly Can Teach Us About Isolation Is it just fancy write skew? In the first draft of yesterday’s post, the example I used was one that showed Fekete’s anomaly. After drafting, I realized the example distracted too much from the story. But there’s still something I want to say about the anomaly, and so now we’re here. What is Fekete’s anomaly? It’s an example of a snapshot isolation behavior first described in Fekete, O’Neil, and O’Neil’s paper A Read-Only Transaction Anomaly…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2025-02-05T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:jbOBdwtOlFWLz2G1QdtS71GFrzBuL3qIfMBFqfVzQN0HfWQ4klOcOz0oF4xj4DKXkAK807D1lGb66CAwRWvFDQ"
    },
    {
      "kind": "entry",
      "id": "itl:3f4f0640fa875fe6b72735fb31868341",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2025/02/04/versioning",
      "url": "http://brooker.co.za/blog/2025/02/04/versioning.html",
      "title": "Versioning versus Coordination",
      "summary_text": "Versioning versus Coordination Spoiler: Versioning Wins. Today, we’re going to build a little database system. For availability, latency, and scalability, we’re going to divide our data into multiple shards, have multiple replicas of each shard, and allow multiple concurrent queries. As a block diagram, it’s going to look something like this: Next, borrowing heavily from Hermitage, we’re going to run some SQL. begin; -- T0 create table test (id int primary key, value int); -- T0 insert into…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2025-02-04T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:lqkwz34HGz86hbp_0ZNZkCzU4wYFWn8gCjH-Jpixp1uf3jBK4vUwK_m_JCPB4-vE5v_xfXt_IArO7ZiYkT9PAg"
    },
    {
      "kind": "entry",
      "id": "itl:d9eb18a940bcbe060d38c909f33e875b",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2024/12/17/occ-and-isolation",
      "url": "http://brooker.co.za/blog/2024/12/17/occ-and-isolation.html",
      "title": "Snapshot Isolation vs Serializability",
      "summary_text": "Snapshot Isolation vs Serializability Getting into some fundamentals. In my re:Invent talk on the internals of Aurora DSQL I mentioned that I think snapshot isolation is a sweet spot in the database isolation spectrum for most kinds of applications. Today, I want to dive in a little deeper into why I think that, and some of the trade-offs of going stronger and weaker. This post is going to be a little deeper than the last few. If you’re not deeply familiar with SQL’s isolation levels, I…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2024-12-17T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:6hvs_E9GhIA6HNrThtgAR6wcgGGbRUZLqK0wYlvlfcxu_4Qve_-jO8HxKfgZ_qm-7-rYznMRtjrtaovaeh0NAw"
    },
    {
      "kind": "entry",
      "id": "itl:9ddd45903a29d39179bf00bb67eea575",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2024/12/06/inside-dsql-cap",
      "url": "http://brooker.co.za/blog/2024/12/06/inside-dsql-cap.html",
      "title": "DSQL Vignette: Wait! Isn't That Impossible?",
      "summary_text": "DSQL Vignette: Wait! Isn’t That Impossible? Laws of physics are real. In today’s post, I’m going to look at how Aurora DSQL is designed for availability, and how we work within the constraints of the laws of physics. If you’d like to learn more about the product first, check out the official documentation, which is always a great place to go for the latest information on Aurora DSQL, and how to fit it into your architecture. In yesterday’s post, I mentioned that Aurora DSQL is designed to…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2024-12-06T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:6CGPLVSsaSzfaRN4G1V2ueBy783T-TcyfT_HHtYzbk2lcl6_odURJSNJbQs4MPB7BZeONuygCBIHcjegbTtlDg"
    },
    {
      "kind": "entry",
      "id": "itl:88e732181864f39163493f908dd446f2",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2024/12/05/inside-dsql-writes",
      "url": "http://brooker.co.za/blog/2024/12/05/inside-dsql-writes.html",
      "title": "DSQL Vignette: Transactions and Durability",
      "summary_text": "DSQL Vignette: Transactions and Durability The hard half of a database system? In today’s post, I’m going to look at the other half of what’s under the covers of Aurora DSQL, our new scalable, active-active, SQL database. If you’d like to learn more about the product first, check out the official documentation, which is always a great place to go for the latest information on Aurora DSQL, and how to fit it into your architecture. Today, we’re going to focus on writes (INSERTS, UPDATES, etc),…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2024-12-05T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:Rl0YGM6Z4GTfGS6cI2ASkoHJbIriYdXMaAvVPh4CrO39_PyuMyeggjMb374cocqGlGUEPMXoLNwKa-DWppQKDA"
    },
    {
      "kind": "entry",
      "id": "itl:344badd0fbc4f4b5343c5d50f6d7eb95",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2024/12/04/inside-dsql",
      "url": "http://brooker.co.za/blog/2024/12/04/inside-dsql.html",
      "title": "DSQL Vignette: Reads and Compute",
      "summary_text": "DSQL Vignette: Reads and Compute The easy half of a database system? In today’s post, I’m going to look at half of what’s under the covers of Aurora DSQL, our new scalable, active-active, SQL database. If you’d like to learn more about the product first, check out the official documentation, which is always a great place to go for the latest information on Aurora DSQL, and how to fit it into your architecture. Today, we’re going to focus on running SQL and doing transactional reads. But first,…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2024-12-04T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:9uoYVQSaFvaAqrNG52DCtmjOJZrILJtv5sZTsSQphLOquuuUv-ay6Gda0Fz01tlMxdeqzxZYwr5sZhCTjsFcDw"
    },
    {
      "kind": "entry",
      "id": "itl:11362c74e319db03ac7a35f1a7dba98f",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2024/12/03/aurora-dsql",
      "url": "http://brooker.co.za/blog/2024/12/03/aurora-dsql.html",
      "title": "DSQL Vignette: Aurora DSQL, and A Personal Story",
      "summary_text": "DSQL Vignette: Aurora DSQL, and A Personal Story It's happening. In this morning’s re:Invent keynote, Matt Garman announced Aurora DSQL. We’re all excited, and some extremely excited, to have this preview release in customers’ hands. Over the next few days, I’m going to be writing a few posts about what DSQL is, how it works, and how to make the best use of it. This post is going to look at the product itself, and a little bit of a personal story. The official AWS documentation for Aurora DSQL…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2024-12-03T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:3_nBU75wQZF4Y8S0oqqB7-3sBc9-o-uqll9Zt_QgrU_jPWJeFL1bpO1Taju6KK0-nQqwVYU0s9POJa-Oc9G4Cg"
    },
    {
      "kind": "entry",
      "id": "itl:2e6816404eb864c4ddd0c66956e4f082",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2024/11/14/lambda-ten-years",
      "url": "http://brooker.co.za/blog/2024/11/14/lambda-ten-years.html",
      "title": "Ten Years of AWS Lambda",
      "summary_text": "Ten Years of AWS Lambda Everything starts somewhere. Today, Werner Vogels shared his annotated version of the original AWS Lambda PRFAQ. This is a great inside look into how product development happens at AWS - the real working backwards process in action. This was, in some ways, the start of serverless computing2. Tim Wagner, Ajay Nair, and others really saw the future when they wrote this PRFAQ3. I wanted to take the opportunity to dive a little deeper into some of the things Werner…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2024-11-14T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:6UFKrW1URuqCgIYKQTJu5P_MbM6FqvUBhsLh86QKU2supjxMxsPeYoPr0SbzrSIcACou1wqhvwBc1-TOWoH6Cg"
    },
    {
      "kind": "entry",
      "id": "itl:0deb684847a2d5fa14a0c11d1d9c704d",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2024/08/14/gc-metastable",
      "url": "http://brooker.co.za/blog/2024/08/14/gc-metastable.html",
      "title": "Garbage Collection and Metastability",
      "summary_text": "Garbage Collection and Metastability Cleaning up is hard to do. I’ve written a lot about stability and metastability, but haven’t touched on one other common cause of metastability in large-scale systems: garbage collection. GC is great. Garbage collected languages like Javascript, Java, Python, and Go power a big chunk of the internet’s infrastructure. Until Rust came along, choosing memory safety typically implied choosing garbage collection. For almost all applications, languages with…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2024-08-14T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:0-SoqWZT2EOYbjkIdURtGRnXSrBzdpNY5vyCZbG1QvGl6XO4R7kB5Sgj4XkoKltoqo0ttfziDozxGpT-eTqCCg"
    },
    {
      "kind": "entry",
      "id": "itl:da53e9e10ce45266a96d6d5b682fed27",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2024/07/29/aurora-serverless",
      "url": "http://brooker.co.za/blog/2024/07/29/aurora-serverless.html",
      "title": "Resource Management in Aurora Serverless",
      "summary_text": "Resource Management in Aurora Serverless Systems, big and small. My favorite thing about distributed systems is how they allow us to solve problems at multiple levels: single process problems, single machine problems, multi-machine problems, and large-scale cluster problems. Our new paper Resource management in Aurora Serverless1 describes what this looks like in context of a large-scale running system: Amazon Aurora Serverless. What is Aurora Serverless? Aurora Serverless (or, rather, Aurora…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2024-07-29T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:splFPJovPsF4EAn4HlWFsP_RzBqOtmCGnc2hoOKORmvujlsNNQ8n75BNJFYBLeclIl7PBCI2BFO7gVAZ2eRpCg"
    },
    {
      "kind": "entry",
      "id": "itl:be9175ec5ca9694db7ce89d412a3066f",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2024/07/25/cap-again",
      "url": "http://brooker.co.za/blog/2024/07/25/cap-again.html",
      "title": "Let's Consign CAP to the Cabinet of Curiosities",
      "summary_text": "Let’s Consign CAP to the Cabinet of Curiosities CAP? Again? Still? Brewer’s CAP theorem, and Gilbert and Lynch’s formalization of it, is the first introduction to hard trade-offs for many distributed systems engineers. Going by the vast amounts of ink and bile spent on the topic, it is not unreasonable for new folks to conclude that it’s an important, foundational, idea. The reality is that CAP is nearly irrelevant for almost all engineers building cloud-style distributed systems, and…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2024-07-25T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:v7UNVElOAwCUzvAiV5G6WyYGXtDaoVI-SLztLhGA7sYFMxE3y0h0myy6FVjzmXL2bOpWUzLkyXE-cD4aNbmzBw"
    },
    {
      "kind": "entry",
      "id": "itl:5213f4aab89dece1335ad26cd4ada5e8",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2024/06/04/scale",
      "url": "http://brooker.co.za/blog/2024/06/04/scale.html",
      "title": "Not Just Scale",
      "summary_text": "Not Just Scale Bookmarking this so I can stop writing it over and over. It seems like everywhere I look on the internet these days, somebody’s making some form of the following argument: You don’t need distributed systems! Computers are so fast these days you can serve all your customers off a single machine! This argument is silly and reductive. But first, let’s look for the kernel of truth. One Machine Is All You Need? This argument is based on a kernel of truth: modern machines are…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2024-06-04T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:Pb565u07A_8IirVfnmiuBKZ921e9juhKuh_t8viqDp8y2ZIAtJDed0BtQhXUn2HiHmMx2NVvQ-bvAY8aywK3AQ"
    },
    {
      "kind": "entry",
      "id": "itl:64436906804e07a6c54f931806c92355",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2024/05/09/nagle",
      "url": "http://brooker.co.za/blog/2024/05/09/nagle.html",
      "title": "It's always TCP_NODELAY. Every damn time.",
      "summary_text": "It’s always TCP_NODELAY. Every damn time. It's not the 1980s anymore, thankfully. The first thing I check when debugging latency issues in distributed systems is whether TCP_NODELAY is enabled. And it’s not just me. Every distributed system builder I know has lost hours to latency issues quickly fixed by enabling this simple socket option, suggesting that the default behavior is wrong, and perhaps that the whole concept is outmoded. First, let’s be clear about what we’re talking about. There’s…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2024-05-09T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:iM-gY4kPUwa7DbC6VV8OxLf5gnAE-EYAgWx1_NWrFps7hG-6vax6_A7_yElRz4q3UZ8TLJTHNqda7JxKDdrmDA"
    },
    {
      "kind": "entry",
      "id": "itl:55e6804256cd8fb58e81ae7b90032c74",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2024/04/25/memorydb",
      "url": "http://brooker.co.za/blog/2024/04/25/memorydb.html",
      "title": "MemoryDB: Speed, Durability, and Composition.",
      "summary_text": "MemoryDB: Speed, Durability, and Composition. Blocks are fun. Earlier this week, my colleagues Yacine Taleb, Kevin McGehee, Nan Yan, Shawn Wang, Stefan Mueller, and Allen Samuels published Amazon MemoryDB: A fast and durable memory-first cloud database1. I’m excited about this paper, both because its a very cool system, and because it gives us an opportunity to talk about the power of composition in distributed systems, and about the power of distributed systems in general. But first, what is…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2024-04-25T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:YYzMVVfvz9yUwZGPXjsffYNDTvZHi4hlaznPFpwpShriaz1eU68lhBPj6gjZX0uu2e9iCcxZ_pAigxc_Mi0lAA"
    },
    {
      "kind": "entry",
      "id": "itl:c7e988e7d1f92da9ebf3ef89c25dd2ad",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2024/04/17/formal",
      "url": "http://brooker.co.za/blog/2024/04/17/formal.html",
      "title": "Formal Methods: Just Good Engineering Practice?",
      "summary_text": "Formal Methods: Just Good Engineering Practice? Yes. The answer is yes. In your face, Betteridge. Earlier this week, I did the keynote at TLA+ conf 2024 (watch the video or check out the slides). My message in the keynote was something I have believed to be true for a long time: formal methods are an important part of good software engineering practice. If you’re a software engineer, especially one working on large-scale systems, distributed systems, or critical low-level system, and are not…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2024-04-17T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:Xi0MH2tsD9k5Mt7DUBLlA6QbG-Mjls_jnMy3IUSA_tNmB-5VLcFojd3jP5B84YR_6kj5OejgqSlUXjYGgn2KCQ"
    },
    {
      "kind": "entry",
      "id": "itl:dcd4745c000d07d868493df5e66323ec",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2024/03/25/needles",
      "url": "http://brooker.co.za/blog/2024/03/25/needles.html",
      "title": "Finding Needles in a Haystack with Best-of-K",
      "summary_text": "Finding Needles in a Haystack with Best-of-K Keep track of those needles. As I’ve written about before, best of two and best of k are surprisingly powerful tools for load balancing in distributed systems. I have deployed them many times in large-scale production systems, and been happy with the performance nearly every time. There is one case where they don’t perform so well, though: when the bins are very limited in size. Reminder: Best-of-K Consider a load balancing problem in a distributed…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2024-03-25T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:6bhar2UBiwJDUp7cVHj0WZFFbiWKN6M338BEpISCy3KHE0_iDOiYjKvVzawwVMxexF8-HSVONF52gmr_9qVRCQ"
    },
    {
      "kind": "entry",
      "id": "itl:6fc282b373223830d0a24f8c04821b87",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2024/03/04/mousetrap",
      "url": "http://brooker.co.za/blog/2024/03/04/mousetrap.html",
      "title": "The Builder's Guide to Better Mousetraps",
      "summary_text": "The Builder’s Guide to Better Mousetraps A little rubric for making a tough decision. Some people who ask me for advice at work get very long responses. Sometimes, those responses aren’t specific to my particular workplace, and so I share them here. In the past, I’ve written about writing, writing for an audience, heuristics, getting big things done, and how to spend your time. This is another of those emails. So, you’re thinking of building a new thing. It’s going to be a lot like that other…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2024-03-04T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:vtzzwBlN-B_07yGH77OoiAW88H97SIut_joLEcAS_18Wd9eXym3dgeirv9V0e9y8bteXIJ6bhlew_AMmxh2jBQ"
    },
    {
      "kind": "entry",
      "id": "itl:fd967017ad7312a0f72e2c8b8447958c",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2024/02/12/parameters",
      "url": "http://brooker.co.za/blog/2024/02/12/parameters.html",
      "title": "Better Benchmarks Through Graphs",
      "summary_text": "Better Benchmarks Through Graphs Isn't the ambiguity in the word *graphs* fun? This is a blog post version of a talk I gave at the Northwest Database Society meeting last week. The slides are here, but I don’t believe the talk was recorded. I believe that one of the things that’s holding back databases as an engineering discipline (and why so much remains stubbornly opinion-based) is a lack of good benchmarks, especially ones available at the design stage. The gold standard is designing for…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2024-02-12T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:4IJhZcXv5yJUusG-KpH_rcIMggAgY5MmyBbuPOWqjapE4CIJl_ztfVBW2I84UdkaZG-jPDZh4HIQ90dhJOsADw"
    },
    {
      "kind": "entry",
      "id": "itl:81908290979fb1f536792a8f7d6f1a77",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2024/02/06/time",
      "url": "http://brooker.co.za/blog/2024/02/06/time.html",
      "title": "How Do You Spend Your Time?",
      "summary_text": "How Do You Spend Your Time? Career advice, or something like it. Some people who ask me for advice at work get very long responses. Sometimes, those responses aren’t specific to my particular workplace, and so I share them here. In the past, I’ve written about writing, writing for an audience, heuristics, and getting big things done. This is another of those emails. When we spoke, you mentioned that you weren’t happy with the things you were getting done. You thought you were productive, and…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2024-02-06T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:RAdb4ztBc3b2APwh1uh9gyGzpVUf7JFTocGDriCHWTvp0mQuVgZqP6uEB2bmO4I1S9LWIRHPUQrThbRmea3NAw"
    },
    {
      "kind": "entry",
      "id": "itl:bf24338f28048ff33ea3d24099745b4a",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2024/01/23/big-deal",
      "url": "http://brooker.co.za/blog/2024/01/23/big-deal.html",
      "title": "Pat's Big Deal, and Transaction Coordination",
      "summary_text": "Pat’s Big Deal, and Transaction Coordination Working together towards a common goal. I have a lot of opinions about Pat Helland’s CIDR’24 paper Scalable OLTP in the Cloud: What’s the BIG DEAL?1. Most importantly, I like the BIG DEAL that he proposes: Scalable apps don’t concurrently update the same key. Scalable DBs don’t coordinate across disjoint TXs updating different keys. In exchange for fulfilling their sides of this big deal2 the application gets a database that can scale6, and the…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2024-01-23T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:dqhClsLGn0Yh-wIO-FFEdKEXC25tA1iKgVGYuBng3miYdRLKLl7Rr5cWxGZI9si5ot9tLdd-xXvVpyAGiSggDg"
    },
    {
      "kind": "entry",
      "id": "itl:2052b21adf9d623e02b9e99863ba61d7",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2024/01/18/scalability",
      "url": "http://brooker.co.za/blog/2024/01/18/scalability.html",
      "title": "What is Scalability Anyway?",
      "summary_text": "What is Scalability Anyway? Do words mean things? Why? What does scalable mean? As systems designers, builders, and researchers, we use that word a lot. We kind of all use it to mean that same thing, but not super consistently. Some include scaling both up and down, some just up, and some just down. Some include both scaling on a box (vertical) and across boxes (horizontal), some just across boxes. Some include big rack-level systems, some don’t. Here’s my definition: A system is scalable in…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2024-01-18T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:vWIIN63Z-lUQEpQfnPMGW-4lEQv0PnUFa_H_rdtkqlTAUVM8Rgdvl4Jnm1LUCmBWl33oNcXkB1SRGBZ5MWK7CA"
    },
    {
      "kind": "entry",
      "id": "itl:fc884bec0a55195d3eee616764246d6b",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2023/12/15/sieve",
      "url": "http://brooker.co.za/blog/2023/12/15/sieve.html",
      "title": "Why Aren't We SIEVE-ing?",
      "summary_text": "Why Aren’t We SIEVE-ing? Captain, we are being scanned! Long-time readers of this blog will know that I have mixed feelings about caches. One on hand, caching is critical to the performance of systems at every layer, from CPUs to storage to whole distributed architectures. On the other hand, caching being this critical means that designers need to carefully consider what happens when the cache is emptied, and they don’t always do that well1. Because of how important caches are, I follow the…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2023-12-15T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:-h07f0JHY7kqC3kGe0ux2EIsV4ykKts45XQSb_9KJoezWTVjrft3x2Hm6hqS-MJ2Az_kE2F0BTDbdFK63ptXAA"
    },
    {
      "kind": "entry",
      "id": "itl:b51ec88cb33a041a120681b92b699b0c",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2023/11/27/about-time",
      "url": "http://brooker.co.za/blog/2023/11/27/about-time.html",
      "title": "It's About Time!",
      "summary_text": "It’s About Time! What's the time? Time to get a watch. My friend Al Vermeulen used to say time is for the amusement of humans1. Al’s sentiment is still the common one among distributed systems builders: real wall-clock physical time is great for human-consumption (like log timestamps and UI presentation), but shouldn’t be relied on by computer for things like actually affect the operation of the system. This remains a solid starting point, the right default position, but the picture has always…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2023-11-27T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:sBuZOek61qkKwV_Z1Ww9dK946qM5RrWioG4DXK1wAgpF4A9gfs5iXm1_6aikWqI40nCBs2dfIIsNLna8hbB2DQ"
    },
    {
      "kind": "entry",
      "id": "itl:6133bbb532ba6b6cadec84035b9418c7",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2023/10/18/optimism",
      "url": "http://brooker.co.za/blog/2023/10/18/optimism.html",
      "title": "Optimism vs Pessimism in Distributed Systems",
      "summary_text": "Optimism vs Pessimism in Distributed Systems What—Me Worry? Avoiding coordination is the one fundamental thing that allows us to build distributed systems that out-scale the performance of a single machine1. When we build systems that avoid coordinating, we end up building components that make assumptions about what other components are doing. This, too, is fundamental. If two components can’t check in with each other after every single step, they need to make assumptions about the ongoing…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2023-10-18T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:bUk1ay3uGNyP1WgIaPB9oDkxEiD15EIXdHJ3xvQZoTIXpiZ7mla3uQ5xp99NkeWAHBosrYBaH2MgNP587vxiDw"
    },
    {
      "kind": "entry",
      "id": "itl:0d11aa3781a8e5624ea0be20a24ff1c0",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2023/09/21/audience",
      "url": "http://brooker.co.za/blog/2023/09/21/audience.html",
      "title": "Writing For Somebody",
      "summary_text": "Writing For Somebody Who's there? Sometimes I write long emails to people at work. Sometimes those emails are generally interesting, and not work-specific at all. Sometimes I share those emails here on my blog. This may be one of those times. Always write for somebody. Always have an idea in your head, as you’re writing, who your writing is intended to communicate with. Sometimes, that’s a particular person. Your boss. A mentee, or mentor. Bob from legal. Sometimes it’s a group of people, or a…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2023-09-21T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:uAlGIpSaVHvuFiUOfdQWvx8mI0E9-oThSj4CGw9Y8DL8lNjbYzCmBImmbws-Kgh1HqInMnmF3gceSOUL7psNDA"
    },
    {
      "kind": "entry",
      "id": "itl:d7b4afcef26205b0ec814419b83a8e5b",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2023/09/08/exponential",
      "url": "http://brooker.co.za/blog/2023/09/08/exponential.html",
      "title": "Exponential Value at Linear Cost",
      "summary_text": "Exponential Value at Linear Cost What a deal! Binary search is kind a of a magical thing. With each additional search step, the size of the haystack we can search doubles. In other words, the value of a search is exponential in the amount of effort. That’s a great deal. There are a few similar deals like that in computing, but not many. How often, in life, do you get exponential value at linear cost? Here’s another important one: redundancy. If we have $N$ hosts, each with availability $A$,…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2023-09-08T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:h4jOL6YupAiU8PJHh8plQim_eiU298KjgASeGXDpli_Bs6jpqPqhhuU6EDd5Th8eJgg2jlRnHAc2MPDUB6PoAQ"
    },
    {
      "kind": "entry",
      "id": "itl:74ba0a8333eed77a3835484a529e8a52",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2023/08/25/party-time",
      "url": "http://brooker.co.za/blog/2023/08/25/party-time.html",
      "title": "On The Acoustics of Cocktail Parties",
      "summary_text": "On The Acoustics of Cocktail Parties Only parties of well-mannered guests will be considered. If you, like me, tend to practice punctual arrival at parties, you’ve likely noticed that most parties start out quiet. Folks are talking in small groups, using their normal voices, and having productive conversations. As more people arrive, the background noise increase. First a little, allowing guests to continue to use a conventional volume. Then, at some point, the background noise will exceed a…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2023-08-25T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:mQWbw4OHh-fJIWDJfYR4LlqNxWWP837jBrsXCI_z_Fxx2lqinf9HdTwR5C3b4lKBK7xkAhMOG2IijTzIZln8BQ"
    },
    {
      "kind": "entry",
      "id": "itl:781c9a490ca42829656ffbd16b5d1bb1",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2023/07/28/ds-testing",
      "url": "http://brooker.co.za/blog/2023/07/28/ds-testing.html",
      "title": "Invariants: A Better Debugger?",
      "summary_text": "Invariants: A Better Debugger? 🎵Some things never change🎵 Like many of my blog posts, this started out as a long email to a colleague. I expanded it here because I thought folks might find it interesting. I don’t tend to use debuggers. I’m not against them. I’ve seen folks do amazing things with gdb, and envy their skills. I just don’t tend to reach for a debugger very often. I’m also not a huge fan of printf debugging. It can be useful, it’s easy to implement, and works well in both one-box…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2023-07-28T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:1L3TEOm1dj3MMHGk2JqBb23jrlgFQJ5_9V8vkTwDKrYwvp3LtF-st5l9s1E7oJyzVq3ZjBXeD04LLwaYOGRHBw"
    },
    {
      "kind": "entry",
      "id": "itl:1bd19d1b124f9a571fb9718ea5cd6155",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2023/07/13/osdi",
      "url": "http://brooker.co.za/blog/2023/07/13/osdi.html",
      "title": "My Favorite Bits of OSDI/ATC'23",
      "summary_text": "My Favorite Bits of OSDI/ATC’23 Talking to 3D people is cool again. This week brought USENIX ATC’23 and OSDI’23 together in Boston. While I’ve followed OSDI and ATC papers for years, it’s the first time I’ve been to either of them (I’ve have been to NSDI a couple times). It was a really good time. In this post I’ll cover a couple of my favorite papers1, and trends I noticed. Overall, it was great to meet a bunch of folks in person who I’ve only interacted with online, and nice to be back to…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2023-07-13T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:jlWNpQvq-jBU3vOVBZbu-sG93QMefwbIylHP1FmVvG4LOuXy8saawO8N9rKktkBIF6Js4bAnlokHQlJeTMTgDQ"
    },
    {
      "kind": "entry",
      "id": "itl:fb08b3f1982c3c853773f7e4725a537f",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2023/06/23/belady",
      "url": "http://brooker.co.za/blog/2023/06/23/belady.html",
      "title": "Bélády's Anomaly Doesn't Happen Often",
      "summary_text": "Bélády’s Anomaly Doesn’t Happen Often Anomaly is a really fun word. Try saying it ten times. It was 1969. The Summer of Love wasn’t raging4, Hendrix was playing the anthem, and Forest Gump was running rampant. In New York, IBM researchers Bélády, Nelson, and Schedler were hot on the trail of something strange. They had a paging machine, a computer which kept its memory in pages, and sometimes moved those pages to storage. Weird1. It wasn’t only the machine that was weird, it was their…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2023-06-23T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:WLeP_xn7QXjo9SycGEfHA8KeCvDpfNEWLCMLrIZqGNZFolQpXLBi8aLFwbh2WABi4YydcsACTgIEs3nkZDd8DA"
    },
    {
      "kind": "entry",
      "id": "itl:f3793fd8bf80b1edc852920c3b05f879",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2023/06/19/container",
      "url": "http://brooker.co.za/blog/2023/06/19/container.html",
      "title": "What is a container?",
      "summary_text": "What is a container? What are words, anyway? A common cause of confusion and miscommunication I see is different people using different definitions of words. Sometimes the definitions are subtly different (as with availability). Sometimes they’re completely different, and we’re just talking about different things entirely. A common example is the word container, a popular term for a popular technology that means at least four different things. An approach to packaging an application along with…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2023-06-19T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:mFb0aJKiLu_s8YAGuo7DQBEqrjC6rnLrEZYPuLy11aHur8ilVWdLQZfuCt4YvUZfJK3Bm8UUyP5y0ko6bYj-Dw"
    },
    {
      "kind": "entry",
      "id": "itl:3b9eb6069f0a5ab867a19c77fa162962",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2023/05/23/snapshot-loading",
      "url": "http://brooker.co.za/blog/2023/05/23/snapshot-loading.html",
      "title": "Container Loading in AWS Lambda",
      "summary_text": "Container Loading in AWS Lambda Slap shot? Back in 2019, we started thinking about how allow Lambda customers to use container images to deploy their Lambda functions. In theory this is easy enough: a container image is an image of a filesystem, just like the zip files we already supported. The difficulty, as usual with big systems, was performance. Specifically latency. More specifically cold start latency. For eight years cold start latency has been one of our biggest investment areas in…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2023-05-23T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:y1hGPRkFliS-mrcFbtR5dVbzrLnKIHH6gsXig4MJyU25hzjJwsU_A8Et5RQrL6oFjIeCA4zrs66H3xv1URRfDg"
    },
    {
      "kind": "entry",
      "id": "itl:72bf7246b4e0dfd650a64d2e2794da53",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2023/05/10/open-closed",
      "url": "http://brooker.co.za/blog/2023/05/10/open-closed.html",
      "title": "Open and Closed, Omission and Collapse",
      "summary_text": "Open and Closed, Omission and Collapse Were you born in a cave? This, from Open Versus Closed: A Cautionary Tale by Schroeder et al1 is one of the most important concepts in systems performance: Workload generators may be classified as based on a closed system model, where new job arrivals are only triggered by job completions (followed by think time), or an open system model, where new jobs arrive independently of job completions. In general, system designers pay little attention to whether a…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2023-05-10T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:889DArZJl8VWPODEe2qVR0TC5H0bAzLQl4pGi77mInvkdq6lDLfPJ4-TcEIuADwuiWz-ey4M-SrMrwCa-euXAQ"
    },
    {
      "kind": "entry",
      "id": "itl:e03d1a71e67c786cfe499df7939997c5",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2023/04/20/hobbies",
      "url": "http://brooker.co.za/blog/2023/04/20/hobbies.html",
      "title": "The Four Hobbies, and Apparent Expertise",
      "summary_text": "The Four Hobbies, and Apparent Expertise Around the end of high school, I started to get really into photography. My friend (let’s call him T) was also into it, which should have been great fun. But it wasn’t. Going shooting with him was never great, for a reason I didn’t figure out till much later. I wanted to take photos. T mostly enjoyed tinkering with cameras. As I’ve spent more time on different hobbies, it’s become clear that this is a common pattern. Every hobby, pastime1, or sport, is…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2023-04-20T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:kQHtJeCrDVvF5GArxyLrnArhVDYKQ7H802WDj114rOjv7v3By-8h9d3_Xc7QK6LRWI7PtTeW1TQw9kcGxG9PDQ"
    },
    {
      "kind": "entry",
      "id": "itl:f150b2632cf1646a02fd352d59f97a2e",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2023/03/23/economics",
      "url": "http://brooker.co.za/blog/2023/03/23/economics.html",
      "title": "Surprising Scalability of Multitenancy",
      "summary_text": "Surprising Scalability of Multitenancy When most folks talk about the economics of cloud systems, their focus is on automatically scaling for long-term seasonality: changes on the order of days (fewer people buy things at night), weeks (fewer people visit the resort on weekdays), seasons, and holidays. Scaling for this kind of seasonality is useful and important, but there’s another factor that can be even more important and is often overlooked: short-term peak-to-average. Roughly speaking,…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2023-03-23T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:trr1LDa3Kl9upyB5iuZLvLUE5TzlclFFelm6NLvwLq4F-jq4X9v6eVoggpNRrQI-4ZwlDbVjYlTXloDh44W5Bg"
    },
    {
      "kind": "entry",
      "id": "itl:8a4e7271fb98e81116d877d59f5e613c",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2023/03/07/false-sharing",
      "url": "http://brooker.co.za/blog/2023/03/07/false-sharing.html",
      "title": "False Sharing versus Perfect Placement",
      "summary_text": "False Sharing versus Perfect Placement This is part 3 of an informal series on database scalability. The previous parts were on NoSQL, and Hot Keys. In the last installment, we looked at hot keys and how they affect the theoretical peak scale a database can achieve. Hidden in that post was an underlying assumption: that can successfully isolate the hottest key onto a shard of its own. If the key distribution is slow moving (hot keys now will still be hot keys later) then this is achievable.…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2023-03-07T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:s33dDuGrzlAk4r6Eca3oq1XsTm3czcKRxPUfXYrNbr9oWcek_RFLPSsmkFRMe_BnYu6oFlN1pZ79adid1865Dw"
    },
    {
      "kind": "entry",
      "id": "itl:c3d081809d710adbbcc715f1b43595c7",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2023/02/07/hot-keys",
      "url": "http://brooker.co.za/blog/2023/02/07/hot-keys.html",
      "title": "Hot Keys, Scalability, and the Zipf Distribution",
      "summary_text": "Hot Keys, Scalability, and the Zipf Distribution the: so hot right now. Does your distributed database (or microservices architecture, or queue, or whatever) scale? It’s a good question, and often a relevant one, but almost impossible to answer. To make it a meaningful question, you also need to specify the workload and the data in the system. Given this workload, over this data, does this database scale? One common reason systems don’t scale is because of hot keys or hot items: things in the…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2023-02-07T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:lbHfsm8lKAviEQ90vbccfP1Xqxuk_WZaB7_UzU0-ti6cUk4ibTBG257hDBbJCGnB22icgPLxiAb8EUYgkeBmCQ"
    },
    {
      "kind": "entry",
      "id": "itl:92a1d52cb62a77e9bb1b88ae95fd4718",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2023/01/30/nosql",
      "url": "http://brooker.co.za/blog/2023/01/30/nosql.html",
      "title": "NoSQL: The Baby and the Bathwater",
      "summary_text": "NoSQL: The Baby and the Bathwater Is this a database? This is a bit of an introduction to a long series of posts I’ve been writing about what, fundamentally, it is that makes databases scale. The whole series is going to take me a long time, but hopefully there’s something here folks will enjoy. On March 12 2006, Australia set South Africa the massive target of 434 runs to chase in a one-day international at the Wanderers in Johannesburg. South Africa, in reply, set a record that stands to…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2023-01-30T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:ggkw4BZI3GoFe6t_9QTgrjz8NteA0nDlaYIJUIJrkb0F4e1BOjRbSP5zIq_fmnlPi5I768Ws1Leb7vqfrD_fAw"
    },
    {
      "kind": "entry",
      "id": "itl:ddf869bf8208dcbc5023c2220c18361b",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2023/01/06/erasure",
      "url": "http://brooker.co.za/blog/2023/01/06/erasure.html",
      "title": "Erasure Coding versus Tail Latency",
      "summary_text": "Erasure Coding versus Tail Latency There are zero FEC puns in this post, against my better judgement. Jeff Dean and Luiz Barroso’s paper The Tail At Scale popularized an idea they called hedging, simply sending the same request to multiple places and using the first one to return. That can be done immediately, or after some delay: One such approach is to defer sending a secondary request until the first request has been outstanding for more than the 95th-percentile expected latency for this…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2023-01-06T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:CRn9z-H6yGB-gc1x_dh52hhyOBq0z5lIdrxq-jrvou-76GfXyuY__6dTBy6GNYQDLZaydFAy6gUE63H-h1ZdDw"
    },
    {
      "kind": "entry",
      "id": "itl:4e107622bd058c072e2f165549cb153e",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2022/12/15/thumb",
      "url": "http://brooker.co.za/blog/2022/12/15/thumb.html",
      "title": "Under My Thumb: Insight Behind the Rules",
      "summary_text": "Under My Thumb: Insight Behind the Rules My left thumb is exactly 25.4mm wide. Starting off in a new field, you hear a lot of rules of thumb. Rules for estimating things, thinking about things, and (ideally) simplifying tough decisions. When I started in Radar, I heard: the transmitter makes up three quarters of the cost of a radar system and when I started building computer systems, I heard a lot of things like: hardware is free, developers are expensive and, the ubiquitous: premature…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2022-12-15T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:uJW4Bj596tIrNMTpburz5sbd2gUmoJtymLwluY6uVbmwYiLExjFz2FmDf5nSXcH1q5v1X2jrghfCGwbneujqAw"
    },
    {
      "kind": "entry",
      "id": "itl:7e115c69ec2ae84ee75949ace60a5ea5",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2022/11/29/snapstart",
      "url": "http://brooker.co.za/blog/2022/11/29/snapstart.html",
      "title": "Lambda Snapstart, and snapshots as a tool for system builders",
      "summary_text": "Lambda Snapstart, and snapshots as a tool for system builders Clones. Yesterday, AWS announced Lambda Snapstart, which uses VM snapshots to reduce cold start times for Lambda functions that need to do a lot of work on start (starting up a language runtime3, loading classes, running static code, initializing caches, etc). Here’s a short 1 minute video about it: Or, for a lot more context on Lambda and how we got here5: Snapstart is a super useful capability for Lambda customers. I’m extremely…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2022-11-29T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:v8euStbrAR5eJl1ce7-fP5lqFcwialP86bL64V_WCee91pdfsYQVo96J16ua8xRTK8gbNz-Ttrj4HFsapze5Bg"
    },
    {
      "kind": "entry",
      "id": "itl:cfa24e56af535da23fe302e88f9d5e34",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2022/11/22/manifesto",
      "url": "http://brooker.co.za/blog/2022/11/22/manifesto.html",
      "title": "Amazon's Distributed Computing Manifesto",
      "summary_text": "Amazon’s Distributed Computing Manifesto Manifesto made manifest. In the Johannesburg of 1998, I was rocking a middle parting, my friend group was abuzz about the news that there was water (and therefore monsters) on Europa, and all the cool kids were getting satellite TV at home1. Over in Seattle, the folks at Amazon.com had started to notice that their architecture was in need of rethinking. $147 million in sales in 1997, and over $600 million in 1998, were proving to be challenging to deal…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2022-11-22T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:TiBPtIAXPkDYHrE4OdvMm6KTw_SNYkDA_ZyLzqxI_H3dHprBwODY2jCYUKy47oKzuaYldwRVidPqIDyLpZIrBg"
    },
    {
      "kind": "entry",
      "id": "itl:1e3aeba2a4a513a8d7efe7e437e2db82",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2022/11/08/writing",
      "url": "http://brooker.co.za/blog/2022/11/08/writing.html",
      "title": "Writing Is Magic",
      "summary_text": "Writing Is Magic Magic can be dangerous. Sometimes when folks ask me for advice at work, I write them very long emails to answer their question. Sometimes, those emails are generally interesting and not work-specific, so I share them here. A couple days ago somebody asked me about how to get better at communicating their ideas and opinions, how to extend their influence, and how to drive consensus. This was my reply. There are many ways to be influential. You can form 1:1 relationships with…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2022-11-08T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:PiC2VTmAgvSD66JwnT8Jmu_Fl6npY5XGmqbSDGsg1OXKRVm8F1NW28YwN3DA6-fzqCnwPF2sEAxWggAHn93DDQ"
    },
    {
      "kind": "entry",
      "id": "itl:c489cd24ef030257477a6596b0e867be",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2022/10/21/nudge",
      "url": "http://brooker.co.za/blog/2022/10/21/nudge.html",
      "title": "Give Your Tail a Nudge",
      "summary_text": "Give Your Tail a Nudge Tricks are fun. We all care about tail latency (also called high percentile latency, also called those times when your system is weirdly slow). Simple changes that can bring it down are valuable, especially if they don’t come with difficult tradeoffs. Nudge: Stochastically Improving upon FCFS presents one such trick. The Nudge paper interests itself in tail latency compared to First Come First Served (FCFS)1, for a good reason: While advanced scheduling algorithms are a…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2022-10-21T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:0SHK3p-cOH14sXopOA6El3ALPCje9KSfXLsvmaGinhqBQT84vxHqd7Um6P8hM1P_tlNzPU4XwGCIEMa6tvp1CA"
    },
    {
      "kind": "entry",
      "id": "itl:bebb986eef8923d66dd055796e8a960d",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2022/10/04/commitment",
      "url": "http://brooker.co.za/blog/2022/10/04/commitment.html",
      "title": "Atomic Commitment: The Unscalability Protocol",
      "summary_text": "Atomic Commitment: The Unscalability Protocol 2PC is my enemy. Let’s consider a single database system, running on one box, good for 500 requests per second. ┌───────────────────┐ │ Database │ │(good for 500 rps) │ └───────────────────┘ What if we want to access that data more often than 500 times a second? If by access we mean read, we have a lot of options. If be access, we mean write or even perform arbitrary transactions on, we’re in a trickier situation. Tricky problems aside, we forge…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2022-10-04T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:_nQHW7wPvR1wsfFxibh4tj5V8NetaM0UEJ6foQ1mknu_Fe5xCbpCridN7LkJcOQ8E9Pzw6YxQ3a3Zf0NmljzBw"
    },
    {
      "kind": "entry",
      "id": "itl:23e71ac7a0d6f4fa58a53c45b36671b2",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2022/09/02/ecdf",
      "url": "http://brooker.co.za/blog/2022/09/02/ecdf.html",
      "title": "Histogram vs eCDF",
      "summary_text": "Histogram vs eCDF Accumulation is a fun word. Histograms are a rightfully popular way to present data like latency, throughput, object size, and so on. Histograms avoid some of the difficulties of picking a summary statistic, or group of statistics, which is hard to do right. I think, though, that there’s nearly always a better choice than histograms: the empirical cumulative distribution function (eCDF). To understand why, let’s look at an example, starting with the histogram1. This latency…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2022-09-02T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:Qcam8W-wHrjWtZNc54XsKen_EB5xe9SmTZHCnsjAfNW9hR1TUgjV7AR_BHBKKwvHrMk-947ACjuPFE1fCTE4Aw"
    },
    {
      "kind": "entry",
      "id": "itl:247d7a4815c7b08b1010d722e300041f",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2022/08/11/backoff",
      "url": "http://brooker.co.za/blog/2022/08/11/backoff.html",
      "title": "What is Backoff For?",
      "summary_text": "What is Backoff For? Back off man, I'm a scientist. Years ago I wrote a blog post about exponential backoff and jitter, which has turned out to be enduringly popular. I like to believe that it’s influenced at least a couple of systems to add jitter, and become more stable. However, I do feel a little guilty about pushing the popularity of jitter without clearly explaining what backoff and jitter do, and do not do. Here’s the pithy statement about backoff: Backoff helps in the short term. It is…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2022-08-11T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:yULLHqPRBqPEIyHPkqJetbmbA0v-uG0SyoYF5Azp4kAp6Wtn_2D_tZZVHJL3UyIJWpEiCqepUn0rbujC7k3aBw"
    },
    {
      "kind": "entry",
      "id": "itl:be0c97f78490e3032a181b646f077161",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2022/07/29/getting-into-tla",
      "url": "http://brooker.co.za/blog/2022/07/29/getting-into-tla.html",
      "title": "Getting into formal specification, and getting my team into it too",
      "summary_text": "Getting into formal specification, and getting my team into it too Getting started is the hard part Sometimes I write long email replies to people at work asking me questions. Sometimes those emails seem like they could be useful to more than just the recipient. This is one of those emails: a reply to a software engineer asking me how they could adopt formal specification in their team, and how I got into it. Sometime around 2011 I was working on some major changes to the EBS control plane. We…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2022-07-29T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:OD-eNOUNBpZw4UiLHyChGwFtPGGhJ48s2Ej30YkVvJCYhluxWs1QSPWdh8gADTGaVPTauuLRdZwvH38bnnKwBQ"
    },
    {
      "kind": "entry",
      "id": "itl:e65305dc0a22fa182c5e33b998909dab",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2022/07/12/dynamodb",
      "url": "http://brooker.co.za/blog/2022/07/12/dynamodb.html",
      "title": "The DynamoDB paper",
      "summary_text": "The DynamoDB paper The other database called Dynamo This week at USENIX ATC’22, a group of my colleagues1 from the AWS DynamoDB team are going to be presenting their paper Amazon DynamoDB: A Scalable, Predictably Performant, and Fully Managed NoSQL Database Service. This paper is a rare look at a real-world distributed system that runs at massive scale. From the paper: In 2021, during the 66-hour Amazon Prime Day shopping event, Amazon systems … made trillions of API calls to DynamoDB, peaking…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2022-07-12T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:I0Bn34-vwMDkU2W-v8NTiGX9W-Jcf77ba58qgOjdeRm4Z2gd7FK5BJ8nZyTC3P-RoivDUmO5ESsmH8SFhf8dDA"
    },
    {
      "kind": "entry",
      "id": "itl:4812bdbd85a3a9fd506ba12642cb1ebd",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2022/06/02/formal",
      "url": "http://brooker.co.za/blog/2022/06/02/formal.html",
      "title": "Formal Methods Only Solve Half My Problems",
      "summary_text": "Formal Methods Only Solve Half My Problems At most half my problems. I have a lot of problems. The following is a one-page summary I wrote as a submission to HPTS’22. Hopefully it’s of broader interest. Formal methods, like TLA+ and P, have proven to be extremely valuable to the builders of large scale distributed systems1, and to researchers working on distributed protocols. In industry, these tools typically aren’t used for full verification. Instead, effort is focused on interactions and…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2022-06-02T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:rbAzNxj8SWqlPfUoBe8wloLZ91Ui97Qppl3rOGc6fIscKqbuPnYb-Eoloea-DkH1HaWtvujgYhnSqDAnKsXiAg"
    },
    {
      "kind": "entry",
      "id": "itl:3cfb303f1945d9f8149e9f1f6b4b47f6",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2022/05/03/simplicity",
      "url": "http://brooker.co.za/blog/2022/05/03/simplicity.html",
      "title": "What is a simple system?",
      "summary_text": "What is a simple system? Is this pretentious? Why do I need cryptography when I could simply hide the contents of my communications rotating every letter by 13? Why do I need a distributed storage system when I could simply store my files on this one server? Why do I need a database when I could simply use a flat file? Do any of those things, and feel joy in a job well done. A simple solution. Perhaps you’re hiding your communications from a child, storing little data with low value, and…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2022-05-03T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:yFz6SGgRn0BDaIFxICjs-OpJKacABXJBj0MXf3SZj95ruf0a11smIccFiAUwIAMRRAPNJcBDxmhB8DYvpD76Cw"
    },
    {
      "kind": "entry",
      "id": "itl:1b066705a44b2833028308779360ec74",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2022/04/11/simulation",
      "url": "http://brooker.co.za/blog/2022/04/11/simulation.html",
      "title": "Simple Simulations for System Builders",
      "summary_text": "Simple Simulations for System Builders Even the most basic numerical methods can lead to surprising insights. It’s no secret that I’m a big fan of formal methods. I use P and TLA+ often. I like these tools because they provide clear ways to communicate about even the trickiest protocols, and allow us to use computers to reason about the systems we’re designing before we build them1. These tools are typically focused on safety (Nothing bad happens) and liveness (Something good happens…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2022-04-11T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:omF6St-iqKG4mRI_yZQeunzyrGYjMkkSbymDVaX0rJXNIwu7NWg2zUkz9C7bY7mZAghgpla-tTq17TEpSXL_CA"
    },
    {
      "kind": "entry",
      "id": "itl:411b0b330202ab038468e61938ae6396",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2022/02/28/retries",
      "url": "http://brooker.co.za/blog/2022/02/28/retries.html",
      "title": "Fixing retries with token buckets and circuit breakers",
      "summary_text": "Fixing retries with token buckets and circuit breakers Throttle yourself before you DoS yourself. After my last post on circuit breakers, a couple of people reached out to recommend using circuit breakers only to break retries, and still send normal first try traffic no matter the failure rate. That’s a nice approach. It provides possible solutions to the core problem with client-side circuit breakers (they may make partial outages worse), and to the retry problem (where retries increase load…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2022-02-28T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:T73rV3MU9mvYefn_QwR2M7mIpOHbE_t-XjuKqPwTafkHnq0Gzl80hDHm5DBEZ84hy0iThyHRTW82-WB1pBU1CA"
    },
    {
      "kind": "entry",
      "id": "itl:fecc02ca29af49c1015545b4c7c34acc",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2022/02/16/circuit-breakers",
      "url": "http://brooker.co.za/blog/2022/02/16/circuit-breakers.html",
      "title": "Will circuit breakers solve my problems?",
      "summary_text": "Will circuit breakers solve my problems? Maybe, but you need to know what problem you're trying to solve first. A couple of weeks ago, I started a tiny storm on Twitter by posting this image, and claiming that retries (mostly) make things worse in real-world distributed systems. The bottom line is that retries are often triggered by overload conditions, permanent or transient, and tend to make those conditions worse by increasing traffic. Many people replied saying that I’m ignoring the…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2022-02-16T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:k1VPyy_EvPrVfgh4zxeNfJTHOEda7FbnH0UxnrHA5I6fFFZ2okDC7xmIKAcrO6yp2CLR69FDP8oMSHhKXCIlDg"
    },
    {
      "kind": "entry",
      "id": "itl:634c572ba999eedbd48614152b376c98",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2022/01/31/deployments",
      "url": "http://brooker.co.za/blog/2022/01/31/deployments.html",
      "title": "Software Deployment, Speed, and Safety",
      "summary_text": "Software Deployment, Speed, and Safety There's one right answer that applies in all situations, as always. Disclaimer: Sometime around a 2015, I wrote AWS’s official internal guidance on balancing deployment speed and safety. This blog post is not that. It’s not official guidance from AWS (nothing on this blog is), and certainly not guidance for AWS. Instead, it’s my own take on deployments and safety, and how I think about the space. You’ll find a lot of opinions about deployments on the…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2022-01-31T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:ZB_ggY7PfxbrtpWRXl74GNBxevAFOZx966-mdNSaInoCjw5P0T4HgnXd8oEV0IzanmMpSVN3fbLgUPCqg96ODg"
    },
    {
      "kind": "entry",
      "id": "itl:b8e2242bdab4a8328e91b97607181205",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2022/01/19/predictability",
      "url": "http://brooker.co.za/blog/2022/01/19/predictability.html",
      "title": "DynamoDB's Best Feature: Predictability",
      "summary_text": "DynamoDB’s Best Feature: Predictability Happy birthday! It’s 10 years since the launch of DynamoDB, Amazon’s fast, scalable, NoSQL database. Back when DynamoDB launched, I was leading the team rethinking the control plane of EBS. At the time, we had a large number of manually-administered MySQL replication trees, which were giving us a lot of operational pain. Writes went to a single primary, and reads came from replicas, with lots of eventual consistency and weird anomalies in the mix. Our…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2022-01-19T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:QzHu_b0tHc1ltfxbtf2_EKro_FmBno_Vl5YgU1gYr6IY9B50uT22BrlGk1fZ2-7SdH520Azvg6errkROV2XBCw"
    },
    {
      "kind": "entry",
      "id": "itl:150fff88cf13ad75be67ed58538c5130",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2021/11/16/paxos",
      "url": "http://brooker.co.za/blog/2021/11/16/paxos.html",
      "title": "The Bug in Paxos Made Simple",
      "summary_text": "The Bug in Paxos Made Simple There's not really a bug in Paxos, but clickbait is fun. Over the last few weeks, I’ve been picking up the excellent P programming language, a language for modelling and specifying distributed systems. One of the first things I did in P was implement Paxos - an algorithm I know well, has a lot of subtle failure modes, and is easy to get wrong. Perfect for practicing specification. To test out P’s model checker, I intentionally implemented a subtly buggy version of…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2021-11-16T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:gY2DsPzFu4Jni70ACCtNQUnqIvMQFVDWPEiAyZaR8Ht3CGG_w0OUH7vpJ41lIC9EXA2KbCEwT7r2J2lPD2qNAw"
    },
    {
      "kind": "entry",
      "id": "itl:2449950301a507566e71a479e3202c18",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2021/10/20/simulation",
      "url": "http://brooker.co.za/blog/2021/10/20/simulation.html",
      "title": "Serial, Parallel, and Quorum Latencies",
      "summary_text": "Serial, Parallel, and Quorum Latencies Why are they letting me write Javascript? I’ve written before about the latency effects of series (do X, then Y), parallel (do X and Y, wait for them both), and quorum (do X, Y and Z, return when two of them are done) systems. The effects of these different approaches to doing multiple things are quite intuitive. What may not be intuitive, though, is the impact of quorums, and how much quorums can reduce tail latency. So I put together this little toy…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2021-10-20T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:3VNybwaFpDjdlkw1vLPpjjUa_RxQr3Onn7JCRrGS3Li7Cmz07f8EsBOnAcQbM5NUHBvClFA50mZslZ7SkJz-Ag"
    },
    {
      "kind": "entry",
      "id": "itl:0277a151a8c8a09bc667903b105ee2a1",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2021/08/27/caches",
      "url": "http://brooker.co.za/blog/2021/08/27/caches.html",
      "title": "Caches, Modes, and Unstable Systems",
      "summary_text": "Caches, Modes, and Unstable Systems Best practices are seldom the best. Is your system having scaling trouble? A bit too slow? Sending too much traffic to the database? Add a caching layer! After all, caches are a best practice and a standard way to build systems. What trouble could following a best practice cause? Lots of trouble, as it turns out. In the context of distributed systems, caches are a powerful and useful tool. Unfortunately, applied incorrectly, caching can introduce some highly…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2021-08-27T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:eSyrb5SlMJCtwZ3v0VS1gYy7qOhy0meuKb7jYwFcsTyPLvWa1M9Nhi2kCnU9kbPBpkOMER0zkacC7BZNOygVDQ"
    },
    {
      "kind": "entry",
      "id": "itl:a8d7430bf3c14e288c38f127c6a9d8c1",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2021/08/11/arecibo",
      "url": "http://brooker.co.za/blog/2021/08/11/arecibo.html",
      "title": "My Proposal for Arecibo: Drones",
      "summary_text": "My Proposal for Arecibo: Drones With apologies to real radio astronomers Last night I finally got around to watching Grady Hillhouse’s excellent video on the collapse of the Arecibo Telescope. At the end of Grady’s video he says: I hope eventually that they can replace the telescope with an instrument as futuristic and forward-looking as the Arecibo Telescope when first conceived. I hope so too. While I’ve never worked in radio astronomy, my PhD supervisor and a number of my colleagues were…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2021-08-11T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:adPXj6aUHpCgdoSxY1wVzX7rSq1JZjhCTXlUSGMRkiEbfeRJtvDW3OyKnvCnD3EfOWwsrorzEFG3eCw8sUgKCg"
    },
    {
      "kind": "entry",
      "id": "itl:08dd91bfe4d02eb7cc0edd01ef0a030e",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2021/08/05/utilization",
      "url": "http://brooker.co.za/blog/2021/08/05/utilization.html",
      "title": "Latency Sneaks Up On You",
      "summary_text": "Latency Sneaks Up On You And is a bad way to measure efficiency. As systems get big, people very reasonably start investing more in increasing efficiency and decreasing costs. That’s a good thing, for the business, for the environment, and often for the customer. Much of the time efficient systems have lower and more predictable latencies, and everybody enjoys lower and more predictable latencies. Most folks around me think about latency using percentiles or other order statistics. Common…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2021-08-05T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:ir9XNGR2uy1yfgqoVvlz4LD4df3_WaCjlAP-L_GV6NnPrOZxoN8hHPQbjV7vTH3GsEwMT6OAiMqcsk0Ro-HYDA"
    },
    {
      "kind": "entry",
      "id": "itl:d8ff9792ee218fd59fbaab72e908ba56",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2021/05/24/metastable",
      "url": "http://brooker.co.za/blog/2021/05/24/metastable.html",
      "title": "Metastability and Distributed Systems",
      "summary_text": "Metastability and Distributed Systems What if computer science had different parents? There’s no more time-honored way to get things working again, from toasters to global-scale distributed systems, than turning them off and on again. The reasons that works so well are varied, but one reason is especially important for the developers and operators of distributed systems: metastability. I’ll let the authors of Metastable Failures in Distributed Systems define what that means: Metastable…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2021-05-24T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:SKFbg5XBKB1ZbPFQCN36eOtwJXFRNaybwdZcsvkSj1zagVO6ic-hm5sQjgSod6ankxfl8D1vrbNDFSjfFq-HCQ"
    },
    {
      "kind": "entry",
      "id": "itl:4e7bf4adf86676ac1d9351e16ca07e6c",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2021/04/19/latency",
      "url": "http://brooker.co.za/blog/2021/04/19/latency.html",
      "title": "Tail Latency Might Matter More Than You Think",
      "summary_text": "Tail Latency Might Matter More Than You Think A frustratingly qualitative approach. Tail latency, also known as high-percentile latency, refers to high latencies that clients see fairly infrequently. Things like: “my service mostly responds in around 10ms, but sometimes takes around 100ms”. There are many causes of tail latency in the world, including contention, garbage collection, packet loss, host failure, and weird stuff operating systems do in the background. It’s tempting to look at the…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2021-04-19T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:JEnD2jUeLLRDiqiCoiBKXy27bFMGSDUnE3gELrbHEz0wjf33CLaXJjiBAXZ2MLt4PEltR7wxSTvmREDdEUleCQ"
    },
    {
      "kind": "entry",
      "id": "itl:2669852aceba3253955268529572e476",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2021/04/14/redundancy",
      "url": "http://brooker.co.za/blog/2021/04/14/redundancy.html",
      "title": "Redundant against what?",
      "summary_text": "Redundant against what? Threat modeling thinking to distributed systems. There’s basically one fundamental reason that distributed systems can achieve better availability than single-box systems: redundancy. The software, state, and other things needed to run a system are present in multiple places. When one of those places fails, the others can take over. This applies to replicated databases, load-balanced stateless systems, serverless systems, and almost all other common distributed…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2021-04-14T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:l3EYFjeAikf7I7IYZx757QrreQIRQDQPanjXOcxVRRWKTuzDuGgFIWoLMq6cbTFPPQzrvXLRWeMXrgCcYyjkCA"
    },
    {
      "kind": "entry",
      "id": "itl:39b33f7d7d5bbac6b3e28c9f3a9f925d",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2021/03/25/latency-bandwidth",
      "url": "http://brooker.co.za/blog/2021/03/25/latency-bandwidth.html",
      "title": "What You Can Learn From Old Hard Drive Adverts",
      "summary_text": "What You Can Learn From Old Hard Drive Adverts The single most important trend in systems. Adverts for old computer hardware, especially hard drives, are a fun staple of computer forums and the nerdier side of the internet1. For example, a couple days ago, Glenn Lockwood tweeted out this old ad: At least this isn’t an ad for a HAMR drive. $10k in today’s dollars. pic.twitter.com/2h2g3Gnguw — Glenn K. Lockwood (@glennklockwood) March 24, 2021 Apparently from the early ’80s, these drives offered…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2021-03-25T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:HVZVX7XdBCVf2vjVcp3UKVnwRFNjWXKw5WlOaDeTMWq8I2AeKarHEGoMj_NzMVLG9t7wREHZPv8JeUOM4QsmDg"
    },
    {
      "kind": "entry",
      "id": "itl:8174d20323bc2082af709b1429ff2c89",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2021/02/22/postmortem",
      "url": "http://brooker.co.za/blog/2021/02/22/postmortem.html",
      "title": "Incident Response Isn't Enough",
      "summary_text": "Incident Response Isn’t Enough Single points of failure become invisible. Postmortems, COEs, incident reports. Whatever your organization calls them, when done right they are a popular and effective way of formalizing the process of digging into system failures, and driving change. The success of this approach has lead some to believe that postmortems are the best, or even only, way to improve the long-term availability of systems. Unfortunately, that isn’t true. A good availability program…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2021-02-22T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:etZyvzqHVuE4w8OOlHSx3BDUoJLN1rQ8RQ2OZoYaRjST6uZ_rJI1RW-S6vj2m_d2mH-1mfBezBMYkTMblweAAQ"
    },
    {
      "kind": "entry",
      "id": "itl:972204b457417e32dde1e84d9291164c",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2021/01/22/cloud-scale",
      "url": "http://brooker.co.za/blog/2021/01/22/cloud-scale.html",
      "title": "The Fundamental Mechanism of Scaling",
      "summary_text": "The Fundamental Mechanism of Scaling It's not Paxos, unfortunately. A common misconception among people picking up distributed systems is that replication and consensus protocols—Paxos, Raft, and friends—are the tools used to build the largest and most scalable systems. It’s obviously true that these protocols are important building blocks. They’re used to build systems that offer more availability, better durability, and stronger integrity than a single machine. At the most basic level,…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2021-01-22T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:z-6Lm3UxsAZMEr-HwP5R_YRTeQx3b0PjTPgjr7yT401l0CfWrsrylMIkFlcwDPx76FlTGQ1MoKKsmIJ547fhAA"
    },
    {
      "kind": "entry",
      "id": "itl:2fc1bc6d1cc640ec22fd877d7dd285f9",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2021/01/06/quorum-availability",
      "url": "http://brooker.co.za/blog/2021/01/06/quorum-availability.html",
      "title": "Quorum Availability",
      "summary_text": "Quorum Availability It's counterintuitive, but is it right? In our paper Millions of Tiny Databases, we say this about the availability of quorum systems of various sizes: As illustrated in Figure 4, smaller cells offer lower availability in the face of small numbers of uncorrelated node failures, but better availability when the proportion of node failure exceeds 50%. While such high failure rates are rare, they do happen in practice, and a key design concern for Physalia. And this is what…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2021-01-06T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:FRbWRj7b44q0--V8wmQrzNr8zdq_CG59Ft9KyH--eO693qNti2rXi0ktyKyZFucV7oSA6Qr5C4uyHOVh3-7RBg"
    },
    {
      "kind": "entry",
      "id": "itl:a3bd89e8ff6c66d02fcf343329250d96",
      "source_id": "src-brooker",
      "origin_feed": "https://brooker.co.za/blog/rss.xml",
      "origin_id": "http://brooker.co.za/blog/2020/10/19/big-changes",
      "url": "http://brooker.co.za/blog/2020/10/19/big-changes.html",
      "title": "Getting Big Things Done",
      "summary_text": "Getting Big Things Done In one particular context. A while back, a colleague wanted to make a major change in the design of a system, the sort of change that was going to take a year or more, and many tens of person-years of effort. They asked me how to justify the project. This post is part of the email reply I sent. The advice is in context of technical leadership work at a big company, but perhaps it may apply elsewhere. Is it the right solution? I like to pay attention to ways I can easily…",
      "topics": [
        "distributed-systems"
      ],
      "published": "2020-10-19T00:00:00Z",
      "observed_at": "2026-08-22T14:18:01Z",
      "sig": "ed25519:k1:RfTmId-EExOSxzAtfi5eKEaGDxEQCSSz1_wfwHyg_Ioxj-n_cNDSGG78Ej22IaJghXvtHpMfiUF8Te2BO2SfDA"
    }
  ]
}
