The Crash Window: Why I Picked Double-Counting Over Data Loss

A service I worked on processed inventory messages. Millions per minute. Each message carried a list of SKU changes. The service filtered, accumulated counts, and wrote them to a distributed cache. The message queue guaranteed at-least-once delivery. That guarantee is the problem. At-least-once means the queue redelivers a message if it thinks you did not process it. Crash during processing, ack timeout, partition rebalance. All trigger redelivery. Duplicates are not exceptional. They are the baseline. ...

2026-09-26 · 5 min · 1022 words