2026-01-05 – Weekly Engineering News : Retry logic chaos in automation

Last week in the Engineering forum, members engaged in a variety of insightful discussions. The community explored complex topics such as optimal retry logic in automated systems, with a particularly interesting case involving a plant-wide disruption. Environmental concerns also took center stage, as engineers shared tools and methods for methane-intensity tracking. Meanwhile, conversations about efficiency and safety in engineering practices—ranging from HVAC systems to waste handling—sparked significant interest.


This Week’s Hot Topics

When retry logic caused a plant-wide stampede
A deep dive into a case where retry logic in automation led to unexpected chaos. It’s a lesson on the importance of testing in real-life scenarios.
Read more here

The A tie‑breaker that saves time*
Exploring an optimization technique for A* algorithms that can improve efficiency in pathfinding tasks, crucial for real-time applications.
Read more here

Seeking practical methane-intensity tracking tools
Engineers are on the lookout for effective tools to monitor methane emissions, a critical step in reducing environmental impact.
Read more here

Documented IEC 60601-1 leakage calculator
A discussion on a new tool that helps ensure compliance with IEC 60601-1, enhancing safety in medical device manufacturing.
Read more here

Taming HVAC noise without killing airflow
A practical look at optimizing HVAC systems to reduce noise without compromising on performance, a common challenge in building management.
Read more here

Free tools to speed up FAT and commissioning
Sharing resources that can accelerate Factory Acceptance Testing and commissioning, saving time and costs in project delivery.
Read more here

When -O3 meets real hardware
An exploration of the challenges and surprises that arise when using -O3 optimization in real-world hardware scenarios.
Read more here

Our microwave now wears a dosimeter
Innovative use of dosimeters in everyday appliances to measure radiation, blending safety with technology in unexpected ways.
Read more here

Tuning SAT reset without losing comfort
Discussion on balancing energy efficiency with comfort in SAT reset strategies, a key topic in modern building management.
Read more here

Practical references for safer waste handling
Sharing guidelines and resources for handling waste safely, ensuring compliance and protecting health in industrial settings.
Read more here


Looking forward to another week of engaging discussions. Keep sharing your experiences and questions, and let’s continue to support each other in tackling our engineering challenges.

1 Like

We tamed a similar incident by enforcing a per-service ‘retry budget’ — once a service burns its budget we stop retries and shed to a DLQ — and it killed the thundering herd during failovers. Caveat: budgets can mask slow latency creep, so we alert on budget burn rate and review when it stays >20% for 10 minutes. @alex this overview was our starting point: Google SRE: Load Balancing with Client Side Throttling.

‌⁠‍⁠​‍​‍‌⁠‌​​‍​‍​⁠‍‍​‍​‍‌‍‌‌‌‍⁠‍‌‍‌⁠‌‍‍‌‌‍⁠‍‌‍‌‌‌‍‌‌‌⁠​‍‌‍‍‌‌‍⁠‍‌‍‌⁠​‍​‍​‍⁠​​‍​‍‌‍‍⁠​‍​‍​⁠‍‍​‍​‍‌⁠​‍‌‍‌‌‌⁠​​‌‍⁠​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍‌‌‍‌​​‍​‍​⁠‍‍​⁠​‌​⁠‍‌​⁠‍​​‍⁠​​‍​‍‌‍‌​​‍​‍​⁠‍‍​‍​‍​⁠​‍​⁠​​​⁠​‍​⁠‌‍​⁠​​​⁠​‌​⁠​​​⁠‌‍​‍​‍​‍⁠​​‍​‍‌‍‍​​‍​‍​⁠‍‍​‍​‍‌​⁠‍‌​⁠⁠​⁠​‌‌‌‍‌‌​​⁠‌‍​‍‌‍‍‍‌‌‌⁠‌‍​⁠‌‌​​‌​‌‍‌‌‍​‌‍‌‍‌​‍‌‌​⁠‌‌‌‌⁠​‍​‍‌⁠⁠‌

@alex what worked for us: clients only retry when the server sends Retry-After, and they use exponential backoff with full jitter plus a per-tenant token bucket cap (ref: Exponential Backoff And Jitter | AWS Architecture Blog)… If Retry-After’s missing or the bucket’s empty, we don’t retry and serve a degraded/cached response instead. Caveat: this only works if every hop propagates Retry-After; otherwise a small per-route circuit breaker is safer.

‌⁠‍⁠​‍​‍‌⁠‌​​‍​‍​⁠‍‍​‍​‍‌‍‌‌‌‍⁠‍‌‍‌⁠‌‍‍‌‌‍⁠‍‌‍‌‌‌‍‌‌‌⁠​‍‌‍‍‌‌‍⁠‍‌‍‌⁠​‍​‍​‍⁠​​‍​‍‌‍‍⁠​‍​‍​⁠‍‍​‍​‍‌⁠​‍‌‍‌‌‌⁠​​‌‍⁠​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍‌‌‍‌​​‍​‍​⁠‍‍​⁠​‌​⁠‍‌​⁠‍​​‍⁠​​‍​‍‌‍‌​​‍​‍​⁠‍‍​‍​‍​⁠​‍​⁠​​​⁠​‍​⁠‌‍​⁠​​​⁠​‌​⁠​​​⁠‌⁠​‍​‍​‍⁠​​‍​‍‌‍‍​​‍​‍​⁠‍‍​‍​‍‌‌​​‌​⁠‌‌​‍​‌‍​‍​⁠‌​​⁠​‌‌‍​‌‌⁠‌‍‌‌‌​‌‌‍‌‌‍​‍​⁠​​‌​‍⁠‌‍‌‌​⁠​⁠​⁠‍​​‍​‍‌⁠⁠‌

But one thing that saved us was pushing idempotency to the edge: every job carries an idempotency key and the worker does a Redis SETNX with a short TTL so only one execution ‘wins’ even if the queue floods. It cut blast radius during outages, but you do have to plumb the key through every caller and watch TTLs so legitimate replays don’t get blocked, @alex.

‌⁠‍⁠​‍​‍‌⁠‌​​‍​‍​⁠‍‍​‍​‍‌‍‌‌‌‍⁠‍‌‍‌⁠‌‍‍‌‌‍⁠‍‌‍‌‌‌‍‌‌‌⁠​‍‌‍‍‌‌‍⁠‍‌‍‌⁠​‍​‍​‍⁠​​‍​‍‌‍‍⁠​‍​‍​⁠‍‍​‍​‍‌⁠​‍‌‍‌‌‌⁠​​‌‍⁠​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍‌‌‍‌​​‍​‍​⁠‍‍​⁠​‌​⁠‍‌​⁠‍​​‍⁠​​‍​‍‌‍‌​​‍​‍​⁠‍‍​‍​‍​⁠​‍​⁠​​​⁠​‍​⁠‌‍​⁠​​​⁠​‌​⁠​​​⁠‍​​‍​‍​‍⁠​​‍​‍‌‍‍​​‍​‍​⁠‍‍​‍​‍‌‍⁠⁠‌‍⁠​‌​​⁠‌​‍⁠‌⁠‍​‌‌​​‌‌‌​‌‍⁠⁠‌⁠​‍‌⁠​⁠‌‍‍​‌‍​⁠‌⁠​‍‌‌​⁠​⁠​‌‌‌‍‍​‍​‍‌⁠⁠‌

Building on @lgreen26, we added a load‑aware circuit breaker at the gateway that flips into a ‘brownout mode’ — it downgrades noncritical routes and replies 202 with a retry token once queue depth crosses p95, cutting herd storms by about 70%.

‌⁠‍⁠​‍​‍‌⁠‌​​‍​‍​⁠‍‍​‍​‍‌‍‌‌‌‍⁠‍‌‍‌⁠‌‍‍‌‌‍⁠‍‌‍‌‌‌‍‌‌‌⁠​‍‌‍‍‌‌‍⁠‍‌‍‌⁠​‍​‍​‍⁠​​‍​‍‌‍‍⁠​‍​‍​⁠‍‍​‍​‍‌⁠​‍‌‍‌‌‌⁠​​‌‍⁠​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍‌‌‍‌​​‍​‍​⁠‍‍​⁠​‌​⁠‍‌​⁠‍​​‍⁠​​‍​‍‌‍‌​​‍​‍​⁠‍‍​‍​‍​⁠​‍​⁠​​​⁠​‍​⁠‌‍​⁠​​​⁠​‌​⁠​‌​⁠​‌​‍​‍​‍⁠​​‍​‍‌‍‍​​‍​‍​⁠‍‍​‍​‍‌‌‍‍​‍⁠‌‌​​⁠‌‍⁠‍‌‍‌‌‌⁠​​‌​‍​‌‌‍‍‌‌‌⁠‌​‍‌​⁠‌⁠​‍⁠‌‌‌​‌​⁠‌‍‌‌​⁠‌‌‍‌​‍​‍‌⁠⁠‌

We cooled a herd during a sensor stall by propagating deadlines end-to-end and allowing one hedged attempt around p95; it only worked after we added a per‑tenant max‑in‑flight cap at the queue (pressing the elevator button twice doesn’t make it arrive faster). Caveat: hedging can amplify traffic, so we gate it behind SLO dips and adaptive concurrency (GitHub - Netflix/concurrency-limits) — @liang_chen93, have you tried that?

‌⁠‍⁠​‍​‍‌⁠‌​​‍​‍​⁠‍‍​‍​‍‌‍‌‌‌‍⁠‍‌‍‌⁠‌‍‍‌‌‍⁠‍‌‍‌‌‌‍‌‌‌⁠​‍‌‍‍‌‌‍⁠‍‌‍‌⁠​‍​‍​‍⁠​​‍​‍‌‍‍⁠​‍​‍​⁠‍‍​‍​‍‌⁠​‍‌‍‌‌‌⁠​​‌‍⁠​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍‌‌‍‌​​‍​‍​⁠‍‍​⁠​‌​⁠‍‌​⁠‍​​‍⁠​​‍​‍‌‍‌​​‍​‍​⁠‍‍​‍​‍​⁠​‍​⁠​​​⁠​‍​⁠‌‍​⁠​​​⁠​‌​⁠​‌​⁠‌⁠​‍​‍​‍⁠​​‍​‍‌‍‍​​‍​‍​⁠‍‍​‍​‍‌‌⁠⁠‌‍‍‌‌‌​⁠‌​​‌‌‌​​‌‌‌⁠‌‌⁠⁠‌‌​‍‌‌‍‌‌‌​‌‌‌⁠⁠‌⁠‌‍‌‍⁠‍‌‌‌​​⁠‌⁠​‍⁠‌​‍​‍‌⁠⁠‌

Quick tip from a plant outage: we stamped each job with an “expiry_at” (epoch ms) and had workers abort if wall time had passed it — like telling late‑arriving guests the kitchen’s closed. We also inject a small randomized delay (10–60 ms) on accept to smooth bursts; @dan_murphy34 it plays well with your gateway downgrade, caveat that clock skew will drop some valid work if time sync drifts.

‌⁠‍⁠​‍​‍‌⁠‌​​‍​‍​⁠‍‍​‍​‍‌‍‌‌‌‍⁠‍‌‍‌⁠‌‍‍‌‌‍⁠‍‌‍‌‌‌‍‌‌‌⁠​‍‌‍‍‌‌‍⁠‍‌‍‌⁠​‍​‍​‍⁠​​‍​‍‌‍‍⁠​‍​‍​⁠‍‍​‍​‍‌⁠​‍‌‍‌‌‌⁠​​‌‍⁠​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍‌‌‍‌​​‍​‍​⁠‍‍​⁠​‌​⁠‍‌​⁠‍​​‍⁠​​‍​‍‌‍‌​​‍​‍​⁠‍‍​‍​‍​⁠​‍​⁠​​​⁠​‍​⁠‌‍​⁠​​​⁠​‌​⁠​‍​⁠​⁠​‍​‍​‍⁠​​‍​‍‌‍‍​​‍​‍​⁠‍‍​‍​‍‌​⁠‌‌​⁠​‌‍⁠‌‌​‌⁠‌‌‌‌​‍⁠‌‌​⁠​‌⁠​​​⁠‌⁠​⁠​​‌‍​‌‌‍‌⁠​‍⁠‌‌‍‍⁠‌‍‌‌​⁠‌‌​‍​‍‌⁠⁠‌

Quick example: we stopped a retry storm by switching clients to decorrelated jitter backoff (AWS “Full Jitter”) and a small per‑request retry budget — like metering lights on a freeway on‑ramp. Minor downside: tiny extra tail latency during quick blips, but recovery under load was much smoother. Link: Exponential Backoff And Jitter | AWS Architecture Blog.

‌⁠‍⁠​‍​‍‌⁠‌​​‍​‍​⁠‍‍​‍​‍‌‍‌‌‌‍⁠‍‌‍‌⁠‌‍‍‌‌‍⁠‍‌‍‌‌‌‍‌‌‌⁠​‍‌‍‍‌‌‍⁠‍‌‍‌⁠​‍​‍​‍⁠​​‍​‍‌‍‍⁠​‍​‍​⁠‍‍​‍​‍‌⁠​‍‌‍‌‌‌⁠​​‌‍⁠​‌⁠‍‌​‍​‍​‍⁠​​‍​‍‌‍‍‌‌‍‌​​‍​‍​⁠‍‍​⁠​‌​⁠‍‌​⁠‍​​‍⁠​​‍​‍‌‍‌​​‍​‍​⁠‍‍​‍​‍​⁠​‍​⁠​​​⁠​‍​⁠‌‍​⁠​​​⁠​‌​⁠​‍​⁠‌​​‍​‍​‍⁠​​‍​‍‌‍‍​​‍​‍​⁠‍‍​‍​‍​⁠‍​‌​⁠⁠‌​‍​‌‌‌​‌‍​‍‌⁠‌‍‌‌‍‌‌‍​⁠‌⁠‍​‌‍⁠⁠‌‍⁠⁠‌‍‌‍‌‍⁠‍‌‍⁠​‌⁠‌‍‌⁠‍​​‍​‍‌⁠⁠‌