The one piece of monitoring I added after an incident, and now put in every service on day one, is queue depth over time. I care about the trend far more than the current number.
A large queue that drains fast is fine. A smaller one that has grown steadily for an hour is a problem: consumers are falling behind producers, and that won't fix itself.
The slope tells you whether you have twenty minutes or two. Divide the remaining headroom by the growth rate and you get a rough time until it hurts, which is what decides whether you page someone, scale consumers, or keep watching.
Building something like this?
I'm Ahmed Mamdouh, a senior full-stack & AI engineer. I reply within one working day.

Scaling 100k WebSocket connections: the reconnect storm
At 100k+ concurrent sockets the count is easy. The reconnect storm is what breaks, and jittered backoff, load shedding and resumable sessions fix it.

Splitting a monolith: what it actually bought us
Split a monolith for failure isolation and team ownership. If speed is the goal, profile first; the cause is usually a missing index.

Caching mistakes I keep debugging
Before adding a cache, decide what invalidates it, measure first, set a TTL, and cache the expensive part instead of the whole result.
