On a platform I worked on, we split a monolith into services and p95 latency dropped about 40 percent. That's the least interesting result. What we actually got was failure isolation, and what it cost us is the part conference talks leave out.
What stopped happening
One bad deploy used to take the whole product down. After the split, a bad deploy took down one service and the rest kept serving. That's the reason I'd split a monolith, along with clearer ownership for each team. The latency drop was a side effect.
What it cost
Local development got worse. Running the full product on a laptop went from one command to a compose file nobody fully understood.
Debugging got worse too. A stack trace turned into correlating three logs across two services. Distributed tracing went from nice to have to something you install before the split, because adding it afterwards means debugging blind in the meantime.
When splitting won't help
If your monolith is slow because one endpoint is doing an N+1 query, splitting won't fix it. You'll have the same slow query in a smaller box, with a network hop in front of it. If speed is your reason for splitting, profile first. The answer is usually a missing index.
Building something like this?
I'm Ahmed Mamdouh, a senior full-stack and AI engineer. I reply within one working day.

Watch queue depth trend, not the current number
Track queue depth over time from day one. The current number says little; the slope tells you whether you have twenty minutes or two.

Scaling 100k WebSocket connections: the reconnect storm
At 100k+ concurrent sockets the count is easy. The reconnect storm is what breaks, and jittered backoff, load shedding and resumable sessions fix it.

npm v12 blocks install scripts by default
npm v12 no longer runs preinstall, install or postinstall scripts unless you approve them, closing a common supply chain attack path.
