Most systems do not fail because someone chose the wrong database. They fail because a reasonable decision kept being reasonable long after the conditions that justified it had changed.
This piece walks through how we think about it at Paycux, what we have changed our minds about, and where the sharp edges are.
How it breaks
That brings us to how it breaks. Make the boundary explicit. When one part of the system can only talk to another through a named interface, you can change either side without a meeting; when it cannot, every change becomes a negotiation.
Measure before you optimise, then measure the thing users feel rather than the thing that is easy to instrument. A p50 that looks fine while the p99 is unusable is a reporting failure, not a performance one.
What to do instead
What to do instead is where this gets concrete. Make the boundary explicit. When one part of the system can only talk to another through a named interface, you can change either side without a meeting; when it cannot, every change becomes a negotiation.
Make the boundary explicit. When one part of the system can only talk to another through a named interface, you can change either side without a meeting; when it cannot, every change becomes a negotiation.
- Name the boundary before you cross it
- Measure what users feel, not what is easy to instrument
- Make every retried operation idempotent
- Write down the assumption that would invalidate the design
Make the boundary explicit.
Where this leaves us
Consider where this leaves us. Idempotency is not a nice-to-have in any system with retries. Key on an identifier the caller supplies, store the outcome, and return the same answer to the same key rather than doing the work twice.
Idempotency is not a nice-to-have in any system with retries. Key on an identifier the caller supplies, store the outcome, and return the same answer to the same key rather than doing the work twice.
Where this leaves us
The pattern repeats across every system we have looked at: the hard part is not the mechanism, it is keeping the mechanism honest as the surrounding assumptions change.
If you are working through the same problem and want to compare notes, the docs cover the mechanics and the console shows the behaviour on your own data.
Everything here, already built
Sign-in, enterprise SSO, directory provisioning, roles and an audit trail behind one API. Start with the quickstart and have a working sign-in this afternoon.