Most systems do not fail because someone chose the wrong database. They fail because a reasonable decision kept being reasonable long after the conditions that justified it had changed.
This piece walks through how we think about it at Paycux, what we have changed our minds about, and where the sharp edges are.
The infrastructure layer you didn’t plan to build
The infrastructure layer you didn’t plan to build is where this gets concrete. Make the boundary explicit. When one part of the system can only talk to another through a named interface, you can change either side without a meeting; when it cannot, every change becomes a negotiation.
Idempotency is not a nice-to-have in any system with retries. Key on an identifier the caller supplies, store the outcome, and return the same answer to the same key rather than doing the work twice.
Authentication: not just users anymore
Authentication: not just users anymore is where this gets concrete. Idempotency is not a nice-to-have in any system with retries. Key on an identifier the caller supplies, store the outcome, and return the same answer to the same key rather than doing the work twice.
Idempotency is not a nice-to-have in any system with retries. Key on an identifier the caller supplies, store the outcome, and return the same answer to the same key rather than doing the work twice.
- Name the boundary before you cross it
- Measure what users feel, not what is easy to instrument
- Make every retried operation idempotent
- Write down the assumption that would invalidate the design
Authorization: where real complexity lives
Authorization: where real complexity lives is where this gets concrete. Measure before you optimise, then measure the thing users feel rather than the thing that is easy to instrument. A p50 that looks fine while the p99 is unusable is a reporting failure, not a performance one.
Idempotency is not a nice-to-have in any system with retries. Key on an identifier the caller supplies, store the outcome, and return the same answer to the same key rather than doing the work twice.
Multi-tenancy: isolation as a first-class constraint
That brings us to multi-tenancy: isolation as a first-class constraint. Idempotency is not a nice-to-have in any system with retries. Key on an identifier the caller supplies, store the outcome, and return the same answer to the same key rather than doing the work twice.
Measure before you optimise, then measure the thing users feel rather than the thing that is easy to instrument. A p50 that looks fine while the p99 is unusable is a reporting failure, not a performance one.
Idempotency is not a nice-to-have in any system with retries.
User provisioning and lifecycle management
Consider user provisioning and lifecycle management. Measure before you optimise, then measure the thing users feel rather than the thing that is easy to instrument. A p50 that looks fine while the p99 is unusable is a reporting failure, not a performance one.
Idempotency is not a nice-to-have in any system with retries. Key on an identifier the caller supplies, store the outcome, and return the same answer to the same key rather than doing the work twice.
Self-serve onboarding, without losing control
Self-serve onboarding, without losing control is where this gets concrete. Idempotency is not a nice-to-have in any system with retries. Key on an identifier the caller supplies, store the outcome, and return the same answer to the same key rather than doing the work twice.
Make the boundary explicit. When one part of the system can only talk to another through a named interface, you can change either side without a meeting; when it cannot, every change becomes a negotiation.
Real-time protection: enforcing trust as the system scales
Real-time protection: enforcing trust as the system scales is where this gets concrete. Measure before you optimise, then measure the thing users feel rather than the thing that is easy to instrument. A p50 that looks fine while the p99 is unusable is a reporting failure, not a performance one.
Measure before you optimise, then measure the thing users feel rather than the thing that is easy to instrument. A p50 that looks fine while the p99 is unusable is a reporting failure, not a performance one.
Where this leaves us
If there is one thing worth taking away, it is that the expensive decisions are the ones made implicitly. Making them on purpose costs an afternoon.
If you are working through the same problem and want to compare notes, the docs cover the mechanics and the console shows the behaviour on your own data.
Everything here, already built
Sign-in, enterprise SSO, directory provisioning, roles and an audit trail behind one API. Start with the quickstart and have a working sign-in this afternoon.