Skip to content
All articles
Engineering29 May 2025·4 min read

OpenAI's Codex wants to become your AI coworker

The uncomfortable part of building for agents is that the old assumptions quietly stop holding.

Paycux engineering

The uncomfortable part of building for agents is that the old assumptions quietly stop holding. There is no browser to redirect, no person to read a consent screen at the moment it matters, and no obvious place to put the word "no".

This piece walks through how we think about it at Paycux, what we have changed our minds about, and where the sharp edges are.

From Autocomplete to autonomous agent

That brings us to from autocomplete to autonomous agent. An agent's credential should describe what it may do, not who owns it. Scope it to the narrowest set of operations that make the task possible, bind it to a single principal, and give it a lifetime measured in minutes rather than days.

The audit trail matters more here than in any human flow. When something goes wrong, the question is never "did the user intend this" in the abstract — it is which agent, acting under whose authority, made which call, and whether the record can prove it.

Trained for real-world development

Trained for real-world development deserves its own treatment. The audit trail matters more here than in any human flow. When something goes wrong, the question is never "did the user intend this" in the abstract — it is which agent, acting under whose authority, made which call, and whether the record can prove it.

An agent's credential should describe what it may do, not who owns it. Scope it to the narrowest set of operations that make the task possible, bind it to a single principal, and give it a lifetime measured in minutes rather than days.

  • Scope every agent credential to one principal and one task
  • Give tokens minutes of life, not days
  • Record which agent acted, under whose authority, on what
  • Make revocation a single call that takes effect immediately

How it actually works

How it actually works deserves its own treatment. Consent is the part most implementations get wrong. Asking once at install time and then acting indefinitely is not consent; it is a standing grant with no expiry and no visibility.

Consent is the part most implementations get wrong. Asking once at install time and then acting indefinitely is not consent; it is a standing grant with no expiry and no visibility.

Secure sandboxes

Consider secure sandboxes. Consent is the part most implementations get wrong. Asking once at install time and then acting indefinitely is not consent; it is a standing grant with no expiry and no visibility.

An agent's credential should describe what it may do, not who owns it. Scope it to the narrowest set of operations that make the task possible, bind it to a single principal, and give it a lifetime measured in minutes rather than days.

The audit trail matters more here than in any human flow.

Project-aware intelligence

That brings us to project-aware intelligence. Consent is the part most implementations get wrong. Asking once at install time and then acting indefinitely is not consent; it is a standing grant with no expiry and no visibility.

Consent is the part most implementations get wrong. Asking once at install time and then acting indefinitely is not consent; it is a standing grant with no expiry and no visibility.

End-to-end workflows

End-to-end workflows is where this gets concrete. Consent is the part most implementations get wrong. Asking once at install time and then acting indefinitely is not consent; it is a standing grant with no expiry and no visibility.

An agent's credential should describe what it may do, not who owns it. Scope it to the narrowest set of operations that make the task possible, bind it to a single principal, and give it a lifetime measured in minutes rather than days.

OpenAI Codex use cases

OpenAI Codex use cases deserves its own treatment. The audit trail matters more here than in any human flow. When something goes wrong, the question is never "did the user intend this" in the abstract — it is which agent, acting under whose authority, made which call, and whether the record can prove it.

The audit trail matters more here than in any human flow. When something goes wrong, the question is never "did the user intend this" in the abstract — it is which agent, acting under whose authority, made which call, and whether the record can prove it.

Where this leaves us

If there is one thing worth taking away, it is that the expensive decisions are the ones made implicitly. Making them on purpose costs an afternoon.

If you are working through the same problem and want to compare notes, the docs cover the mechanics and the console shows the behaviour on your own data.

Everything here, already built

Sign-in, enterprise SSO, directory provisioning, roles and an audit trail behind one API. Start with the quickstart and have a working sign-in this afternoon.

Start selling to enterprise customers

Create an account, point sign-in at Paycux, and get back to the part of the product that is actually yours.