Skip to content
All articles
Engineering8 January 2026·4 min read

Fireworks.ai: The PyTorch Team's Bet on Inference as the New Runtime

Most of the difficulty with agents is not the model.

Paycux engineering

Most of the difficulty with agents is not the model. It is that an agent sits between a user and a system that was designed on the assumption a user would be there in person.

This piece walks through how we think about it at Paycux, what we have changed our minds about, and where the sharp edges are.

The founding story

That brings us to the founding story. An agent's credential should describe what it may do, not who owns it. Scope it to the narrowest set of operations that make the task possible, bind it to a single principal, and give it a lifetime measured in minutes rather than days.

The audit trail matters more here than in any human flow. When something goes wrong, the question is never "did the user intend this" in the abstract — it is which agent, acting under whose authority, made which call, and whether the record can prove it.

Why inference is the battleground

Consider why inference is the battleground. The audit trail matters more here than in any human flow. When something goes wrong, the question is never "did the user intend this" in the abstract — it is which agent, acting under whose authority, made which call, and whether the record can prove it.

Consent is the part most implementations get wrong. Asking once at install time and then acting indefinitely is not consent; it is a standing grant with no expiry and no visibility.

  • Scope every agent credential to one principal and one task
  • Give tokens minutes of life, not days
  • Record which agent acted, under whose authority, on what
  • Make revocation a single call that takes effect immediately

Technical differentiation: performance plus orchestration

Consider technical differentiation: performance plus orchestration. Consent is the part most implementations get wrong. Asking once at install time and then acting indefinitely is not consent; it is a standing grant with no expiry and no visibility.

The audit trail matters more here than in any human flow. When something goes wrong, the question is never "did the user intend this" in the abstract — it is which agent, acting under whose authority, made which call, and whether the record can prove it.

Compound AI systems

Compound AI systems deserves its own treatment. The audit trail matters more here than in any human flow. When something goes wrong, the question is never "did the user intend this" in the abstract — it is which agent, acting under whose authority, made which call, and whether the record can prove it.

An agent's credential should describe what it may do, not who owns it. Scope it to the narrowest set of operations that make the task possible, bind it to a single principal, and give it a lifetime measured in minutes rather than days.

The audit trail matters more here than in any human flow.

The product stack

That brings us to the product stack. The audit trail matters more here than in any human flow. When something goes wrong, the question is never "did the user intend this" in the abstract — it is which agent, acting under whose authority, made which call, and whether the record can prove it.

Consent is the part most implementations get wrong. Asking once at install time and then acting indefinitely is not consent; it is a standing grant with no expiry and no visibility.

Serverless inference

Consider serverless inference. Consent is the part most implementations get wrong. Asking once at install time and then acting indefinitely is not consent; it is a standing grant with no expiry and no visibility.

Consent is the part most implementations get wrong. Asking once at install time and then acting indefinitely is not consent; it is a standing grant with no expiry and no visibility.

On-demand deployments

Consider on-demand deployments. Consent is the part most implementations get wrong. Asking once at install time and then acting indefinitely is not consent; it is a standing grant with no expiry and no visibility.

The audit trail matters more here than in any human flow. When something goes wrong, the question is never "did the user intend this" in the abstract — it is which agent, acting under whose authority, made which call, and whether the record can prove it.

Where this leaves us

None of this is exotic. It is the ordinary discipline of deciding what you own, writing down what you assume, and making the failures loud enough to notice.

If you are working through the same problem and want to compare notes, the docs cover the mechanics and the console shows the behaviour on your own data.

Everything here, already built

Sign-in, enterprise SSO, directory provisioning, roles and an audit trail behind one API. Start with the quickstart and have a working sign-in this afternoon.

Start selling to enterprise customers

Create an account, point sign-in at Paycux, and get back to the part of the product that is actually yours.