One task.
Two interfaces.
No soft scoring.
A paired agent benchmark asks whether Fiber turns a literal nine-tool Stripe surface into a more reliable workflow—without changing the model, task, or transport.
Does the interface make the same agent better?
Using the same model and Stripe transport, Fiber's reviewed workflow interface will improve end-to-end completion and reduce tool calls, argument errors, and unsafe cleanup behavior compared with a literal one-operation-per-tool interface.
Literal baseline
Nine one-operation tools expose raw Stripe product, price, and payment-link arguments to the model.
Fiber surface
One approval-gated workflow plans, verifies, and safely deactivates the same nine underlying operations.
Deterministic judge
Code—not another model—checks ownership, relationships, completion, errors, and final cleanup state.
The benchmark is preregistered. Comparative results remain hidden until all paired model runs and cleanup checks are verified.