FiberBench / Stripe / 01

One task.
Two interfaces.
No soft scoring.

A paired agent benchmark asks whether Fiber turns a literal nine-tool Stripe surface into a more reliable workflow—without changing the model, task, or transport.

25 paired runs50 total trialsEvidence pending
Paired result
locked until evidence gate passes
MetricLiteralFiber
Completion
Tool calls / run
Argument errors
Tokens / run
Remediation runs
Result cells are intentionally blank. The page cannot render comparative numbers from mock or partial evidence.
The question

Does the interface make the same agent better?

Using the same model and Stripe transport, Fiber's reviewed workflow interface will improve end-to-end completion and reduce tool calls, argument errors, and unsafe cleanup behavior compared with a literal one-operation-per-tool interface.

Same model and settings
Same five frozen tasks
Randomized surface order
Run-owned Stripe objects
Deterministic cleanup grading
Redacted trace per trial
01

Literal baseline

Nine one-operation tools expose raw Stripe product, price, and payment-link arguments to the model.

02

Fiber surface

One approval-gated workflow plans, verifies, and safely deactivates the same nine underlying operations.

03

Deterministic judge

Code—not another model—checks ownership, relationships, completion, errors, and final cleanup state.

Publication gate

The benchmark is preregistered. Comparative results remain hidden until all paired model runs and cleanup checks are verified.

Inspect protocol