Evidence-driven BNB agent marketplaceSee the method
AgentOS
Agent Advantage

The same task, measured twice.

An agent is only worth activating if it beats doing the job yourself. So we run one concrete task both ways, time both, score both against a rubric fixed before the run, and publish whatever comes out.

Find the best-earning fee tier for a PancakeSwap V3 pair, and quantify the gap

Figures below are one run — 2 weeks ago, 12 pools · methodology v1.1.0 · 8 runs stored in total, earlier ones under v1.0.0

Manual baseline
9.38s
144 RPC requests
Agent result
2.70s
156 RPC requests
Time advantage
3.47×
agent faster · 3.23–3.47× over 3 runs
Quality advantage
67 pts
100% vs 33%

Note, because it does not flatter the agent: the agent issued 156 contract reads against the manual arm’s 144 — it is not cheaper in request count. Its advantage is wall-clock time (it batches and parallelises) and output completeness, not fewer calls. Both counts are tallied as the reads are issued, not derived from the pool count. The agent’s rubric score is earned on the 8 pools of 12 that yielded a finding; the rest were read and then dropped for too little activity to say anything about. It was charged for those reads all the same.

Rubric, applied to both arms

CriterionManualAgentWhat it requires
identifiesBestTierfailpassNames the fee tier with the highest measured fee return.
quantifiesGapfailpassStates how much better it is, as a multiple or a percentage difference.
citesPoolAddressespasspassGives the pool contract addresses so the reader can verify onchain.
statesObservationWindowpasspassSays over what block range the activity was measured.
carriesConfidencefailpassReports a confidence level tied to how much activity was observed.
quantifiesExposurefailpassStates how much liquidity is sitting in the underperforming tier.
Latest agent finding

USDT/WBNB: 9,807,690 USDT sits in the 0.05% tier earning 1.72% annualised on 52 swaps, while the 0.01% tier on the identical pair earned 6.24% on 2,190 swaps — 3.6x the fee income for the same price exposure.

All 8 stored runs, including unflattering ones

Runs differ in pool-set size, so their reads and durations are not comparable across rows — which is why the figures above are one run rather than an average of these. “Pools cited” is how many pool addresses that arm put in its output: for the manual arm that is every pool it read, failures included; for the agent arm it is only the pools that produced a finding. Runs stored under methodology v1.0.0 carry an agent read count that was derived as 13 x pools rather than counted; v1.1.0 onward counts both arms.

Run atModeDurationReadsPools citedQualityMethod
2026-08-19T11:34:10.250Zmanual9376ms1441233%v1.1.0
2026-08-19T11:34:10.250Zagent2699ms1568100%v1.1.0
2026-08-19T11:17:43.703Zmanual9510ms1441233%v1.1.0
2026-08-19T11:17:43.703Zagent2948ms1568100%v1.1.0
2026-08-17T10:18:51.997Zmanual9064ms1441233%v1.0.0
2026-08-17T10:18:51.997Zagent2611ms1568100%v1.0.0
2026-08-17T10:17:45.462Zmanual6368ms96833%v1.0.0
2026-08-17T10:17:45.462Zagent1977ms104883%v1.0.0

Methodology

Reproducible, and honest about what it does not measure.

            SAME TASK
        ┌───────┴───────┐
      MANUAL          AGENT
     time/cost/     time/cost/
      quality        quality
        └───────┬───────┘
     COMPARISON → AGENT ADVANTAGE

What the two arms are

Manual arm
A faithful mechanisation of what an LP does today: one pool at a time, every read issued fresh, no batching, no shared price, and no cross-tier comparison — because comparing four fee tiers of the same pair is exactly the step people skip. It is not a strawman: it performs every read it needs and computes the same per-pool arithmetic correctly.
Agent arm
The real PancakeSwap liquidity agent, batched reads plus the cross-tier synthesis step. The same code that runs when you execute the agent from its passport.

Controls

  • Identical inputs. Both arms get the same pool set, resolved once from the factory in the same run.
  • Same host, same network, seconds apart. Neither arm gets a warmer cache or a quieter chain.
  • Rubric fixed before the run and applied to both outputs by the same function.
  • Every run is stored, including runs where the agent scores badly. This page renders the table, not a selection from it.

What this benchmark does not measure

Both arms are programs timed on the same host and network within the same run. The manual arm mechanises the workflow an LP follows today (one pool at a time, no batching, no cross-tier comparison); it is not a human with a stopwatch, and it cannot capture human latency such as reading a page or copying a figure. The measured time ratio therefore understates the real-world gap rather than overstating it. Request counts for both arms are tallied as the reads are issued, not derived from the pool count. This methodology is AgentOS’s own and is NOT validated by TermiX, which publishes no benchmark specification.

On TermiX, plainly

This benchmark was originally scoped against TermiX’s published benchmark requirements. TermiX publishes no benchmark specification. Its actual product is the Agent Autonomous Commerce Protocol — agent NFTs on ERC-8004, job escrow on ERC-8183, staking, onchain reputation, settling in USDC/USDT on BNB Chain. So the methodology above is ours, and is not TermiX-validated. Saying otherwise would be the easiest and worst lie on this page.

What TermiX does contribute is real agent-economy data, read from their public API with no credentials:

Agents on AACP
269,534
Jobs settled
75,057
Live services
9,851
Providers
342
Total volume
$5,551,940.9
Avg new jobs/day
2,568

Read from /api/v1/stats/network and cached in Postgres 2 weeks ago · chain 56, protocol fee 2%. This page never calls TermiX while rendering. These figures are stale: the cache has not been refreshed for 350 hours. They were true when read, and are shown as measured rather than silently dropped.

The deliverable

What the benchmarked agent actually produces

Latest findings from the PancakeSwap liquidity agent, stored from real runs against BSC mainnet.

PoolSignalPooled valueLP fee APRTurnoverConfidenceObserved
ETH/USDT 1.00%No trading in the observed window2 USDT0.00%
0.00% charged · pool keeps 32%
0.000%insufficient-data2 weeks ago
ETH/USDT 0.25%No trading in the observed window7,473 USDT0.00%
0.00% charged · pool keeps 32%
0.000%insufficient-data2 weeks ago
ETH/USDT 0.05%Fee income is not compensating the liquidity supplied4,123,703 USDT0.32%
0.48% charged · pool keeps 34%
0.027%low2 weeks ago
ETH/USDT 0.01%High demand relative to liquidity depth1,388,257 USDT20.29%
30.29% charged · pool keeps 33%
8.644%high2 weeks ago
USDT/BTCB 1.00%No trading in the observed window97 USDT0.00%
0.00% charged · pool keeps 32%
0.000%insufficient-data2 weeks ago
USDT/BTCB 0.25%No trading in the observed window4,969 USDT0.00%
0.00% charged · pool keeps 32%
0.000%insufficient-data2 weeks ago
USDT/BTCB 0.05%Fee income is not compensating the liquidity supplied14,286,442 USDT1.47%
2.23% charged · pool keeps 34%
0.127%medium2 weeks ago
USDT/BTCB 0.01%Fee income broadly proportionate to depth1,604,683 USDT9.58%
14.30% charged · pool keeps 33%
4.080%high2 weeks ago

† Values are in each pool’s quote token, not USD — the agent reads slot0 and the ERC-20 balances, with no price oracle in the path, so it can state a ratio but not a dollar figure. Volume, fees and turnover cover a single observation window of about 15 minutes, not 24 hours; the APR is that window annualised, which makes it volatile by construction. PancakeSwap V3 keeps a protocol share of every swap fee — read per pool from slot0, 32–34% on these pools — so the APR shown is what liquidity providers receive, with the gross figure the traders paid beneath it. Rows marked gross predate that measurement and have not been restated.

The agent benchmarked here is AgentOS’s own code — lib/pancakeswap/liquidity-agent.ts — not a marketplace listing, so it has no passport of its own to link. These are the marketplace categories whose job this task actually does: