MOMO / BENCHMARK CENTER

Agent Evaluation Center

Understand procurement recommendations through task outcomes, measured service performance, and token costs.

Explore the market

Loading evaluation evidence…

01

BFCL V4

Function calling

55 fixed tasks × 2 trials

2026.3.23

02

τ³

Business workflows

15 fixed tasks × 2 trials

1.0.1

03

AgentDojo

Utility & security

12 fixed tasks × 2 trials · attack control

0.1.35

04

Gaia2

Dynamic environments

15 fixed tasks × 2 trials

ARE-1.2.0/78ea3bdbdeec

Published procurement evaluations

Fixed samples · Reproducible · 30-day validity

Preparing the first standard evaluation

The first cohort covers momo_278, momo_280, and momo_247. Results enter procurement recommendations after the shared sample, token ledger, and publication checks are complete.

01 / Fixed tasks

Versioned frameworks, samples, execution, and grading.

02 / Token accounting

Separate target and auxiliary calls, reconcile usage and debits.

03 / Procurement evidence

Native results, service performance, and cost per successful task.

Agent Benchmarks & Token Procurement Evidence | MomoAI