MOMO / BENCHMARK CENTER
Agent Evaluation Center
Understand procurement recommendations through task outcomes, measured service performance, and token costs.
Loading evaluation evidence…
01
BFCL V4
Function calling
55 fixed tasks × 2 trials
2026.3.23
02
τ³
Business workflows
15 fixed tasks × 2 trials
1.0.1
03
AgentDojo
Utility & security
12 fixed tasks × 2 trials · attack control
0.1.35
04
Gaia2
Dynamic environments
15 fixed tasks × 2 trials
ARE-1.2.0/78ea3bdbdeec
Published procurement evaluations
Fixed samples · Reproducible · 30-day validityPreparing the first standard evaluation
The first cohort covers momo_278, momo_280, and momo_247. Results enter procurement recommendations after the shared sample, token ledger, and publication checks are complete.
01 / Fixed tasks
Versioned frameworks, samples, execution, and grading.
02 / Token accounting
Separate target and auxiliary calls, reconcile usage and debits.
03 / Procurement evidence
Native results, service performance, and cost per successful task.