RouterBench

•  ai credit score™  •

A live rating for
every AI model.

Static benchmarks go stale the day a new model ships. The AI Credit Score™ scores every model continuously — per task — so routing decisions rest on today's data, not last quarter's leaderboard.

•  the dimensions  •

Five signals. One score. Per task.

01 — accuracy

Gets it right?

Does the model actually solve the task — measured on live traffic, not synthetic tests.

02 — reliability

Behaves consistently?

Stability across runs, formats and edge cases. Flaky models lose score.

03 — cost efficiency

Worth the price?

Quality per dollar. A cheaper model that's good enough outranks an expensive one that isn't better.

04 — latency

Answers fast?

Time-to-first-token and completion speed, tracked per provider and region.

05 — real-world

Wins in production?

How it performs on real workloads — the signal static leaderboards can never capture.

•  why it matters  •

Benchmarks tell you yesterday. The score tells you now.

  • live Updated from real traffic — the score moves when models move
  • per-task The best model for code isn't the best for extraction
  • weighted Quality × cost × latency, tuned to your routing profile
  • unbiased Every provider — including RouterBench's own models — scored by the same yardstick
  • automatic Powers routerbench/auto — every routed request benefits

Stop routing on gut feel.

Let the live score pick — and see the reasoning in your logs.