$ ls ~/tags

Evaluation

active

Model Bench

Empirical cross-provider LLM benchmark — 22,522 trials × 64 models × 9 tasks, with subscription-vs-public-API cost columns side-by-side and industry-benchmark side tracks.

Read more