<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Evaluation on Herman</title><link>https://hermanity.dev/tags/evaluation/</link><description>Recent content in Evaluation on Herman</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Wed, 15 Jul 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://hermanity.dev/tags/evaluation/index.xml" rel="self" type="application/rss+xml"/><item><title>Model Bench</title><link>https://hermanity.dev/projects/model-bench/</link><pubDate>Wed, 15 Jul 2026 00:00:00 +0000</pubDate><guid>https://hermanity.dev/projects/model-bench/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;A benchmarking site I built to answer one question well: &lt;strong&gt;which
subscription-based LLM is worth its slot on my dashboard, and which
pay-as-you-go API is the better substitute.&lt;/strong&gt; Every public benchmark
either (a) scores only pay-as-you-go APIs (so the results don&amp;rsquo;t tell
me what a Claude Max / ChatGPT Plus / ollama-routed model would
actually do on my workloads) or (b) makes vague quality claims with no
cost column (so I can&amp;rsquo;t compare a $20/month subscription to a $5 in
API spend). This benchmark measures both, on the same tasks, in the
same week, scored by a judge from a different model family.&lt;/p&gt;</description></item></channel></rss>