<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Benchmark on Herman</title><link>https://hermanity.dev/tags/benchmark/</link><description>Recent content in Benchmark on Herman</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Thu, 16 Jul 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://hermanity.dev/tags/benchmark/index.xml" rel="self" type="application/rss+xml"/><item><title>Skynet Breakage Atlas</title><link>https://hermanity.dev/projects/breakage-atlas/</link><pubDate>Thu, 16 Jul 2026 00:00:00 +0000</pubDate><guid>https://hermanity.dev/projects/breakage-atlas/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;The &lt;strong&gt;Skynet Breakage Atlas&lt;/strong&gt; answers a practical routing question: when traffic goes through our proxy fleet instead of first-party subscription surfaces, which models go empty, hard-error, or slow-walk under adversarial families?&lt;/p&gt;
&lt;p&gt;Live at &lt;a href="https://breakage.hermanity.dev/"&gt;breakage.hermanity.dev&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="whats-measured"&gt;What’s measured&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Controls:&lt;/strong&gt; Claude Max, Codex, Grok (subscription / direct)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Fleet:&lt;/strong&gt; Skynet-routed models (deepseek, glm, kimi, minimax, …)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Families:&lt;/strong&gt; cascade seed, context boundary, cost threshold, empty probe, format JSON, multi-intent&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Metrics:&lt;/strong&gt; flag rate, empty rate, hard error rate, p50 latency&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="snapshot-public-table-on-the-site"&gt;Snapshot (public table on the site)&lt;/h2&gt;
&lt;p&gt;glm-5.2 is the known-broken reference (flag rate 1.0 across cells in the frozen N-batch). Controls stay low-flag; fleet variance is the point of the atlas — empty replies are a first-class failure mode, not a footnote.&lt;/p&gt;</description></item><item><title>Model Bench</title><link>https://hermanity.dev/projects/model-bench/</link><pubDate>Wed, 15 Jul 2026 00:00:00 +0000</pubDate><guid>https://hermanity.dev/projects/model-bench/</guid><description>&lt;h2 id="overview"&gt;Overview&lt;/h2&gt;
&lt;p&gt;A benchmarking site I built to answer one question well: &lt;strong&gt;which
subscription-based LLM is worth its slot on my dashboard, and which
pay-as-you-go API is the better substitute.&lt;/strong&gt; Every public benchmark
either (a) scores only pay-as-you-go APIs (so the results don&amp;rsquo;t tell
me what a Claude Max / ChatGPT Plus / ollama-routed model would
actually do on my workloads) or (b) makes vague quality claims with no
cost column (so I can&amp;rsquo;t compare a $20/month subscription to a $5 in
API spend). This benchmark measures both, on the same tasks, in the
same week, scored by a judge from a different model family.&lt;/p&gt;</description></item></channel></rss>