<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Experiments on cphaynes.app</title><link>https://cphaynes.app/tags/experiments/</link><description>Recent content in Experiments on cphaynes.app</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Sun, 20 Sep 2026 07:31:35 -0400</lastBuildDate><atom:link href="https://cphaynes.app/tags/experiments/index.xml" rel="self" type="application/rss+xml"/><item><title>Run the Blank First</title><link>https://cphaynes.app/posts/2026-09-20-run-the-blank-first/</link><pubDate>Sun, 20 Sep 2026 07:31:35 -0400</pubDate><guid>https://cphaynes.app/posts/2026-09-20-run-the-blank-first/</guid><description>&lt;p&gt;In the comment on my last post, Claude Sonnet 5 ended with a line I have not been able to put down: if temperature, system prompt, or model version drift between your ten draws, then &amp;ldquo;seven out of ten&amp;rdquo; isn&amp;rsquo;t measuring belief, it&amp;rsquo;s measuring your test harness. I want to leave the twins alone today and follow that one sentence somewhere else, because it names something that is much older than language models and that I think most people, including me, systematically under-do.&lt;/p&gt;</description></item></channel></rss>