<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Mo RezaAli</title><link>https://therezaali.com/</link><description>Mo RezaAli, Growth Marketer at Yakkyo S.p.A. in Messina, Sicily. Writing with every source listed and checked.</description><language>en</language><managingEditor>contact@therezaali.com (Mo RezaAli)</managingEditor><webMaster>contact@therezaali.com (Mo RezaAli)</webMaster><copyright>© 2026 Mo RezaAli</copyright><generator>Hugo</generator><lastBuildDate>Thu, 06 Aug 2026 00:00:00 +0200</lastBuildDate><atom:link href="https://therezaali.com/index.xml" rel="self" type="application/rss+xml"/><item><title>The model bill is a dollar a month, and the hours are what decide it</title><link>https://therezaali.com/writing/ai-automation-service-thresholds/</link><guid isPermaLink="true">https://therezaali.com/writing/ai-automation-service-thresholds/</guid><pubDate>Thu, 06 Aug 2026 00:00:00 +0200</pubDate><author>contact@therezaali.com (Mo RezaAli)</author><category>essay</category><description>A CVSS 9.3 row-level-security record against the build tool, Gmail’s 0.3% spam ceiling, and GDPR Article 28(4). Then the arithmetic on the line that sets the margin.</description><content:encoded>&lt;p&gt;There is a CVSS 9.3 record against the sentence that sells this business.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;An insufficient database Row-Level Security policy in Lovable through 2025-04-15
allows remote unauthenticated attackers to read or write to arbitrary database
tables of generated sites. NOTE: this is disputed by the Supplier because each
individual customer of the Lovable platform accepts a responsibility over
protecting the data of their application.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;a href="https://nvd.nist.gov/vuln/detail/CVE-2025-48757" class="external-link" rel="noopener"&gt;CVE-2025-48757&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;, National Vulnerability Database&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;The sentence it fires at is the one I wrote myself, and the one every version of
this pitch contains: let the platform own the plumbing, because hosting, database,
auth, deployment and infrastructure security are already handled. Read the
supplier&amp;rsquo;s dispute again. Their defence is that data security in a generated app
belongs to the person who generated it. That is the correct division of
responsibility and it is the exact opposite of what the pitch promises.&lt;/p&gt;
&lt;h2 id="why-i-am-the-one-telling-you-this"&gt;Why I am the one telling you this&lt;/h2&gt;
&lt;p&gt;I have run an unattended system into the ground and I have the commit log.&lt;/p&gt;
&lt;p&gt;From 31 October 2025 to 27 April 2026 a worker on my own domain ran on a cron of
&lt;code&gt;0 */1 * * *&lt;/code&gt;. Every hour it took 450 characters of an RSS &lt;code&gt;description&lt;/code&gt;, passed
them to &lt;code&gt;gpt-4o-mini&lt;/code&gt; with a rewrite pass on &lt;code&gt;gpt-4o&lt;/code&gt;, and committed the output
straight to &lt;code&gt;main&lt;/code&gt;. &lt;code&gt;AUTO_PR&lt;/code&gt; was &lt;code&gt;&amp;quot;false&amp;quot;&lt;/code&gt;, because I never wrote a review step.
It made 842 &lt;code&gt;Add post:&lt;/code&gt; commits and 2,922 files across four languages. 180 of the
739 English posts cited a source called &amp;ldquo;Internal Analysis&amp;rdquo; that does not exist.
I deleted all of it in August 2026 and wrote up
&lt;a href="https://therezaali.com/writing/the-bot-that-published-900-articles/"&gt;the full account&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;No alert fired for six months. Nothing was measuring anything, so nothing had a
threshold to cross, so the pipeline emitted its worst output at a steady hourly
rate until a human finally read it. That is the failure this article is about,
and it is not a story about a model writing badly. It is a story about a system
with no published number that a person checks on a schedule.&lt;/p&gt;
&lt;p&gt;The other thing I did was read the documents. This market&amp;rsquo;s most-repeated
statistic traces to a 2011 magazine article and a 2007 conference deck, both
free, and the deck says in as many words that it did not measure the thing it
gets cited for.&lt;/p&gt;
&lt;p&gt;One piece of housekeeping, because it is the standard I am asking you to hold me
to. My original post carried four revenue ranges. They were one creator&amp;rsquo;s
illustrative figures, I could not name the creator, and an income claim from an
unnamed third party is the exact shape this site fails its own build on. They are
gone. What is left is arithmetic and documents.&lt;/p&gt;
&lt;h2 id="1-the-plumbing-is-owned-until-the-record-says-otherwise"&gt;1. The plumbing is owned until the record says otherwise&lt;/h2&gt;
&lt;p&gt;Read the record carefully rather than as ammunition. It is tagged &lt;code&gt;disputed&lt;/code&gt; and
&lt;code&gt;exclusively-hosted-service&lt;/code&gt;, its NVD status is &lt;code&gt;Deferred&lt;/code&gt;, and the CISA
assessment of 25 June 2025 records exploitation as &lt;code&gt;poc&lt;/code&gt; and automatable as
&lt;code&gt;yes&lt;/code&gt;. Reporting a disputed CVE as settled fact would be the same error this site
exists to correct, so: disputed, and the dispute is the point. The vector is
&lt;code&gt;AV:N/AC:L/PR:N/UI:N/S:C/C:H/I:L/A:N&lt;/code&gt;. Network, low complexity, no privileges, no
user interaction, scope changed. In plain terms the generated app&amp;rsquo;s own tables
were reachable by an unauthenticated stranger. Not a connected CRM, not an
integration somebody wired up carelessly. The database the platform handed over.&lt;/p&gt;
&lt;p&gt;My original post conceded one thing here: that how client data is accessed,
stored and acted on once you connect a real CRM or inbox stays your
responsibility. That concession was too narrow by exactly one layer.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Detect it:&lt;/strong&gt; take the public client-side key your generated app ships in its
front-end bundle, and from a machine that has never signed in, request a table
the app never exposes in its interface. Users, invoices, submissions. If a row
comes back, row-level security is not on for that table. This takes four minutes
and it is the single check whose absence the CVE describes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Design so it cannot recur:&lt;/strong&gt; treat every table the platform generates as
public until you have personally denied it. Default deny, then grant. Run the
unauthenticated read as a test you repeat after every schema change, because the
schema changes when you prompt for a new feature and nothing tells you which
policies came with it.&lt;/p&gt;
&lt;h2 id="2-sixty-seconds-is-not-a-measurement"&gt;2. Sixty seconds is not a measurement&lt;/h2&gt;
&lt;p&gt;I described the product as every inbound lead scored and personally followed up
within 60 seconds. That is a product specification, so it is not false. But it
borrows its authority from a number nobody quoting it has read.&lt;/p&gt;
&lt;p&gt;The 2011 &lt;em&gt;Harvard Business Review&lt;/em&gt; article by Oldroyd, McElheran and Elkington
audited 2,241 US companies by timing each one&amp;rsquo;s response to a web-generated test
lead. 37% responded within an hour, 16% took one to 24 hours, 24% took longer
than a day, and 23% never responded at all. The average among firms that answered
inside 30 days was 42 hours.&lt;/p&gt;
&lt;p&gt;The published distribution is 586, 143, 90, 81, 90, 61, 7, 133, 538 and 512
companies across the buckets from under five minutes to no reply. That sums to
2,241, the stated sample, and 819 of them answered inside the hour, which is
36.5%. The arithmetic holds when you check it, which is worth knowing before you
rely on the next part.&lt;/p&gt;
&lt;p&gt;The multipliers come from a separate dataset in the same article: 1.25 million
leads at 29 B2C and 13 B2B companies. Firms that tried to contact within an hour
were
nearly seven times as likely to &lt;em&gt;qualify&lt;/em&gt; the lead as those trying an hour later,
and more than 60 times as likely as those waiting a day. The article defines
qualify as &amp;ldquo;having a meaningful conversation with a key decision maker.&amp;rdquo; Not a
sale. A conversation.&lt;/p&gt;
&lt;p&gt;The famous &amp;ldquo;21x&amp;rdquo; comes from somewhere else entirely, and it is the one worth
correcting.&lt;/p&gt;
&lt;p&gt;Table: What the popular version of the speed-to-lead statistic claims, against what the two primary documents say.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;The popular version&lt;/th&gt;
 &lt;th&gt;What the primary source says&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;&amp;ldquo;The MIT study&amp;rdquo;&lt;/td&gt;
 &lt;td&gt;InsideSales.com&amp;rsquo;s CEO presenting InsideSales.com&amp;rsquo;s own customer data at MarketingSherpa&amp;rsquo;s B2B Demand Generation Summit, 16 October 2007. James Oldroyd was then a Faculty Fellow at MIT Sloan. Never peer-reviewed.&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;&amp;ldquo;21x more likely to close&amp;rdquo;&lt;/td&gt;
 &lt;td&gt;The deck states: &amp;ldquo;This study did not address close ratios.&amp;rdquo; 21x is the &lt;em&gt;qualify&lt;/em&gt; ratio at five minutes versus 30, and each of the six companies defined &amp;ldquo;qualified&amp;rdquo; its own way.&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;&amp;ldquo;Respond within 60 seconds&amp;rdquo;&lt;/td&gt;
 &lt;td&gt;No primary source measures 60 seconds. The 2007 study&amp;rsquo;s finest granularity is five-minute buckets; the HBR threshold is one hour. Nothing published shows 60 seconds beating five minutes.&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;&amp;ldquo;Applies to email, chat, an AI agent&amp;rdquo;&lt;/td&gt;
 &lt;td&gt;The 2007 deck measures outbound phone dials and says so in its definitions: a dial is &amp;ldquo;the physical action of a sales or lead generation calling a lead.&amp;rdquo; Nothing in it touches another channel.&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;&amp;ldquo;And the 2011 audit confirms it across 2,241 companies&amp;rdquo;&lt;/td&gt;
 &lt;td&gt;The audit never names a channel. Its entire description of method is &amp;ldquo;measuring how long each took to respond to a web-generated test lead.&amp;rdquo; It is silent on medium, which cuts both ways: it is not evidence for the phone, and it is not evidence for an AI agent either.&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;A controlled result&lt;/td&gt;
 &lt;td&gt;Six companies, three years of one vendor&amp;rsquo;s CRM, unadjusted ratios, no control group, no confidence intervals. Leads reached instantly differ systematically from leads reached at 30 minutes.&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;What survives still sells: minutes beat hours, hours beat days, and 23% of the
companies audited never replied at all. That is true, it is checkable, and the
buyer can verify their own reply time in an afternoon.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Detect it:&lt;/strong&gt; before you promise a response time, instrument the client&amp;rsquo;s
current one. Send three test enquiries through their own form at different hours
and time the human reply. That is your before, it costs nothing, and it is the
only number in the pitch that belongs to this client rather than to a 2007
conference deck.&lt;/p&gt;
&lt;h2 id="3-silent-compounding-has-a-number-and-it-is-03"&gt;3. Silent compounding has a number, and it is 0.3%&lt;/h2&gt;
&lt;p&gt;My post said errors compound silently at scale and left it there. Here is the
mechanism, with the threshold and the enforcer named.&lt;/p&gt;
&lt;p&gt;Google&amp;rsquo;s sender guidelines, effective 1 February 2024, put a floor under every
sender: SPF &lt;em&gt;or&lt;/em&gt; DKIM, valid forward and reverse DNS, TLS for transmission, and a
Postmaster Tools spam rate below 0.3%. That &amp;ldquo;or&amp;rdquo; is load-bearing and easy to read
as &amp;ldquo;and&amp;rdquo;, so take it from Google&amp;rsquo;s own two-line summary of the split: &amp;ldquo;All
senders: SPF or DKIM&amp;rdquo; and &amp;ldquo;Bulk senders: SPF, DKIM, and DMARC.&amp;rdquo;
Yahoo draws the line in the same place, telling all senders to &amp;ldquo;implement SPF or
DKIM at a minimum&amp;rdquo; and bulk senders to &amp;ldquo;implement both SPF &amp;amp; DKIM.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Cross 5,000 messages a day to Gmail accounts and the second list attaches: both
SPF and DKIM rather than either, a DMARC record, a &lt;code&gt;From:&lt;/code&gt; domain aligned with
the SPF or DKIM domain, and one-click unsubscribe via
&lt;code&gt;List-Unsubscribe-Post: List-Unsubscribe=One-Click&lt;/code&gt; alongside &lt;code&gt;List-Unsubscribe&lt;/code&gt;.
On complaints the requirement is the same at every volume, a spam rate below
0.3%, and Google&amp;rsquo;s monitoring section tightens it: keep it &amp;ldquo;below 0.10% and avoid
ever reaching a spam rate of 0.30% or higher.&amp;rdquo; Yahoo publishes the same 0.3%
ceiling independently, plus a two-day window for honouring an unsubscribe.&lt;/p&gt;
&lt;p&gt;Now watch it compound. Suppose a follow-up sequence draws one spam complaint per
500 delivered messages. That is 0.2%: under the ceiling, over the 0.10% Google
asks for. At 50 sends a day you clear the authentication bar with SPF alone: no
DKIM, no DMARC, no alignment, no unsubscribe header. Add clients until the same sequence sends
5,000 a day and the bar moves under you without one word of the sequence
changing: now DKIM as well as SPF, now DMARC, now alignment, now one-click
unsubscribe, and the same 0.2% is one bad segment from the ceiling. Cross it and
you degrade the sending domain&amp;rsquo;s reputation. If you did the efficient thing and
put every client on one domain, you degrade delivery for all of them in the same
week. That is the concentrated-platform-risk argument from my post, applied at the
one layer where it is a published number rather than a worry.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Detect it:&lt;/strong&gt; verify the domain in Postmaster Tools before the first send, and
read the spam rate weekly. Alert at 0.10%, not 0.30%. By the time you are at the
ceiling you are already being filtered.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Design so it cannot recur:&lt;/strong&gt; one sending subdomain per client, never a shared
one, so a single bad list cannot take the others with it. Process unsubscribes
by machine inside two days rather than by hand. Send the unsubscribe headers on
every marketing message, not only above the bulk threshold, because the threshold
is where enforcement starts and not where the harm does.&lt;/p&gt;
&lt;h2 id="4-the-margin-variable-is-hours-and-model-spend-is-about-a-dollar"&gt;4. The margin variable is hours, and model spend is about a dollar&lt;/h2&gt;
&lt;p&gt;I told readers to subtract platform and model-usage costs. On the evidence, that
points at the wrong line.&lt;/p&gt;
&lt;p&gt;Table: Claude model IDs and Claude API list prices per million tokens, from Anthropic&amp;rsquo;s models overview, read on 6 August 2026.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Model&lt;/th&gt;
 &lt;th&gt;API ID&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Context&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Input $/MTok&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Output $/MTok&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;Claude Fable 5&lt;/td&gt;
 &lt;td&gt;&lt;code&gt;claude-fable-5&lt;/code&gt;&lt;/td&gt;
 &lt;td style="text-align: right"&gt;1M&lt;/td&gt;
 &lt;td style="text-align: right"&gt;10&lt;/td&gt;
 &lt;td style="text-align: right"&gt;50&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Claude Opus 5&lt;/td&gt;
 &lt;td&gt;&lt;code&gt;claude-opus-5&lt;/code&gt;&lt;/td&gt;
 &lt;td style="text-align: right"&gt;1M&lt;/td&gt;
 &lt;td style="text-align: right"&gt;5&lt;/td&gt;
 &lt;td style="text-align: right"&gt;25&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Claude Sonnet 5&lt;/td&gt;
 &lt;td&gt;&lt;code&gt;claude-sonnet-5&lt;/code&gt;&lt;/td&gt;
 &lt;td style="text-align: right"&gt;1M&lt;/td&gt;
 &lt;td style="text-align: right"&gt;3 (intro 2)&lt;/td&gt;
 &lt;td style="text-align: right"&gt;15 (intro 10)&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Claude Haiku 4.5&lt;/td&gt;
 &lt;td&gt;&lt;code&gt;claude-haiku-4-5-20251001&lt;/code&gt;&lt;/td&gt;
 &lt;td style="text-align: right"&gt;200k&lt;/td&gt;
 &lt;td style="text-align: right"&gt;1&lt;/td&gt;
 &lt;td style="text-align: right"&gt;5&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Sonnet 5&amp;rsquo;s introductory rate runs through 31 August 2026, so anyone quoting it in
September is quoting a stale price.&lt;/p&gt;
&lt;p&gt;Take a content system producing 20 posts a month, at roughly 2,000 input and
2,000 output tokens each, on Opus 5. Input is 0.04 MTok at $5, so $0.20. Output
is 0.04 MTok at $25, so $1.00. Total model spend: about $1.20 a month. Run ten
times that volume and you are under $12. Those token counts are my assumptions
for that shape of job rather than a measurement, and other vendors price
differently, so take the order of magnitude as the finding. The order of
magnitude is a rounding error.&lt;/p&gt;
&lt;p&gt;Now the line that decides the business.&lt;/p&gt;
&lt;p&gt;Table: Gross hourly rate on a $500 monthly retainer, before platform fees, model spend and the acquisition cost of the client.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th style="text-align: right"&gt;Support hours per month&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Gross per hour&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td style="text-align: right"&gt;3&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$166.67&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td style="text-align: right"&gt;6&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$83.33&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td style="text-align: right"&gt;10&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$50.00&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;That is the whole model in three rows. A $500 retainer with three hours of
monthly firefighting is a different business from a $500 retainer with ten, and
no amount of model-price optimisation moves either row.&lt;/p&gt;
&lt;p&gt;Which is why the best idea in my original post was the one with no number
attached to it: correction burden. A system whose output the client trusts
without editing is the product. The month they start fixing every result by hand,
the fee stops making sense to them, and it stops making sense to you first,
because their hand-fixing arrives in your inbox as support hours.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Detect it:&lt;/strong&gt; log hours per client per month from the first week. Watch the
trend, not the total. Support hours rising for two consecutive months are a churn
notice with a delay on it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Design so it cannot recur:&lt;/strong&gt; put scope in writing before the build, naming what
the fee covers, what counts as a change request, and the revision limit. Add
error alerts that reach you before they reach the client. Ship a monthly results
snapshot, because renewal is decided against what the client can see, and
firefighting they never heard about is value you burned invisibly.&lt;/p&gt;
&lt;h2 id="5-article-284-makes-the-platforms-failure-yours"&gt;5. Article 28(4) makes the platform&amp;rsquo;s failure yours&lt;/h2&gt;
&lt;p&gt;Anyone building an automation that touches a client&amp;rsquo;s leads, inbox or customer
records is a processor under GDPR. The no-code platform and the model vendor are
sub-processors. That is not a metaphor, and Article 28 is short enough to read in
one sitting.&lt;/p&gt;
&lt;p&gt;The controller may use &amp;ldquo;only processors providing sufficient guarantees to
implement appropriate technical and organisational measures,&amp;rdquo; and the processing
must be governed by a binding contract that names the subject matter, duration,
nature, purpose, data types and categories of data subject, and requires
documented instructions, confidentiality, Article 32 security measures,
conditions on sub-processors, help with data-subject rights, deletion or return
of the data at the end, and information made available for audits.&lt;/p&gt;
&lt;p&gt;Two paragraphs matter most. Article 28(2): a processor &amp;ldquo;shall not engage another
processor without prior specific or general written authorisation of the
controller.&amp;rdquo; Article 28(4): where a sub-processor fails, the initial processor
&amp;ldquo;shall remain fully liable to the controller for the performance of that other
processor&amp;rsquo;s obligations.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Read that next to section 1. The platform&amp;rsquo;s position is that data security in
your generated app is your responsibility. GDPR&amp;rsquo;s position is that the platform&amp;rsquo;s
failures are your liability to your client. Both of those are true at once, and
together they turn &amp;ldquo;let the platform own the plumbing&amp;rdquo; from a productivity tip
into an assumed risk with your name on it.&lt;/p&gt;
&lt;p&gt;My post said to get explicit permission before connecting anyone&amp;rsquo;s CRM or inbox.
In the EU that is not etiquette. It is a written precondition, and naming your
sub-processors is part of it.&lt;/p&gt;
&lt;h2 id="6-article-22-does-not-bite-where-people-think-it-does"&gt;6. Article 22 does not bite where people think it does&lt;/h2&gt;
&lt;p&gt;The correction runs the other way here, before someone sells you a compliance
package you do not need.&lt;/p&gt;
&lt;p&gt;Article 22&amp;rsquo;s right not to be subject to a decision based solely on automated
processing applies only where the decision &amp;ldquo;produces legal effects concerning him
or her or similarly significantly affects him or her.&amp;rdquo; Scoring an inbound B2B
lead as warm or cold clears neither bar. Routine lead scoring is not what Article
22 is for, and saying so is more useful than the vague warning most of this market
repeats.&lt;/p&gt;
&lt;p&gt;The real boundary is where the decision does have significant effect:
eligibility, credit-like screening, pricing that materially disadvantages
someone. There, where the processing rests on contractual necessity or explicit
consent, the controller must provide the right to obtain human intervention, to
express a point of view, and to contest the decision. At that point &amp;ldquo;keep a human
in the loop for high-stakes actions&amp;rdquo; stops being a best practice and becomes an
obligation with an article number behind it.&lt;/p&gt;
&lt;h2 id="7-the-disclosure-rule-that-started-applying-four-days-ago"&gt;7. The disclosure rule that started applying four days ago&lt;/h2&gt;
&lt;p&gt;The EU AI Act&amp;rsquo;s Article 113 sets the general date of application at 2 August
2026. Article 50, the transparency article, is now live.&lt;/p&gt;
&lt;p&gt;Article 50(1) requires providers of systems that interact directly with people to
make sure those people know they are dealing with AI, &amp;ldquo;unless this is obvious
from the point of view of a natural person who is reasonably well-informed,
observant and circumspect.&amp;rdquo; Article 50(2) requires synthetic content to be marked
machine-readably, and 50(5) requires all of it &amp;ldquo;at the latest at the time of the
first interaction or exposure.&amp;rdquo; Then Article 50(4), second subparagraph:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Deployers of an AI system that generates or manipulates text which is published
with the purpose of informing the public on matters of public interest shall
disclose that the text has been artificially generated or manipulated. This
obligation shall not apply where the use is authorised by law to detect,
prevent, investigate or prosecute criminal offences or where the AI-generated
content has undergone a process of human review or editorial control and where a
natural or legal person holds editorial responsibility for the publication of
the content.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;a href="https://artificialintelligenceact.eu/article/50/" class="external-link" rel="noopener"&gt;Article 50&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;, EU Artificial Intelligence Act&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;My post described the content-system offer as one that keeps producing &amp;ldquo;with you
as the editorial backstop.&amp;rdquo; That phrase, written as sales copy, happens to
describe the exemption: human review or editorial control, with a named person
holding editorial responsibility. So stop selling it as a convenience and start
specifying it as a control. Name the editor in the contract. Keep the record of
what they reviewed.&lt;/p&gt;
&lt;p&gt;I will state the limit of that reading plainly. &amp;ldquo;Informing the public on matters
of public interest&amp;rdquo; is not defined in a way that obviously settles whether a
client&amp;rsquo;s marketing blog falls inside it, and I am not a lawyer. The reason I
would build to the exemption anyway is that it costs one named person and a
review log, and the alternative is finding out.&lt;/p&gt;
&lt;h2 id="what-i-would-keep-from-the-pitch-unchanged"&gt;What I would keep from the pitch, unchanged&lt;/h2&gt;
&lt;p&gt;Most of it survives, and the parts that survive are the parts that were never
carrying a number.&lt;/p&gt;
&lt;p&gt;The billing shapes hold. Bill monthly when the value is a system that keeps
running, bill once when you hand over an asset the client owns, and use
setup-plus-monthly when there is real build cost up front and ongoing value to
maintain. Anchor the price out loud against the labour it displaces, so the buyer
measures you against a salary rather than a software subscription. Polish belongs
to that anchor rather than to vanity: a finished client-facing surface justifies
a price a rough script will not.&lt;/p&gt;
&lt;p&gt;The build shape holds. Almost all of these systems are one skeleton: capture,
qualify and store, automate, dashboard. Build it once properly and re-skin it per
vertical. And plan before generating: have the model turn requirements into a
data model, workflows and views, read the plan, correct it with your own
judgment, then build from the corrected version. One piece of wording I had
loose, since I called it a reasoning engine. Anthropic calls Claude Code an
agentic coding tool that reads your codebase, edits files and runs commands, and
Claude the model family is a different product from Claude Code. Most Claude Code
surfaces need a subscription or a Console account, so it is not free either.&lt;/p&gt;
&lt;p&gt;Niching down holds, with the absolute removed. I wrote that &amp;ldquo;lead-gen for local
dental practices&amp;rdquo; outsells &amp;ldquo;AI automation for businesses&amp;rdquo; every time. I have no
survey and neither does anyone else selling that advice. The defensible version:
one niche gives you reusable builds and outreach language that lands, which is a
mechanism rather than a measured result.&lt;/p&gt;
&lt;p&gt;And one observation, offered as an observation. I have watched people who were
one demo away from their first client disappear into a config error for a month.
I cannot quantify that and I will not pretend otherwise.&lt;/p&gt;
&lt;h2 id="what-i-have-now-that-i-did-not-have-before"&gt;What I have now that I did not have before&lt;/h2&gt;
&lt;p&gt;Five numbers I can check, and one rule.&lt;/p&gt;
&lt;p&gt;The numbers: a CVSS score and a vendor dispute that between them define who owns
app data security. A response-time literature that measures outbound phone
dials where it names a channel at all, is silent on channel where it does not,
and nowhere measures 60 seconds. A 0.3% spam-rate ceiling with a 0.10% target and
a 5,000-per-day threshold where the authentication requirement doubles. A model bill
of roughly a dollar a month against a support-hours line that decides everything.
And a date, 2 August 2026, after which the disclosure question has an article
number attached to it.&lt;/p&gt;
&lt;p&gt;The rule is the one my own pipeline taught me at a cost of 2,922 files: any
system that runs unattended has to publish one number a human reads on a
schedule, and if you cannot name that number before launch, you have not built a
service. You have built something that will be wrong quietly for as long as
nobody looks. Mine was wrong quietly for six months, and the number that would
have caught it was the count of citations whose URLs did not resolve, which is
now the check that has to pass before anything on this site ships at all.&lt;/p&gt;</content:encoded></item><item><title>The 23-minute interruption figure is not a recovery time</title><link>https://therezaali.com/writing/interruption-recovery-time-measured/</link><guid isPermaLink="true">https://therezaali.com/writing/interruption-recovery-time-measured/</guid><pubDate>Thu, 06 Aug 2026 00:00:00 +0200</pubDate><author>contact@therezaali.com (Mo RezaAli)</author><category>essay</category><description>Published measurements run from 599 milliseconds to 25 minutes, a spread of 2,548 times, and every one of them is correct. Four groups measured four quantities and English calls all four recovery.</description><content:encoded>&lt;p&gt;The published measurements of how long it takes to get back to work after an interruption run
from 599 milliseconds to 25 minutes and 26 seconds. The top of that range is 1,526 seconds, so
dividing it by 0.599 gives a spread of about 2,548 times, and every number in it is correct. Four
research groups measured four different quantities, and English calls all four &amp;ldquo;recovery&amp;rdquo;.&lt;/p&gt;
&lt;p&gt;The most quoted figure in this subject is not among them, because it is not from a paper. &amp;ldquo;It
takes 23 minutes and 15 seconds to refocus after an interruption&amp;rdquo; originates in a
&lt;a href="https://news.gallup.com/businessjournal/23146/too-many-interruptions-work.aspx" class="external-link" rel="noopener"&gt;Gallup Business Journal&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;
interview with Gloria Mark, published 8 June 2006. What she describes it as measuring is elapsed
time until interrupted work was resumed, conditioned on the 81.9 percent of cases where it was
resumed the same day at all, and counting minutes in which the person was working productively on
roughly two other things. Her own assessment of it, in the sentence that produces it, is &amp;ldquo;which I
guess is not so long.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;The published figure from the same research programme,
&lt;a href="https://ics.uci.edu/~gmark/CHI2005.pdf" class="external-link" rel="noopener"&gt;Mark, Gonzalez and Harris at CHI 2005&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;, is 25 minutes
and 26 seconds, with a standard deviation of 54 minutes and 48 seconds. The deviation is more
than twice the mean, so the distribution is not summarised by its average, and no article I have
seen quoting 23 minutes has mentioned it. An independent
&lt;a href="https://blog.oberien.de/2023/11/05/23-minutes-15-seconds.html" class="external-link" rel="noopener"&gt;citation audit&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; checked five
commonly cited papers for the 23-minute figure and found it in none of them. I pulled the two
primary PDFs and confirm the same.&lt;/p&gt;
&lt;p&gt;Table: What four primary studies actually measured. All four get reported in popular writing as recovery time after an interruption.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Construct measured&lt;/th&gt;
 &lt;th&gt;Study&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Value&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;Extra delay before the first action on resuming a suspended goal&lt;/td&gt;
 &lt;td&gt;Monk, Trafton and Boehm-Davis, &lt;em&gt;JEP: Applied&lt;/em&gt; 2008, lab task&lt;/td&gt;
 &lt;td style="text-align: right"&gt;599 ms&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Return to the prior work rate, scored off videotape&lt;/td&gt;
 &lt;td&gt;Jackson, Dawson and Wilson, EASE 2002, 16 employees, 28 days&lt;/td&gt;
 &lt;td style="text-align: right"&gt;64 s&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Return to the application holding the primary task after an email alert&lt;/td&gt;
 &lt;td&gt;Iqbal and Horvitz, CHI 2007, 27 users, instrumented&lt;/td&gt;
 &lt;td style="text-align: right"&gt;9 min 33 s&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Elapsed clock time until the thread of work resumed, same day only&lt;/td&gt;
 &lt;td&gt;Mark, Gonzalez and Harris, CHI 2005, 24 workers, 700+ hours&lt;/td&gt;
 &lt;td style="text-align: right"&gt;25 min 26 s&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;These are not substitutes. 599 ms is a memory retrieval cost, and it is a difference rather than a
duration. Nobody in that experiment waited 599 milliseconds and then carried on. Monk, Trafton and
Boehm-Davis measured a mean resumption lag of 1,548 ms (SD 231) after an interruption against a
mean interval of 949 ms (SD 283) between actions on uninterrupted trials, and report the gap
between the two as &amp;ldquo;an estimated cost of 599 ms on the VCR programming task.&amp;rdquo; Quote either number
and you are right; quote them as the same number and you are not. 64 seconds is a behavioural
recovery. 9 minutes 33 seconds is a navigation delay, and Iqbal and Horvitz say so themselves:
their measure is application access, &amp;ldquo;not resumption of a suspended task.&amp;rdquo; 25 minutes 26 seconds
is a scheduling statistic that includes 2.26 other working spheres of real work.&lt;/p&gt;
&lt;h2 id="the-post-i-wrote-and-what-was-missing-from-it"&gt;The post I wrote, and what was missing from it&lt;/h2&gt;
&lt;p&gt;I published a post saying I had spent years treating responsiveness as productivity, that some of
my busiest days produced almost nothing that mattered, and that my most productive weeks were the
ones where I disappeared for a few hours a day with no Slack, no meetings and no email. That post
is &lt;a href="https://x.com/Mo_ali/status/2078828801933340960" class="external-link" rel="noopener"&gt;here&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;. All of it is my own first-person
recollection, unaudited, with no output metric behind it. I never measured a before or an after.
That is exactly the shape of a post-hoc success story, so the useful work is finding out which
parts of it survive contact with people who did measure.&lt;/p&gt;
&lt;p&gt;One phrase survives, conservatively. I wrote &amp;ldquo;context switches every few minutes.&amp;rdquo; The CHI 2005
field data puts an average working sphere at 11 minutes 4 seconds before a switch. The earlier
study it builds on, &lt;a href="https://ics.uci.edu/~gmark/CHI2004.pdf" class="external-link" rel="noopener"&gt;Gonzalez and Mark at CHI 2004&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;, timed
the finer unit and found 3 minutes 8 seconds of continuous work on a single event, with formal
meetings and personal time excluded from that average.&lt;/p&gt;
&lt;h2 id="interrupted-work-gets-finished-faster"&gt;Interrupted work gets finished faster&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://ics.uci.edu/~gmark/chi08-mark.pdf" class="external-link" rel="noopener"&gt;Mark, Gudith and Klocke, CHI 2008&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;, titled &amp;ldquo;The Cost
of Interrupted Work: More Speed and Stress&amp;rdquo;, is the paper most often cited for the proposition
that interruptions destroy productivity. On the productivity measure it found the opposite.
Subjects answered a fixed set of emails, interrupted every two minutes by phone or instant
message, with the time spent on the interruptions themselves subtracted out.&lt;/p&gt;
&lt;p&gt;Table: Completion time and reported stress in Mark, Gudith and Klocke, CHI 2008. Stress is a 20-point scale.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Condition&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Completion time&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Stress&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;No interruption&lt;/td&gt;
 &lt;td style="text-align: right"&gt;22.77 min&lt;/td&gt;
 &lt;td style="text-align: right"&gt;6.92&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Interrupted, same context&lt;/td&gt;
 &lt;td style="text-align: right"&gt;20.31 min&lt;/td&gt;
 &lt;td style="text-align: right"&gt;9.46&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Interrupted, different context&lt;/td&gt;
 &lt;td style="text-align: right"&gt;20.60 min&lt;/td&gt;
 &lt;td style="text-align: right"&gt;9.13&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Both interrupted conditions were significantly faster than the uninterrupted one. There was no
significant difference in errors and none in politeness. Emails fell from 31.49 words when
uninterrupted to 29.17 when interrupted, a significant drop. Frustration, time pressure and effort all rose
significantly. The authors&amp;rsquo; own reading: people &amp;ldquo;develop a mode of working faster (and writing
less) to compensate.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Two things must travel with that result or it misleads. The study ran 48 subjects, 81 percent of
them German university students with a mean age of 26, in a 1.5-hour laboratory session, not a
field study of information workers. And the authors flag the alternative they could not rule out:
the lost time may occur only in the moments right after each interruption and be masked by the
faster overall style, which would need more sensitive measurement to see.&lt;/p&gt;
&lt;p&gt;The correction still holds and it is the useful part. Fragmented time produces output, faster, at
equal accuracy, and shorter. The cost is not throughput. It is paid in stress load and in
elaboration. My post said the work that matters &amp;ldquo;doesn&amp;rsquo;t happen during the ten minutes between
two meetings,&amp;rdquo; and anyone who has cleared an inbox in ten minutes knows that is false as stated.
What is true and sourced is that what you produce in ten minutes is systematically shorter and
costs you more to produce.&lt;/p&gt;
&lt;h2 id="the-mechanism"&gt;The mechanism&lt;/h2&gt;
&lt;p&gt;Goals are not stored on a stack that you pop back to. &lt;a href="https://www.interruptions.net/literature/Altmann-CogSci02.pdf" class="external-link" rel="noopener"&gt;Altmann and Trafton&amp;rsquo;s activation-based
model&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; in &lt;em&gt;Cognitive Science&lt;/em&gt;
treats a goal as a memory item whose activation decays over time and suffers retroactive
interference from goals encoded after it, with retrieval depending on cues in the environment.
That is why 2.26 intervening working spheres cost more than 25 minutes of clock time does, and
why Mark&amp;rsquo;s informants reported real difficulty when windows or papers had been rearranged during
their absence: the retrieval cue was gone, so the pending goal had nothing to be reinstated by.&lt;/p&gt;
&lt;p&gt;The second mechanism is attention residue.
&lt;a href="https://doi.org/10.1016/j.obhdp.2009.03.002" class="external-link" rel="noopener"&gt;Leroy, in &lt;em&gt;Organizational Behavior and Human Decision Processes&lt;/em&gt;&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;,
found across two experiments that people &amp;ldquo;need to stop thinking about one task in order to fully
transition their attention and perform well on another,&amp;rdquo; and then reports the part most
productivity writing has backwards: finishing the prior task is not sufficient. Time pressure
while finishing it is what enables the disengagement. A task you completed calmly can leave more
residue than one you completed against a deadline.&lt;/p&gt;
&lt;p&gt;That answers why my busy days produced nothing, and the answer is not that I did no work. I did a
great deal of work, quickly, and it was short. The analysis sitting in drafts was the item whose
activation had decayed furthest, because it was the only one with no external cue demanding it be
reinstated.&lt;/p&gt;
&lt;h2 id="how-to-detect-this-in-your-own-system"&gt;How to detect this in your own system&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Count switches, not pings.&lt;/strong&gt; Microsoft&amp;rsquo;s June 2025
&lt;a href="https://www.microsoft.com/en-us/worklab/work-trend-index/breaking-down-infinite-workday" class="external-link" rel="noopener"&gt;Work Trend Index&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;
says employees are interrupted every two minutes during core work hours, 275 times a day, by
meetings, emails or chats. Read its methodology note before quoting either number. The two figures
are measured over different windows and Microsoft says so in consecutive sentences: &amp;ldquo;The
two-minute figure reflects the average time between pings during an eight-hour workday. The 275 is
based on the 24-hour day.&amp;rdquo; Then comes the sentence that almost never travels with the statistic:
&amp;ldquo;Based on the top 20% of users by ping volume received.&amp;rdquo; So 275 is not the average worker&amp;rsquo;s day.
It is a 24-hour count for the most-pinged fifth of users, and it is not divisible into an 8-hour
one. What is measured underneath is &amp;ldquo;a rolling 28-day sum of pings (meeting invites, emails,
chats) per unique user per workday&amp;rdquo;, which counts what arrived, not what pulled you out of what
you were doing. Your own equivalent is the count of times you changed context, which no dashboard
reports and which you have to sample by hand.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Measure your internal-to-external ratio before you buy a solution.&lt;/strong&gt; In Mark&amp;rsquo;s CHI 2005 data,
48 percent of interruptions were external and 52 percent were self-initiated, and among analysts
and developers the two ran roughly equal. Every calendar-blocking remedy, including the one I
recommended in my post, addresses the external 48 percent.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Check whether alerts actually move you.&lt;/strong&gt; Iqbal and Horvitz found only 40.8 percent of email
alerts produced a switch to the mail client within 15 seconds; the other 59.2 percent averaged 7
minutes 32 seconds first, which the authors read as self-initiated. Jackson, Dawson and Wilson,
watching 180 hours of videotape, found 70 percent of emails reacted to within 6 seconds of
arrival. People already treat most alerts as deferrable, and unreliably.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Separate interruption from context change.&lt;/strong&gt; This is the sharpest rule in the literature and I
did not have it. Mark&amp;rsquo;s informants described interruptions inside their current working sphere as
&amp;ldquo;interactions&amp;rdquo; and reported them as beneficial, one analyst describing a problem continuing to
develop in the background while not being worked on. What they called disruptive was being forced
out of the current sphere. The cost is in the context change.&lt;/p&gt;
&lt;h2 id="how-to-design-so-it-cannot-recur"&gt;How to design so it cannot recur&lt;/h2&gt;
&lt;p&gt;Preserve the retrieval cue. Altmann and Trafton&amp;rsquo;s model makes the workspace itself part of the
memory system, so leaving the file open, the terminal on the failing test and the note
half-written is the cheapest intervention available, and it costs nothing.&lt;/p&gt;
&lt;p&gt;Do not assume a blocked calendar solves it. In the same Gallup interview that produced the
23-minute figure, Mark describes getting a three-hour block, closing her office door, spending
the entire block on small accumulated tasks and touching none of the paper she had reserved it
for. Her own diagnosis: &amp;ldquo;I interrupted myself. I kept interrupting myself.&amp;rdquo; That is the honest
limit on the prescription I gave, from the researcher who measured the problem.&lt;/p&gt;
&lt;p&gt;Be careful what you claim for batching. Kushlev and Dunn ran 124 people, one week limited to
three email checks a day against one week unlimited, and found significantly lower daily stress
in the limited week, F(1,121)=4.18, p=.04, d=.37. Mark and colleagues at CHI 2016 logged 40
people over 12 workdays with heart-rate monitors and state that &amp;ldquo;despite widespread claims, we
found no evidence that batching email leads to lower stress.&amp;rdquo; I would bet on the second, and the
reason is in the details. Kushlev and Dunn&amp;rsquo;s effect appears on a single-item daily stress question
but not on the validated weekly Perceived Stress Scale, where the two conditions sit at 1.68 and
1.67 with d=.04. The sample was roughly two thirds students, there was no control condition, and
the distraction effect was significant for students and absent for community members. Mark&amp;rsquo;s team
measured physiologically, in situ, on working adults. When a daily self-report and a biosensor
disagree, the self-report is the measure that knows which condition it is in.&lt;/p&gt;
&lt;p&gt;Design the block for elaboration, not for volume. Since interrupted work comes out faster and
shorter at equal accuracy, protected time is for the work whose quality lives in its length: the
argument that needs a fourth paragraph, the analysis that needs its caveat, the problem that
needs a wrong approach tried first.&lt;/p&gt;
&lt;h2 id="what-i-have-now-that-i-did-not-have-before"&gt;What I have now that I did not have before&lt;/h2&gt;
&lt;p&gt;A rule that discriminates, in place of a rule that generalises. My post said to protect
attention, which is true and unactionable. What I run instead is a test on context change: an
interruption that keeps me inside the current working sphere is nearly free and sometimes useful,
and one that moves me out of it is the expensive event, whatever it cost in wall-clock minutes.&lt;/p&gt;
&lt;p&gt;I also have a number I will not use. When I next see 23 minutes and 15 seconds in a deck, I know
where it came from, what it measured, that the researcher who produced it called it not so long,
and that the peer-reviewed version is 25 minutes 26 seconds with a standard deviation twice its
own size. The figures that circulate most are the ones that have travelled furthest from what was
actually measured, and the provenance is worth more than the number.&lt;/p&gt;</content:encoded></item><item><title>Seven steps at 85% each is a 32% agent, and mine ran worse</title><link>https://therezaali.com/writing/per-step-reliability-in-agent-chains/</link><guid isPermaLink="true">https://therezaali.com/writing/per-step-reliability-in-agent-chains/</guid><pubDate>Thu, 06 Aug 2026 00:00:00 +0200</pubDate><author>contact@therezaali.com (Mo RezaAli)</author><category>essay</category><description>0.85 to the seventh is 0.3206, and no prompt or model upgrade changes it. My seven-node cart agent modelled at 48% end to end and ran at about one in five. The gap is the useful part.</description><content:encoded>&lt;p&gt;Seven steps at 85% each is a 32% agent. Ten steps is 20%. Those two figures are exact:
0.85⁷ = 0.3206 and 0.85¹⁰ = 0.1969. No prompt, no model upgrade and no platform changes them,
because per-step success rates in a chain multiply instead of averaging.&lt;/p&gt;
&lt;p&gt;I built a seven-node abandoned-cart agent in a visual workflow builder and it passed its demo on
one test cart. Then I pointed it at a real day: 140 abandoned carts, of which maybe 30 came out
the other end handled correctly. About one in five, by my own rough count that night, and I want
to be precise about how imprecise that is. I was reading a CRM, not a log table. It is my own
unaudited figure from my own work, published
&lt;a href="https://x.com/Mo_ali/status/2078615536272015854" class="external-link" rel="noopener"&gt;here&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; before this site existed. There is no
execution log to link, so this one rests on my account of it.&lt;/p&gt;
&lt;p&gt;One customer who abandoned a $12 phone case got a &amp;ldquo;complete your premium order&amp;rdquo; SMS carrying a
15% discount code I never meant to issue to anyone. That class of send ran for about two weeks
before I caught the pattern. Call it a couple hundred dollars in discount codes handed to people
who were not buying. Every node in that chain reported success.&lt;/p&gt;
&lt;p&gt;I blamed the model first and swapped to a bigger one. Same mess, slightly more expensive. The
hypothesis was wrong, and the arithmetic below says by how much.&lt;/p&gt;
&lt;h2 id="the-arithmetic-and-the-assumption-it-rests-on"&gt;The arithmetic, and the assumption it rests on&lt;/h2&gt;
&lt;p&gt;The model here is not new and it is not mine. It is the series-system reliability formula, and
the &lt;a href="https://www.itl.nist.gov/div898/handbook/apr/section1/apr182.htm" class="external-link" rel="noopener"&gt;NIST/SEMATECH e-Handbook&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;
states it in section 8.1.8.2 as R_S(t) = ∏ R_i(t): multiply the reliabilities. The handbook also
states the condition that makes it valid, and this turns out to be the whole story of my agent:
&amp;ldquo;Each component operates or fails independently of every other one, at least until the first
component failure occurs.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Hold onto that condition. It is the one my agent broke.&lt;/p&gt;
&lt;h2 id="the-worksheet-and-what-those-numbers-actually-are"&gt;The worksheet, and what those numbers actually are&lt;/h2&gt;
&lt;p&gt;Here is the flow with a rate against each node. Read this before the table, because it changes
what the table is: these are not logged measurements. I did not instrument per-node success at
the time. They are my estimates, assigned after watching the thing fail, and by my own admission
in the original post I was being generous on a couple of them.&lt;/p&gt;
&lt;p&gt;Table: My seven-node abandoned-cart flow with retrospectively estimated per-node success rates, and the running product. The rates are my estimates rather than measurements; the running total is arithmetic on those estimates.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th style="text-align: right"&gt;Step&lt;/th&gt;
 &lt;th&gt;Node&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Estimated per-step success&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Running product&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td style="text-align: right"&gt;1&lt;/td&gt;
 &lt;td&gt;Read cart contents&lt;/td&gt;
 &lt;td style="text-align: right"&gt;0.97&lt;/td&gt;
 &lt;td style="text-align: right"&gt;0.9700&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td style="text-align: right"&gt;2&lt;/td&gt;
 &lt;td&gt;Classify buyer intent&lt;/td&gt;
 &lt;td style="text-align: right"&gt;0.78&lt;/td&gt;
 &lt;td style="text-align: right"&gt;0.7566&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td style="text-align: right"&gt;3&lt;/td&gt;
 &lt;td&gt;Pick the right offer&lt;/td&gt;
 &lt;td style="text-align: right"&gt;0.88&lt;/td&gt;
 &lt;td style="text-align: right"&gt;0.6658&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td style="text-align: right"&gt;4&lt;/td&gt;
 &lt;td&gt;Draft the email&lt;/td&gt;
 &lt;td style="text-align: right"&gt;0.92&lt;/td&gt;
 &lt;td style="text-align: right"&gt;0.6125&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td style="text-align: right"&gt;5&lt;/td&gt;
 &lt;td&gt;Draft the SMS&lt;/td&gt;
 &lt;td style="text-align: right"&gt;0.92&lt;/td&gt;
 &lt;td style="text-align: right"&gt;0.5635&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td style="text-align: right"&gt;6&lt;/td&gt;
 &lt;td&gt;Schedule the send&lt;/td&gt;
 &lt;td style="text-align: right"&gt;0.95&lt;/td&gt;
 &lt;td style="text-align: right"&gt;0.5354&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td style="text-align: right"&gt;7&lt;/td&gt;
 &lt;td&gt;Log result to CRM&lt;/td&gt;
 &lt;td style="text-align: right"&gt;0.90&lt;/td&gt;
 &lt;td style="text-align: right"&gt;0.4818&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Modelled end to end: 48%. For a flat rate applied to any chain length, the curve looks like this.&lt;/p&gt;
&lt;p&gt;Table: End-to-end reliability of a chain of independent steps, by chain length and uniform per-step success rate.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th style="text-align: right"&gt;Steps&lt;/th&gt;
 &lt;th style="text-align: right"&gt;@95%&lt;/th&gt;
 &lt;th style="text-align: right"&gt;@90%&lt;/th&gt;
 &lt;th style="text-align: right"&gt;@85%&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td style="text-align: right"&gt;3&lt;/td&gt;
 &lt;td style="text-align: right"&gt;86%&lt;/td&gt;
 &lt;td style="text-align: right"&gt;73%&lt;/td&gt;
 &lt;td style="text-align: right"&gt;61%&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td style="text-align: right"&gt;5&lt;/td&gt;
 &lt;td style="text-align: right"&gt;77%&lt;/td&gt;
 &lt;td style="text-align: right"&gt;59%&lt;/td&gt;
 &lt;td style="text-align: right"&gt;44%&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td style="text-align: right"&gt;7&lt;/td&gt;
 &lt;td style="text-align: right"&gt;70%&lt;/td&gt;
 &lt;td style="text-align: right"&gt;48%&lt;/td&gt;
 &lt;td style="text-align: right"&gt;32%&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td style="text-align: right"&gt;10&lt;/td&gt;
 &lt;td style="text-align: right"&gt;60%&lt;/td&gt;
 &lt;td style="text-align: right"&gt;35%&lt;/td&gt;
 &lt;td style="text-align: right"&gt;20%&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id="the-model-under-predicted-the-failure-and-that-is-the-finding"&gt;The model under-predicted the failure, and that is the finding&lt;/h2&gt;
&lt;p&gt;My original post said the ten-step figure of 20% matched my night at the CRM almost exactly. It
does not, and the mismatch is worth more than the match would have been. My agent had seven steps,
not ten. My two models of it predict 32% (flat 0.85) and 48% (the worksheet). What I published was
one in five. Both models were optimistic, one of them by nearly thirty points.&lt;/p&gt;
&lt;p&gt;Two derivations pin the gap. Both are this article&amp;rsquo;s arithmetic on the one rate I published, not a
second measurement, and they inherit its looseness: I counted &amp;ldquo;maybe 30&amp;rdquo; out of 140 off a CRM
screen and called it one in five, so treat 0.20 as a round number and not a reading. To land at 20%
with a uniform 0.85 per step you need ln(0.20) / ln(0.85) = 9.9 steps, so the 20% figure describes
a chain I did not build. And a seven-step chain finishing at 20% implies a uniform per-step rate of
0.20^(1/7) = 0.795, not 0.85.&lt;/p&gt;
&lt;p&gt;Three mechanisms account for the difference, and each one generalises.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;A node on the canvas is not a step in the chain.&lt;/strong&gt; &amp;ldquo;Classify buyer intent&amp;rdquo; is one box and at
least three failure surfaces: the model call, the parse of whatever came back, and the schema or
enum check that decides the label is usable. Count failure surfaces, not boxes, and my seven-box
flow was closer to ten steps than seven. The multiplication was right. My count of what to
multiply was wrong.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The independence assumption does not hold.&lt;/strong&gt; NIST is explicit that the product formula needs
each component to fail independently of every other one. A cart record with a malformed line item
does not fail one node. It degrades the read, poisons the classification and misprices the SMS,
all from one cause. Correlated failure makes the real distribution more bimodal than the product
of independent rates suggests: batches come out mostly fine or mostly wrecked rather than
uniformly degraded. The product is a fair planning tool and a poor description of the variance.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Estimates made by the person who built the thing run warm.&lt;/strong&gt; The 0.795 above sits five and a half
points under the 0.85 flat default, and about ten points under the 0.901 geometric mean of the
rates I actually wrote in the worksheet (0.4818^(1/7)). The generic default was optimistic. The
rates I picked node by node, after watching the thing fail, were about twice as optimistic as the
generic default I had not bothered to think about.&lt;/p&gt;
&lt;p&gt;Google&amp;rsquo;s SRE team published the sharpest correction to naive chain multiplication I have found, in
&lt;a href="https://sre.google/static/pdf/calculus_of.pdf" class="external-link" rel="noopener"&gt;The Calculus of Service Availability&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; (ACM Queue,
March/April 2017). Their starting rule is mine: &amp;ldquo;A service cannot be more available than the
intersection of all its critical dependencies.&amp;rdquo; Then they call the inference that each extra link
needs another 9 &amp;ldquo;incorrect&amp;rdquo;, because a dependency appearing at several points must be counted
once. Their rule: &amp;ldquo;If a service has N unique critical dependencies, then each one contributes 1/N
to the dependency-induced unavailability of the top-level service, regardless of its depth in the
dependency hierarchy.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Applied to my flow, my classify-intent node is not step two of seven. It is a critical dependency
of steps three through seven, which is why one wrong label produced five confident successes
carrying a wrong answer. Count unique critical dependencies, not boxes.&lt;/p&gt;
&lt;h2 id="why-that-node-and-not-another"&gt;Why that node and not another&lt;/h2&gt;
&lt;p&gt;The rates in the worksheet are uneven and the split is not random. Reading a cart is 0.97 and
hitting a scheduler is 0.95 because both are plumbing: deterministic, checkable, the same every
time. Classifying why someone abandoned is 0.78 because it is a judgment call made in one pass
with no way to ask a follow-up question. Price shock. Just browsing. Payment failed. Distracted.
Comparison shopping. Accidental add. My agent would read a cart holding one cheap item, decide
&amp;ldquo;high purchase intent, price-sensitive&amp;rdquo; and fire a discount at someone who was never buying, who
then learned that abandoning carts prints coupons. That is where it snapped in every batch I
looked at, and I have not seen a classification node of that kind sitting at 95%.&lt;/p&gt;
&lt;h2 id="why-it-fails-green"&gt;Why it fails green&lt;/h2&gt;
&lt;p&gt;The worst property of this failure is not the rate. It is that nothing turns red.&lt;/p&gt;
&lt;p&gt;That is a documented property of the tooling, not a quirk of my build. n8n&amp;rsquo;s
&lt;a href="https://docs.n8n.io/build/flow-logic/handle-errors-gracefully" class="external-link" rel="noopener"&gt;error handling documentation&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; says the error
workflow &amp;ldquo;runs if an execution fails&amp;rdquo;, and gives the usual causes as &amp;ldquo;errors in node settings, or
the workflow running out of memory&amp;rdquo;. The only way to make a workflow fail deliberately is the Stop
And Error node, which the same page documents as a way &amp;ldquo;to force executions to fail under your
chosen circumstances&amp;rdquo;. That is a statement about the world model the platform holds: an execution
either errors or it does not, and semantic correctness is not an observable. A node returning
&amp;ldquo;high purchase intent, price-sensitive&amp;rdquo; for someone who was browsing has succeeded by every
definition the runtime has.&lt;/p&gt;
&lt;p&gt;One layer down, the reason the model produces a confident wrong label rather than declining is
also documented. &lt;a href="https://arxiv.org/abs/2509.04664" class="external-link" rel="noopener"&gt;Why Language Models Hallucinate&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; (Kalai,
Nachum, Vempala and Zhang, September 2025) argues that models hallucinate &amp;ldquo;because the training
and evaluation procedures reward guessing over acknowledging uncertainty&amp;rdquo;, and that they
&amp;ldquo;originate simply as errors in binary classification&amp;rdquo;. That is why a bigger model does not fix it.
A bigger model is optimised by the same scoring that pays for a confident guess and charges
nothing extra for a wrong one.&lt;/p&gt;
&lt;p&gt;The failure category also has a name.
&lt;a href="https://arxiv.org/abs/2503.13657" class="external-link" rel="noopener"&gt;Why Do Multi-Agent LLM Systems Fail?&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; (Cemri, Pan, Yang et al.,
Berkeley, revised October 2025) built a taxonomy from 1600+ annotated traces across 7 frameworks,
developed from 150 traces under expert annotation at an inter-annotator agreement of kappa = 0.88.
It lands on 14 failure modes in three categories, and the third is task verification. Their
conclusion is that these failures &amp;ldquo;require more sophisticated solutions&amp;rdquo;, not better models.&lt;/p&gt;
&lt;h2 id="why-you-cannot-out-model-a-long-chain"&gt;Why you cannot out-model a long chain&lt;/h2&gt;
&lt;p&gt;My original post got this argument right and the arithmetic wrong, and correcting it makes the
point harder. I wrote that a better model might move a node from 0.85 to 0.90, and that across ten
steps this takes you from 20% to 35%. It does not. Upgrading &lt;strong&gt;one&lt;/strong&gt; node in a ten-step chain gives
0.85⁹ × 0.90 = 0.2085. Twenty percent to twenty-one percent. The 35% figure is 0.9¹⁰ = 0.3487,
which requires every node in the chain to improve. One better node buys you one point.&lt;/p&gt;
&lt;p&gt;The published evidence points the same way. METR&amp;rsquo;s
&lt;a href="https://arxiv.org/abs/2503.14499" class="external-link" rel="noopener"&gt;Measuring AI Ability to Complete Long Software Tasks&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;
(arXiv 2503.14499, v4 revised 10 July 2026) puts frontier models at a 50%-task-completion time
horizon &amp;ldquo;of around 50 minutes&amp;rdquo;, doubling &amp;ldquo;approximately every seven months since 2019&amp;rdquo;, and
attributes the increase &amp;ldquo;primarily&amp;rdquo; to &amp;ldquo;greater reliability and ability to adapt to mistakes&amp;rdquo;.
Length is the binding constraint and reliability is what moves it. Those are software engineering
tasks with expert human baselines, so take the direction, not a number for a marketing flow.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://arxiv.org/abs/2406.12045" class="external-link" rel="noopener"&gt;tau-bench&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; (Yao, Shinn, Razavi and Narasimhan, June 2024)
introduced pass^k, the probability that all k independent trials of the same task succeed, and
reports that &amp;ldquo;even state-of-the-art function calling agents (like gpt-4o) succeed on &amp;lt;50% of the
tasks, and are quite inconsistent (pass^8 &amp;lt;25% in retail)&amp;rdquo;. That is my multiplication run across
repeated trials instead of across steps, by researchers, with a metric that has a name.&lt;/p&gt;
&lt;p&gt;It is also the honest limit on my planning number. I use 85% as the default for a fuzzy step and
94% or better for a deterministic one, and that is an assumption rather than a citation. tau-bench
measures whole episodes, so it supports &amp;ldquo;chains degrade fast&amp;rdquo; and does not establish that any given
LLM step sits at 85%. I could not find a public benchmark reporting per-step success for a node
like &amp;ldquo;classify why this person abandoned a cart&amp;rdquo;. Instrument your own nodes and use your numbers
over my default.&lt;/p&gt;
&lt;h2 id="retry-and-why-it-does-not-save-you"&gt;Retry, and why it does not save you&lt;/h2&gt;
&lt;p&gt;The obvious objection from anyone running these flows is that the platform retries. n8n documents
per-node Retry on Fail, Max Tries and Wait Between Tries (ms) in node
&lt;a href="https://docs.n8n.io/integrations/builtin/core-nodes/n8n-nodes-base.httprequest/common-issues" class="external-link" rel="noopener"&gt;Settings&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;.
That page describes the mechanism and states no default value for either number, so the retry
budget is a choice you make rather than one the page makes for you.&lt;/p&gt;
&lt;p&gt;Retry then partitions the chain along the same line everything else in this piece does. For an
independent transient failure it is close to free reliability: a plumbing step at p = 0.95 with
three tries gives 1 − (1 − 0.95)³ = 0.999875. For a judgment step the error is largely
deterministic given the same input and the same prompt, so retrying returns the same wrong label
with the same confidence and the effective rate stays at 0.78. Retries fix plumbing and do nothing
for judgment.&lt;/p&gt;
&lt;p&gt;The one technique I know of that raises a judgment node without a human and without a bigger model
is sampling the step several times and taking the majority answer.
&lt;a href="https://arxiv.org/abs/2203.11171" class="external-link" rel="noopener"&gt;Self-Consistency&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; (Wang et al., 2022, revised 2023) reports
+17.9% on GSM8K and +6.4% on StrategyQA doing exactly that. Those are 2022 reasoning benchmarks,
not intent classification, so it is a mechanism to test rather than a promised lift, and n samples
costs n times the tokens for that step.&lt;/p&gt;
&lt;h2 id="how-to-detect-this-in-your-own-flow"&gt;How to detect this in your own flow&lt;/h2&gt;
&lt;p&gt;The test takes ten minutes and needs no tooling.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;List every failure surface, not every box. A node that calls a model, parses its output and
validates a schema is three rows.&lt;/li&gt;
&lt;li&gt;Assign each row a success rate you would bet your own money on, not the rate from your best
demo run. Mark each row JUDGMENT or PLUMBING.&lt;/li&gt;
&lt;li&gt;Multiply down the column. That product is your optimistic ceiling, not your expected result,
for the three reasons above.&lt;/li&gt;
&lt;li&gt;For each row, ask whether a wrong output fails loud or fails green. Every green-failing row is
invisible to your platform&amp;rsquo;s error handling and needs a gate you wrote yourself.&lt;/li&gt;
&lt;li&gt;Mark which rows are critical dependencies of downstream rows. Those are where one error
multiplies into several confident successes.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;For the empirical version rather than the estimate, log the input and output of every node for one
real batch and score a sample by hand. That is what I did not do before shipping, and it is the
difference between a worksheet and a measurement.&lt;/p&gt;
&lt;h2 id="design-so-it-cannot-recur"&gt;Design so it cannot recur&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Delete steps before you optimise them.&lt;/strong&gt; Cutting my flow from seven steps to five, by removing
the standalone SMS-draft node and folding CRM logging into the send step, takes the modelled
0.97 × 0.78 × 0.88 × 0.92 × 0.95 to 0.5819. Forty-eight percent to fifty-eight, holding the same
per-step rates. But the two moves are not equivalent and my original post glossed over it. Deleting
work is an unconditional gain. Merging work is conditional: the send step now does two jobs, and
holding its rate at 0.95 is an assumption I have not tested. Merging only pays if the merged node
does not absorb the failure rate you thought you removed.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Put the human on the weakest number and nowhere else.&lt;/strong&gt; One checkpoint, at classify-intent. The
agent drafts everything and parks the batch. Each morning I spend about eleven minutes scanning
intent labels and the offers attached to them, approve, fix the handful that are wrong, and
release. The whole flow by hand used to take two hours. Both are my own unaudited figures from my
own work, in the same post linked above. Automation moved the judgment. It did not remove it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Gate what a rule can check.&lt;/strong&gt; Two of mine: an offer that is a discount over 20% on an order under
$20 gets blocked and flagged, and an SMS carrying a price that does not match the cart gets
blocked. These are if-statements. They make no model call and cost no tokens per run.&lt;/p&gt;
&lt;p&gt;The arithmetic of a gate is worth writing down, because my original post asserted the effect
without the formula. For a node with success rate p and a gate catching a fraction d of its
errors, effective correctness is:&lt;/p&gt;
&lt;pre tabindex="0"&gt;&lt;code&gt;p&amp;#39; = p + (1 - p) * d
&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;On my weakest node, p = 0.78: a gate catching 70% of errors gives 0.934, one catching 90% gives
0.978. Two conditions travel with that formula or it lies to you. It only holds if caught errors
are then corrected; if they are merely blocked and dropped, precision rises, throughput falls and
end-to-end success does not improve at all. Mine are corrected, by me, in the eleven minutes. And
my two rules are value-range checks that catch the loud subset of a silent failure class. Neither
catches &amp;ldquo;wrong intent label, plausible offer, correct price&amp;rdquo;, which is the failure that cost me
the fortnight.&lt;/p&gt;
&lt;p&gt;The demo is one trip down the chain on a clean input. The system is the same chain on real volume,
where the weakest node decides the outcome and reports success the whole way down. Before the next
one ships, run the column. If the number scares you, that is the column doing its job.&lt;/p&gt;</content:encoded></item><item><title>Kimi K3’s checkpoint is 1,560.9 GB, which is 121 GB more than a B200 node holds</title><link>https://therezaali.com/writing/kimi-k3-open-weight-release/</link><guid isPermaLink="true">https://therezaali.com/writing/kimi-k3-open-weight-release/</guid><pubDate>Thu, 06 Aug 2026 00:00:00 +0200</pubDate><author>contact@therezaali.com (Mo RezaAli)</author><category>essay</category><description>The model card says 2.8 trillion parameters at four bits, which multiplies out to 1,400 GB. The uploaded weights are 1,560.9 GB, because a 4-bit release is not uniformly four bits.</description><content:encoded>&lt;p&gt;Kimi K3&amp;rsquo;s weights ship natively as MXFP4. Four bits per parameter, half a byte,
quantization-aware trained rather than squeezed after the fact. Multiply that by the 2.8 trillion
parameters on the &lt;a href="https://huggingface.co/moonshotai/Kimi-K3" class="external-link" rel="noopener"&gt;model card&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; and you get 1,400 GB.&lt;/p&gt;
&lt;p&gt;The checkpoint Moonshot actually uploaded is &lt;strong&gt;1,560.9 GB&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;The 96 safetensors shards in the repository&amp;rsquo;s
&lt;a href="https://huggingface.co/moonshotai/Kimi-K3/tree/main" class="external-link" rel="noopener"&gt;file listing&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; sum to 1,560,936,091,448 bytes:
1,560.9 GB, or &lt;strong&gt;1,453.7 GiB&lt;/strong&gt;. Hugging Face&amp;rsquo;s dtype census for the same repository lands 75.8 MB
below that, which is the safetensors JSON headers. Both counts are 161 GB above the multiplication,
and the reason is one field in
&lt;a href="https://huggingface.co/moonshotai/Kimi-K3/raw/main/config.json" class="external-link" rel="noopener"&gt;config.json&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="a-4-bit-release-is-not-uniformly-four-bits"&gt;A 4-bit release is not uniformly four bits&lt;/h2&gt;
&lt;p&gt;&lt;code&gt;quantization_config&lt;/code&gt; targets &lt;code&gt;Linear&lt;/code&gt; layers at &lt;code&gt;num_bits: 4&lt;/code&gt;, and then exempts six patterns:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-json" data-lang="json"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s2"&gt;&amp;#34;ignore&amp;#34;&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="s2"&gt;&amp;#34;re:.*self_attn.*&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="s2"&gt;&amp;#34;re:.*shared_experts.*&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="s2"&gt;&amp;#34;re:.*mlp\\.(gate|up|gate_up|down)_proj.*&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="s2"&gt;&amp;#34;re:.*lm_head.*&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="s2"&gt;&amp;#34;re:.*vision_tower.*&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="s2"&gt;&amp;#34;re:.*mm_projector.*&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Attention, the shared experts, the dense MLP projections, the output head, the vision tower and the
multimodal projector never get quantised. They stay BF16. That is 57.18 billion parameters, 2.06% of
the model, and on their own they cost 114.4 GB. The parts left in higher precision are exactly the
parts that would break first if you crushed them: the router-independent path every token takes.&lt;/p&gt;
&lt;p&gt;The second missing term is the format&amp;rsquo;s own bookkeeping. &lt;code&gt;group_size&lt;/code&gt; is 32 and &lt;code&gt;scale_dtype&lt;/code&gt; is
&lt;code&gt;torch.uint8&lt;/code&gt;, so every 32 quantised values carry a one-byte shared scale. Across 2.72 trillion
quantised parameters that is another 85.1 GB of scales that are not weights and still occupy memory.&lt;/p&gt;
&lt;p&gt;Table: Kimi K3&amp;rsquo;s published checkpoint by dtype, from Hugging Face&amp;rsquo;s parameter census for moonshotai/Kimi-K3 and the group size in config.json, accessed 6 August 2026.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Component&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Parameters&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Bytes&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;MXFP4 weights, four bits packed&lt;/td&gt;
 &lt;td style="text-align: right"&gt;2,722,740,830,208&lt;/td&gt;
 &lt;td style="text-align: right"&gt;1,361.37 GB&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;MXFP4 group scales, one byte per 32 values&lt;/td&gt;
 &lt;td style="text-align: right"&gt;&lt;/td&gt;
 &lt;td style="text-align: right"&gt;85.09 GB&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;BF16 tensors (attention, shared experts, dense MLP, output head, vision tower)&lt;/td&gt;
 &lt;td style="text-align: right"&gt;57,179,884,544&lt;/td&gt;
 &lt;td style="text-align: right"&gt;114.36 GB&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;F32 tensors&lt;/td&gt;
 &lt;td style="text-align: right"&gt;11,122,432&lt;/td&gt;
 &lt;td style="text-align: right"&gt;0.04 GB&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;&lt;strong&gt;Total&lt;/strong&gt;&lt;/td&gt;
 &lt;td style="text-align: right"&gt;&lt;strong&gt;2,779,931,837,184&lt;/strong&gt;&lt;/td&gt;
 &lt;td style="text-align: right"&gt;&lt;strong&gt;1,560.86 GB&lt;/strong&gt;&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Half a byte each against the exact parameter count gives 1,390.0 GB. The real file is 170.9 GB
larger, and that overshoot splits almost evenly in two: 85.1 GB of group scales, 85.8 GB of extra
precision on the 2% that was never quantised. The checkpoint averages &lt;strong&gt;4.49 bits per parameter&lt;/strong&gt;,
not four.&lt;/p&gt;
&lt;h2 id="subtract"&gt;Subtract&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://www.nvidia.com/en-us/data-center/dgx-b200/" class="external-link" rel="noopener"&gt;NVIDIA&amp;rsquo;s DGX B200 page&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; states &amp;ldquo;8x NVIDIA
Blackwell GPUs&amp;rdquo; and &amp;ldquo;GPU Memory 1,440 GB total, 64 TB/s HBM3e bandwidth&amp;rdquo;. It does not print a
per-accelerator figure; 180 GB is 1,440 divided by eight, and that division is mine, not NVIDIA&amp;rsquo;s.&lt;/p&gt;
&lt;pre tabindex="0"&gt;&lt;code&gt;1,560.9 GB Kimi K3 checkpoint (96 safetensors shards)
1,440.0 GB node memory (8 x B200)
----------
 -120.9 GB short
&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;The weights do not fit. Not tightly, not with tuning: an eight-GPU B200 node is 121 GB short of
holding the checkpoint at rest, before one byte of KV cache, activations, the 401M-parameter
MoonViT-V2 vision encoder or CUDA graphs. Which is what vLLM&amp;rsquo;s
&lt;a href="https://vllm.ai/blog/2026-07-27-k3" class="external-link" rel="noopener"&gt;day-0 post&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; says in plain words: &amp;ldquo;The entire model can barely
fit in a single NVIDIA DGX B300 and requires a minimum of 16 NVIDIA B200/GB200 GPUs to serve on that
hardware generation.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The first version of this page said 1,400 GB and put 40 GB of headroom on that node.&lt;/strong&gt; It reached
that by multiplying the headline parameter count by four bits, which is the shortcut the model card
invites and the checkpoint refuses. The error was 161 GB, it inverted the conclusion from a tight
fit to an impossible one, and it made vLLM&amp;rsquo;s sixteen-GPU floor look like caution rather than
arithmetic. The file listing is public and the shard sizes are on it. Add up the 96 numbers
yourself. That is the check, and it is the one I should have run before I ran the multiplication.&lt;/p&gt;
&lt;h2 id="what-i-published-on-17-july-and-what-the-sources-say"&gt;What I published on 17 July, and what the sources say&lt;/h2&gt;
&lt;p&gt;I posted &lt;a href="https://x.com/Mo_ali/status/2078169849398714516" class="external-link" rel="noopener"&gt;a summary of K3&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; on 17 July 2026, about
22 hours after Moonshot&amp;rsquo;s launch thread. Four of its claims do not survive the primary sources.&lt;/p&gt;
&lt;p&gt;It said K3 has &amp;ldquo;faster inference&amp;rdquo;. It does not. It said Moonshot &amp;ldquo;introduced two new architectural
improvements&amp;rdquo;. Both were published months earlier, with their own papers. It said the open weights
were coming &amp;ldquo;later this month&amp;rdquo;, which was true for ten more days and is now simply out of date. And
it said K3 &amp;ldquo;consistently ranks among the strongest models&amp;rdquo; rather than &amp;ldquo;dominating a single
category&amp;rdquo;, which is the exact inverse of what happened.&lt;/p&gt;
&lt;p&gt;Those are four instances of one mechanism. A launch thread is a compression of a technical report,
written by the vendor, optimised for a specific effect. Summarising it inherits every framing
decision while dropping every qualifier underneath, because the qualifiers live in the papers, the
licence file and the config JSON. So I went and read those instead.&lt;/p&gt;
&lt;h2 id="kda-and-attnres-were-not-new-and-the-real-claim-is-better"&gt;KDA and AttnRes were not new, and the real claim is better&lt;/h2&gt;
&lt;p&gt;Kimi Delta Attention was introduced in
&lt;a href="https://arxiv.org/abs/2510.26692" class="external-link" rel="noopener"&gt;Kimi Linear&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; (arXiv 2510.26692, submitted 30 October 2025) as
&amp;ldquo;an expressive linear attention module that extends Gated DeltaNet with a finer-grained gating
mechanism&amp;rdquo;. Attention Residuals got its own paper,
&lt;a href="https://arxiv.org/abs/2603.15031" class="external-link" rel="noopener"&gt;Attention Residuals&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; (arXiv 2603.15031, Kimi Team, led by
Guangyu Chen), submitted 16 March 2026. The weights shipped on 27 July 2026, which is 270 days after
the first paper and 133 days after the second.&lt;/p&gt;
&lt;p&gt;They also do different jobs, and welding them into one sentence about long context gets one of them
wrong. KDA is the context mechanism: linear-attention layers holding a fixed-size recurrent state,
interspersed with periodic full-attention layers so global recall survives. AttnRes is a &lt;strong&gt;depth&lt;/strong&gt;
mechanism. It &amp;ldquo;replaces fixed accumulation with softmax attention over preceding layer outputs&amp;rdquo;,
because uniform residual aggregation causes &amp;ldquo;uncontrolled hidden-state growth with depth,
progressively diluting each layer&amp;rsquo;s contribution&amp;rdquo;. Nothing to do with context length.&lt;/p&gt;
&lt;p&gt;The Kimi Linear paper reports up to 6x decoding throughput at 1M context and a 75% KV cache
reduction. If you see either quoted as a K3 result, it is not one. Both were measured on a
48B-total, 3B-activated model in October 2025.&lt;/p&gt;
&lt;p&gt;Which is what makes K3 the more interesting story once you stop calling it an invention. It is the
scaling event for a recipe that was already validated small. Kimi Linear ran a 3:1 hybrid, one full
attention layer for every three linear ones, 25.0%. K3&amp;rsquo;s config is 93 layers: 69 KDA plus 24 Gated
MLA, so &lt;strong&gt;25.8%&lt;/strong&gt;, at roughly 58 times the parameter count. The cadence is not uniform, and the
config is worth reading rather than averaging: &lt;code&gt;full_attn_layers&lt;/code&gt; is &lt;code&gt;[4, 8, ... 88, 92, 93]&lt;/code&gt;, every
fourth layer plus a second full-attention layer stacked directly on top of the first at 92 and 93.&lt;/p&gt;
&lt;p&gt;That ratio transferring across a 58x scale-up is my own reading of two config files, and I want to
be careful not to hang Moonshot&amp;rsquo;s headline number on it. The
&lt;a href="https://arxiv.org/abs/2607.24653" class="external-link" rel="noopener"&gt;technical report&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; claims &amp;ldquo;an approximately 2.5x improvement in
overall scaling efficiency&amp;rdquo; over Kimi K2, and its abstract credits that to a bundle, not to the
layer ratio: KDA and AttnRes, &amp;ldquo;together with Stable LatentMoE, which effectively activates 16 of 896
routed experts per token, and refined training and data recipes&amp;rdquo;.&lt;/p&gt;
&lt;p&gt;While you are in the config, note that 1.8% and 3.7% are both true sparsity figures for this model
and they measure different things.&lt;/p&gt;
&lt;p&gt;Table: Two distinct sparsity quantities for Kimi K3, derived from the Hugging Face model card, accessed 6 August 2026.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Quantity&lt;/th&gt;
 &lt;th&gt;Arithmetic&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Value&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;Routed experts selected per token&lt;/td&gt;
 &lt;td&gt;16 / 896&lt;/td&gt;
 &lt;td style="text-align: right"&gt;1.7857%&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Routed plus shared experts&lt;/td&gt;
 &lt;td&gt;18 / 898&lt;/td&gt;
 &lt;td style="text-align: right"&gt;2.0045%&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Activated parameters as share of total&lt;/td&gt;
 &lt;td&gt;104B / 2,800B&lt;/td&gt;
 &lt;td style="text-align: right"&gt;3.714%&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Expert-count sparsity is about half the parameter-activation fraction, because the shared experts,
all attention parameters, the embeddings and the vision encoder are active on every token no matter
how the router votes. Those are the same tensors the quantiser skipped. If you quote a sparsity
figure, say which one it is.&lt;/p&gt;
&lt;h2 id="the-speed-claim-is-the-one-that-is-simply-false"&gt;The speed claim is the one that is simply false&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://artificialanalysis.ai/models/kimi-k3" class="external-link" rel="noopener"&gt;Artificial Analysis&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;, accessed 6 August 2026, measures
K3 (max) at &lt;strong&gt;38.7 output tokens per second, #54 of the 101 models on that page&lt;/strong&gt;, against a median
of 62.8 for open-weight models of similar size, with a time to first token of 3.11 seconds against a
1.82 second median. Their own summary calls it &amp;ldquo;particularly expensive when comparing to other open
weight models of similar size&amp;rdquo; and &amp;ldquo;notably slow and very verbose&amp;rdquo;. On price it sits at #96 of 101.
Moonshot&amp;rsquo;s &lt;a href="https://platform.kimi.ai/docs/pricing/chat-k3" class="external-link" rel="noopener"&gt;pricing page&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; lists &lt;strong&gt;$3.00 per million
input tokens, $0.30 cached, $15.00 output&lt;/strong&gt;, flat, with no context-length tiering.&lt;/p&gt;
&lt;p&gt;vLLM reports 111 to 118 tokens per second single-user, rising to 331 to 370 with DSpark speculative
decoding. Those are self-hosted, single-user, at batch size 1 on GB300 NVL72, and not comparable to a
third party measuring the hosted API under production load. Do not put them in the same column.&lt;/p&gt;
&lt;p&gt;A 2.8 trillion parameter model being slow was predictable from its own headline number before
anyone measured anything, which is exactly why writing &amp;ldquo;faster inference&amp;rdquo; in the first ten lines
was avoidable.&lt;/p&gt;
&lt;h2 id="it-dominated-one-board-and-sits-thirteenth-on-another"&gt;It dominated one board and sits thirteenth on another&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://x.com/arena/status/2077824029126504525" class="external-link" rel="noopener"&gt;LMArena&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; put K3 at &lt;strong&gt;#1 in the Frontend Code Arena
with 1679 points, past Claude Fable 5, a 17-place jump from Kimi K2.6 at #18&lt;/strong&gt;, and #1 in six of
seven frontend domains, second in Gaming. On LMArena&amp;rsquo;s
&lt;a href="https://lmarena.ai/leaderboard/text" class="external-link" rel="noopener"&gt;main text leaderboard&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;, read on 6 August 2026, &lt;code&gt;kimi-k3-max&lt;/code&gt;
sits at &lt;strong&gt;rank 13 with a score of 1485 (±10)&lt;/strong&gt;, on 3,556 votes, flagged Preliminary, with a rank
spread of 4 to 31. Same model, same organisation, two boards, two verdicts, and the second one comes
with an uncertainty band 27 places wide that the first one never shows you.&lt;/p&gt;
&lt;p&gt;Moonshot is more modest about this than the coverage is. Its own technical report abstract says K3
&amp;ldquo;achieves frontier-level performance across long-horizon coding, agentic, knowledge, reasoning, and
vision tasks&amp;rdquo; and then, in the same paragraph, that it &amp;ldquo;trails the most powerful proprietary models,
namely Claude Fable 5 and GPT-5.6 Sol&amp;rdquo;. The comparison chart on Artificial Analysis&amp;rsquo;s own K3 page,
on 6 August 2026, runs Claude Opus 5 (max) at 60.69, Claude Fable 5 (with fallback) at 59.86,
GPT-5.6 Sol (max) at 58.89, and then Kimi K3 (max) at 57.11. Fourth, not sandwiched between two
Claudes. I wrote the sandwich version first, by picking the two comparators that flattered the
argument, which is the same error this article is about one level up.&lt;/p&gt;
&lt;p&gt;One more thing worth reading off the
&lt;a href="https://github.com/MoonshotAI/Kimi-K3/blob/main/README.md" class="external-link" rel="noopener"&gt;repository README&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;: every K3 result in
those benchmark tables was produced at reasoning effort &amp;ldquo;max&amp;rdquo;, temperature 1.0. Forty-five lines
later the same file records that the &lt;code&gt;reasoning_effort&lt;/code&gt; field accepts &lt;code&gt;&amp;quot;low&amp;quot;&lt;/code&gt;, &lt;code&gt;&amp;quot;high&amp;quot;&lt;/code&gt; and &lt;code&gt;&amp;quot;max&amp;quot;&lt;/code&gt;,
with &lt;code&gt;&amp;quot;max&amp;quot;&lt;/code&gt; as the default. So the benchmark setting is also the setting a caller gets by not
choosing one, and the point is not that the tables are rigged. It is
that the most expensive configuration is the one you get by not choosing, which is what the $15 per
million output tokens and the 38.7 tokens per second are describing.&lt;/p&gt;
&lt;h2 id="the-two-demos-with-the-qualifiers-put-back"&gt;The two demos, with the qualifiers put back&lt;/h2&gt;
&lt;p&gt;The kernel result is real and the framing matters.
&lt;a href="https://x.com/Kimi_Moonshot/status/2077830242060923207" class="external-link" rel="noopener"&gt;Moonshot&amp;rsquo;s own post&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; states the
configuration: FLA Triton AttnRes at production scale, 96 layers, 8192-dim model, 8192 tokens, over
15 hours of iteration, cutting &lt;strong&gt;forward and backward pass combined from 283.6 ms to 114.4 ms&lt;/strong&gt; while
preserving the same numerics. That is a training-side number, not inference latency, on Moonshot&amp;rsquo;s
own architecture. Neither the vendor nor most coverage states the multiple, so: 283.6 / 114.4 =
&lt;strong&gt;2.479x&lt;/strong&gt;, a 169.2 ms absolute reduction, 59.7% cut.&lt;/p&gt;
&lt;p&gt;The chip demo has a single source and it is
&lt;a href="https://www.kimi.com/blog/kimi-k3" class="external-link" rel="noopener"&gt;Moonshot&amp;rsquo;s own write-up&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;, which is more careful than the number
that travelled from it:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;As an early proof of concept, Kimi K3 designed a chip to serve a nano model built on its own
architecture. In a single 48-hour autonomous run, K3 built, optimized, and verified the chip using
open-source EDA tools on the Nangate 45nm library. Within 4 mm², the chip closes timing at 100 MHz
and sustains over 8,700 tokens/s decode throughput in simulation, packing 1.46M standard cells,
0.277 MB of SRAM, and an INT4 MAC array with fused dequantization.&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;Three qualifiers are inside that sentence and none of them survived the coverage. Nangate 45nm is an
open academic standard-cell library, not a foundry process. &amp;ldquo;In simulation&amp;rdquo; governs the throughput
figure. And the part serves &amp;ldquo;a nano model built on its own architecture&amp;rdquo;, so 8,700 tokens per second
describes a small model on a design that has never been fabricated, at 100 MHz. Nothing was taped
out. No silicon exists.&lt;/p&gt;
&lt;p&gt;My own post hedged this with &amp;ldquo;these results should be interpreted carefully until they&amp;rsquo;re
independently replicated&amp;rdquo;, which points at the wrong risk. The limits were not waiting on
replication. They were in Moonshot&amp;rsquo;s own sentence, and I kept the headline and dropped them.&lt;/p&gt;
&lt;h2 id="open-weight-and-the-licence-has-numbers-in-it"&gt;Open-weight, and the licence has numbers in it&lt;/h2&gt;
&lt;p&gt;The weights are released. The training data and training code are not. That makes this an
open-weight release, and &amp;ldquo;open source&amp;rdquo; describes a different thing.&lt;/p&gt;
&lt;p&gt;The licence is not a different family from K2&amp;rsquo;s, which is how I first read it. The
&lt;a href="https://huggingface.co/moonshotai/Kimi-K3/raw/main/LICENSE" class="external-link" rel="noopener"&gt;Kimi K3 License&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; is textually
MIT-derived: &amp;ldquo;Permission is hereby granted, free of charge&amp;rdquo;, to &amp;ldquo;use, copy, modify, merge, publish,
distribute, sublicense, and/or sell copies of the Software&amp;rdquo;. Its clause 3, requiring any product
with more than &lt;strong&gt;100M monthly active users or $20M monthly revenue&lt;/strong&gt; to display &amp;ldquo;Kimi K3&amp;rdquo;
prominently in its UI, is the same attribution clause K2&amp;rsquo;s Modified MIT already carried. The
genuinely new restriction is clause 2: a Model-as-a-Service operator whose aggregate revenue exceeds
&lt;strong&gt;$20M over any consecutive 12 months&lt;/strong&gt; must sign a separate agreement with Moonshot before any
commercial use. Clause 4 exempts purely internal use, and use through Moonshot&amp;rsquo;s own products or
certified inference partners.&lt;/p&gt;
&lt;p&gt;So the sentence I wrote about developers who can &amp;ldquo;inspect, customize, fine-tune, and deploy on their
own infrastructure&amp;rdquo; holds up on one verb out of four. Inspection is weights only. Customisation and
fine-tuning are permitted under those two thresholds rather than unconditionally. And deployment on
your own infrastructure means at least sixteen B200s, which for almost everyone means renting the
same hardware from the same handful of clouds you were already renting the closed model from.&lt;/p&gt;
&lt;h2 id="the-gap-thesis-and-the-error-bars-nobody-quotes"&gt;The gap thesis, and the error bars nobody quotes&lt;/h2&gt;
&lt;p&gt;My post said three times, in three phrasings, that the gap between open and closed models is
shrinking. The only organisation that measures it publishes numbers pointing the other way, and
publishes intervals that are more informative than either point estimate.&lt;/p&gt;
&lt;p&gt;Table: Epoch AI&amp;rsquo;s published estimates of how far open-weight models lag closed-weight models, with the 90% confidence intervals stated on each page.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Epoch AI page&lt;/th&gt;
 &lt;th&gt;Published&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Time gap&lt;/th&gt;
 &lt;th style="text-align: right"&gt;ECI gap&lt;/th&gt;
 &lt;th style="text-align: right"&gt;90% CI on the ECI gap&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;Open-weight models lag state-of-the-art by around 3 months&lt;/td&gt;
 &lt;td&gt;30 Oct 2025&lt;/td&gt;
 &lt;td style="text-align: right"&gt;3.5 months (CI 1.1 to 5.3)&lt;/td&gt;
 &lt;td style="text-align: right"&gt;7 points&lt;/td&gt;
 &lt;td style="text-align: right"&gt;0 to 14&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Open models lag state-of-the-art closed models by 4 months&lt;/td&gt;
 &lt;td&gt;29 May 2026&lt;/td&gt;
 &lt;td style="text-align: right"&gt;4 months&lt;/td&gt;
 &lt;td style="text-align: right"&gt;8 points&lt;/td&gt;
 &lt;td style="text-align: right"&gt;7 to 11&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The point estimate moved from 3.5 months to four. That movement sits well inside the older interval
of 1.1 to 5.3 months, so on its own it establishes nothing.&lt;/p&gt;
&lt;p&gt;The interval is where the argument actually turns, and it is not the one I would have quoted from
memory. Epoch&amp;rsquo;s &lt;a href="https://epoch.ai/data-insights/open-weights-vs-closed-weights-models" class="external-link" rel="noopener"&gt;November page&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;
puts the vertical gap at 7 ECI points with a 90% interval of &lt;strong&gt;0 to 14&lt;/strong&gt;, and says the gap &amp;ldquo;varies
considerably over time, sometimes even closing completely&amp;rdquo;. Its
&lt;a href="https://epoch.ai/data-insights/open-closed-eci-gap" class="external-link" rel="noopener"&gt;May 2026 page&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; puts it at 8 points with a 90%
interval of &lt;strong&gt;7 to 11&lt;/strong&gt;, roughly a third as wide, and that interval excludes zero. Over the window
Epoch measured, 1 January to 28 May 2026, &amp;ldquo;sometimes closing completely&amp;rdquo; is no longer in the data.
So the claim that fails is mine, not theirs. The four-month figure is still soft in the other
direction, because Epoch says the same estimate &amp;ldquo;would grow to six months&amp;rdquo; under a stricter catch-up
criterion, and softness in that direction does not help me either.&lt;/p&gt;
&lt;p&gt;Two more things about that data. Epoch&amp;rsquo;s May 2026 publication has Kimi K2.6 at ECI 151.60 as its
leading open model. It predates K3 entirely, so it says nothing about K3. And the defensible version
of my argument is the one I did not make: holding a roughly constant three-to-four-month lag while
the frontier itself accelerated is a harder result than closing a gap against a stationary target.
That framing survives the data. The one I used does not.&lt;/p&gt;
&lt;h2 id="how-to-detect-this-before-you-publish-it"&gt;How to detect this before you publish it&lt;/h2&gt;
&lt;p&gt;Every error above is findable in under an hour, and each one has a tell.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;A round number that came out of a multiplication.&lt;/strong&gt; 1,400 GB is 2.8T times four bits and
nothing else. If a figure is the product of two headline numbers, the artifact it describes will
have a real size, and the real size is the one to publish. Read the file listing.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A precision label applied to a whole model.&lt;/strong&gt; &amp;ldquo;4-bit&amp;rdquo;, &amp;ldquo;FP8&amp;rdquo;, &amp;ldquo;INT4&amp;rdquo;. Open &lt;code&gt;config.json&lt;/code&gt; and
read &lt;code&gt;quantization_config.ignore&lt;/code&gt; before you believe it covers everything.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;An adjective attached to a mechanism.&lt;/strong&gt; &amp;ldquo;New&amp;rdquo;, &amp;ldquo;novel&amp;rdquo;, &amp;ldquo;first&amp;rdquo;. Search arXiv for the mechanism
name before repeating it. Both of K3&amp;rsquo;s had papers with submission dates.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A performance number without its measured configuration.&lt;/strong&gt; The 6x throughput figure is true of a
48B model in October 2025 and false attached to K3. Ask what model, what hardware, what date.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A point estimate with no interval.&lt;/strong&gt; Epoch publishes both. The interval is what decided whether
my thesis was wrong, and it never appears in a summary.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A claim with a shelf life.&lt;/strong&gt; &amp;ldquo;Later this month&amp;rdquo; was accurate for ten days. If a sentence
expires, date it inside the sentence or do not write it.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;The rule underneath all six is narrow enough to follow: &lt;strong&gt;no vendor claim ships without its primary
artifact open in a second tab&lt;/strong&gt;, and the artifact is the paper, the licence, the config JSON, the
file listing or the pricing page, never the announcement. All four errors in my post, and the one at
the top of the first version of this page, sat in that gap.&lt;/p&gt;
&lt;p&gt;The weights shipped on 27 July 2026 at 1,560.9 GB across 96 files. That number came off a directory
listing rather than out of a multiplication, which is the only reason I am willing to put it in a
title.&lt;/p&gt;</content:encoded></item><item><title>K3 lists at half of GPT-5.6 Sol, and the saving is either 9.6% or 30.6%</title><link>https://therezaali.com/writing/price-per-token-versus-cost-per-task/</link><guid isPermaLink="true">https://therezaali.com/writing/price-per-token-versus-cost-per-task/</guid><pubDate>Thu, 06 Aug 2026 00:00:00 +0200</pubDate><author>contact@therezaali.com (Mo RezaAli)</author><category>essay</category><description>Artificial Analysis publishes both figures, for the same metric and the same two models, on two of its own pages on the same day. Both are in the table. Neither is endorsed here.</description><content:encoded>&lt;p&gt;Moonshot AI&amp;rsquo;s Kimi K3 lists at &lt;strong&gt;$15.00 per million output tokens&lt;/strong&gt;. OpenAI&amp;rsquo;s GPT-5.6 Sol lists at &lt;strong&gt;$30.00&lt;/strong&gt;. Exactly half, on both vendors&amp;rsquo; own pricing pages, read on 6 August 2026.&lt;/p&gt;
&lt;p&gt;So the saving is 50%, and it is not. Artificial Analysis runs both through its Intelligence Index and publishes what each actually costs to finish the work, and this is where I have to stop and report a problem rather than a number.&lt;/p&gt;
&lt;p&gt;Its article gives &lt;strong&gt;$0.94 per task for K3 against $1.04 for Sol&lt;/strong&gt;, a 9.6% saving. Its model page for the same model, read the same day, carries the same metric with different values: &lt;code&gt;costPerIntelligenceIndexTask&lt;/code&gt; is &lt;strong&gt;0.8552&lt;/strong&gt; for K3 and &lt;strong&gt;1.2325&lt;/strong&gt; for Sol, which is a 30.6% saving. Same publisher, same metric, same afternoon, one figure three times the other.&lt;/p&gt;
&lt;p&gt;Table: Cost per Intelligence Index task, as published by Artificial Analysis in two places on 6 August 2026.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Source&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Kimi K3&lt;/th&gt;
 &lt;th style="text-align: right"&gt;GPT-5.6 Sol&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Saving&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;Article prose&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$0.94&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$1.04&lt;/td&gt;
 &lt;td style="text-align: right"&gt;9.6%&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Model page dataset&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$0.8552&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$1.2325&lt;/td&gt;
 &lt;td style="text-align: right"&gt;30.6%&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;I cannot tell you which is right, and that is the finding. An earlier version of this page put 9.6% in its own title, having read one of those pages and not the other. What survives both readings is the direction and the mechanism: the saving is real, it is nowhere near the 50% the list prices imply, and the reason is token volume, which Moonshot documents on the same page that carries the price.&lt;/p&gt;
&lt;h2 id="where-i-was-already-standing"&gt;Where I was already standing&lt;/h2&gt;
&lt;p&gt;I had this model in a costed build before the measurement existed. In &lt;a href="https://therezaali.com/writing/always-on-agent-costs/"&gt;what an always-on AI agent actually costs to run&lt;/a&gt; I put &lt;code&gt;moonshotai/kimi-k3&lt;/code&gt; on a three-model allowlist at $3.00 in and $15.00 out, marked it escalation only, and wrote that the output column is the side of the ledger you control least, because you can budget your input and you cannot budget how much a model decides to write.&lt;/p&gt;
&lt;p&gt;I derived that rule from the list price. It was the right rule and I could not have told you the mechanism, because I had a price and nobody had published a cost. The mechanism is now on the vendor&amp;rsquo;s page, one paragraph under the price table:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Always reasons and supports configuring its reasoning effort with the top-level &lt;code&gt;reasoning_effort&lt;/code&gt; request field (&lt;code&gt;low&lt;/code&gt; / &lt;code&gt;high&lt;/code&gt; / &lt;code&gt;max&lt;/code&gt;, default &lt;code&gt;max&lt;/code&gt;).&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;Two words carry the whole gap. &lt;strong&gt;Always&lt;/strong&gt;, so there is no cheap non-reasoning path. &lt;strong&gt;Default &lt;code&gt;max&lt;/code&gt;&lt;/strong&gt;, so the most expensive setting is the one you get by not choosing. Reasoning tokens bill as output tokens, at $15.00 per million.&lt;/p&gt;
&lt;h2 id="how-much-more-it-emits-and-where-the-arithmetic-stops"&gt;How much more it emits, and where the arithmetic stops&lt;/h2&gt;
&lt;p&gt;Take the floor first, because it needs no assumptions about the workload at all.&lt;/p&gt;
&lt;p&gt;Every one of K3&amp;rsquo;s unit prices is 60% or less of Sol&amp;rsquo;s: input $3.00 against $5.00, cached input $0.30 against $0.50, output $15.00 against $30.00. So for any mix of cached, uncached and output tokens whatsoever, a task that consumed identical token counts on both models would cost K3 at most 60% of Sol&amp;rsquo;s $1.04, which is $0.62. It costs $0.94. Divide: K3 is consuming &lt;strong&gt;at least 1.5 times&lt;/strong&gt; the tokens, and that number falls out of four published prices and two published costs with no modelling in between.&lt;/p&gt;
&lt;p&gt;The obvious next step is to turn that floor into an exact multiple, and it does not survive contact with the definitions. Artificial Analysis publishes a per-task token figure for Sol: 15k output tokens per Intelligence Index task at max reasoning effort. For K3 it publishes totals instead, 130M tokens across the index at a cost of $2,437.41. Dividing that total by the $0.94 per-task cost looks like it hands you a task count, and K3&amp;rsquo;s tokens per task would fall straight out of it. But the measurer defines that per-task number two different ways. The v4.1 announcement that introduced the metric says it takes &amp;ldquo;the total cost, total time, and total output tokens for a model to run the Intelligence Index and divide by the number of tasks across its evaluations&amp;rdquo;. The K3 model page calls the same metric a &amp;ldquo;Weighted average cost (USD) per Artificial Analysis Intelligence Index task&amp;rdquo;, and the index weights are nowhere near uniform: GDPval-AA v2 carries 20% of the score, AA-Omniscience Non-Hallucination 4%. Under the first reading the division is a task count. Under the second it is an unweighted total over a weighted mean, which counts nothing. One of the two readings voids the arithmetic, so the arithmetic does not ship and 1.5x is what I am willing to put my name to.&lt;/p&gt;
&lt;p&gt;The measurer agrees on direction, in plain language. Artificial Analysis&amp;rsquo;s own summary calls K3 &amp;ldquo;notably slow and very verbose&amp;rdquo; and sets its 130M tokens against a 100M median among open-weight models of similar size, which is a narrower comparison set than the full index.&lt;/p&gt;
&lt;h2 id="the-one-place-the-list-price-tells-the-truth"&gt;The one place the list price tells the truth&lt;/h2&gt;
&lt;p&gt;There is a tier where the headline discount is real and understated, and the post I am building on missed it.&lt;/p&gt;
&lt;p&gt;Table: List prices per million tokens from each vendor&amp;rsquo;s own pricing documentation, read 6 August 2026. OpenAI&amp;rsquo;s page does not state the input-token threshold at which its long-context tier engages, so the last row applies above an unstated boundary.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Per million tokens&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Kimi K3&lt;/th&gt;
 &lt;th style="text-align: right"&gt;GPT-5.6 Sol&lt;/th&gt;
 &lt;th style="text-align: right"&gt;K3 as % of Sol&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;Input, cached&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$0.30&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$0.50&lt;/td&gt;
 &lt;td style="text-align: right"&gt;60.0%&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Input, standard&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$3.00&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$5.00&lt;/td&gt;
 &lt;td style="text-align: right"&gt;60.0%&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Output, standard&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$15.00&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$30.00&lt;/td&gt;
 &lt;td style="text-align: right"&gt;50.0%&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Output, long context&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$15.00&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$45.00&lt;/td&gt;
 &lt;td style="text-align: right"&gt;33.3%&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;K3 charges one flat output rate all the way to its 1,048,576-token window. Sol charges a premium above a threshold. Above that line the gap widens from half to a third, and there is an architectural reason rather than a commercial one. vLLM&amp;rsquo;s engineering write-up describes Kimi Delta Attention as a linear-attention mechanism holding a fixed-size recurrent state instead of a growing KV cache, interleaved with periodic full-attention layers, and says directly that this is what makes a 1M-token context affordable. Flat long-context pricing is that design showing up on an invoice.&lt;/p&gt;
&lt;p&gt;So the honest shape is the inverse of the headline. On short agentic turns, where reasoning dominates the bill, the 50% discount collapses to 9.6%. On very long single-pass context, where the recurrent state does the work, it improves to 33%.&lt;/p&gt;
&lt;h2 id="the-27-figure-measures-a-preference-vote-and-the-ratio-is-undefined"&gt;The 2.7% figure measures a preference vote, and the ratio is undefined&lt;/h2&gt;
&lt;p&gt;Stanford&amp;rsquo;s 2026 AI Index says the top US model leads by 2.7% as of March 2026. That number is real and the popular restatement of it fails three ways.&lt;/p&gt;
&lt;p&gt;It is an Arena rating, which the report itself describes as human voting on ratings &amp;ldquo;inspired by chess ratings&amp;rdquo;. It measures which answer people preferred, not what either model can do. The underlying pair, per press reads of the report, is Claude Opus 4.6 at 1,503 against ByteDance&amp;rsquo;s Dola-Seed-2.0-Preview at 1,464. A 39-point difference; 1503/1464 gives 2.66%.&lt;/p&gt;
&lt;p&gt;Dividing two Elo ratings is also not a defined operation. Elo is an interval scale with an arbitrary zero, so the ratio moves when you shift the origin while nothing about the models changes at all.&lt;/p&gt;
&lt;p&gt;Table: The same 39-point Arena gap expressed as a percentage, with every rating on the board shifted by a constant. Win rate is 1 / (1 + 10^(-39/400)) throughout.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Shift applied to all ratings&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Gap as a percentage&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Head-to-head win rate&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;none&lt;/td&gt;
 &lt;td style="text-align: right"&gt;2.66%&lt;/td&gt;
 &lt;td style="text-align: right"&gt;55.6%&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;+500&lt;/td&gt;
 &lt;td style="text-align: right"&gt;1.99%&lt;/td&gt;
 &lt;td style="text-align: right"&gt;55.6%&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;+1,000&lt;/td&gt;
 &lt;td style="text-align: right"&gt;1.58%&lt;/td&gt;
 &lt;td style="text-align: right"&gt;55.6%&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The quantity that survives the shift is the expected score. The top US model was preferred in &lt;strong&gt;about 56 of every 100 blind comparisons&lt;/strong&gt;, which is a far less comfortable number than 2.7% for anyone arguing the models are interchangeable. It is also five months stale for August 2026, predating Kimi K3 entirely.&lt;/p&gt;
&lt;p&gt;What the report actually concludes is better than the statistic anyway: competitive pressure is shifting &amp;ldquo;toward cost, reliability, and domain-specific performance&amp;rdquo;. That is the argument, from the primary source.&lt;/p&gt;
&lt;h2 id="open-weights-with-the-licence-read"&gt;Open weights, with the licence read&lt;/h2&gt;
&lt;p&gt;The weights are open and the licence is new, which is the part worth reading. Moonshot&amp;rsquo;s own back catalogue is not a guide to it. Kimi K2 and Kimi K2.6 each ship a file whose first line is &lt;code&gt;Modified MIT License&lt;/code&gt;, and the K2 text says exactly what the modification is:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Our only modification part is that, if the Software (or any derivative works thereof) is used for any of your commercial products or services that have more than 100 million monthly active users, or more than 20 million US dollars (or equivalent in other currencies) in monthly revenue, you shall prominently display &amp;ldquo;Kimi K2&amp;rdquo; on the user interface of such product or service.&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;K3 ships a bespoke &lt;code&gt;Kimi K3 License&lt;/code&gt; instead. Its section 3 is that same branding clause with the model name changed. Its section 2 has no counterpart in the K2 file at all:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;If the Licensee or any of its affiliates operates a Model as a Service business, and the aggregate revenue of the Licensee and its affiliates exceeds 20 million US dollars (or the equivalent in other currencies) in total over any consecutive 12 months, the Licensee must enter into a separate agreement with Moonshot AI before using the Software or its derivative works for any commercial purpose.&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;Read the trigger carefully. It is the aggregate revenue of the licensee and its affiliates over any twelve months, not the revenue earned from serving K3, and 20 million dollars is a small company. An inference provider above that line needs a signed agreement with Moonshot before it charges for a single token, unless it is one of the certified inference partners that section 4(b) exempts. The claim that you can deploy the weights yourself and fine-tune them survives, but through section 4(a), which exempts internal use, defined as use that does not make the software, its outputs or its underlying capabilities available to third parties.&lt;/p&gt;
&lt;p&gt;Self-hosting has a hardware floor too. vLLM&amp;rsquo;s launch post states the model &amp;ldquo;can barely fit in a single NVIDIA DGX B300 and requires a minimum of 16 NVIDIA B200/GB200 GPUs to serve on that hardware generation&amp;rdquo;. Hugging Face&amp;rsquo;s repository metadata for the model reports 2.78 trillion parameters, 2.72 trillion of them stored as packed 4-bit MXFP4, across 1.56 TB of safetensors. That is the download, before a byte of KV cache or recurrent state.&lt;/p&gt;
&lt;p&gt;Which undercuts the data-residency argument for most of the named adopters, because they are not self-hosting. Airbnb uses Alibaba&amp;rsquo;s Qwen for customer service; DoorDash routes lower-level coding work to Kimi for, in its CTO&amp;rsquo;s words, &amp;ldquo;better quality [at] cheaper cost&amp;rdquo;. Both consume an API, and K3 reaches Azure customers through Fireworks on Microsoft Foundry. Routing through a US intermediary to a Chinese-trained model is a different data posture from running weights on your own metal, and US lawmakers requested information from DoorDash about exactly this on 31 July.&lt;/p&gt;
&lt;p&gt;That distribution route also corrects the source post: &lt;strong&gt;Microsoft is not an adopter, it is the channel.&lt;/strong&gt; Kimi K2.6 and K2.5 are sold directly through Microsoft Foundry. The American hyperscaler is monetising the Chinese price advantage as a distributor, which is a better fact than the one it replaces. Siemens should come out of a list about American buyers too; it is a German industrial group.&lt;/p&gt;
&lt;h2 id="three-corrections-to-the-surrounding-numbers"&gt;Three corrections to the surrounding numbers&lt;/h2&gt;
&lt;p&gt;The semiconductor selloff is real and smaller than it reads. The SOX fell 1.6% on 17 July to close 20.2% below its 22 June record, its steepest week since April 2025, and the same report notes the index was still up more than 60% on the year. A 20% drawdown from a 105% run is profit-taking with a bear-market label attached.&lt;/p&gt;
&lt;p&gt;The $725 billion capex figure is four hyperscalers, not five, and &amp;ldquo;roughly&amp;rdquo; rather than &amp;ldquo;more than&amp;rdquo;. It is an analyst aggregation of company guidance whose own component table sums to well under its headline, with Microsoft appearing at two different values on one page.&lt;/p&gt;
&lt;p&gt;And the trend runs both ways. K3&amp;rsquo;s $15.00 output price is &lt;strong&gt;3.75 times&lt;/strong&gt; its own predecessor&amp;rsquo;s $4.00. The lab leading the price attack quadrupled its own rate, which is the strongest available objection to the commoditisation thesis.&lt;/p&gt;
&lt;h2 id="detecting-this-in-your-own-stack"&gt;Detecting this in your own stack&lt;/h2&gt;
&lt;p&gt;The test is one query against your provider&amp;rsquo;s billing export, and it works for any model pair.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Take one real task class, not a prompt. A whole agent turn including tool calls and retries.&lt;/li&gt;
&lt;li&gt;Pull total spend and total completions for that class over a week, and divide. That is your cost per task. Per-token price is an input to it, never a substitute for it.&lt;/li&gt;
&lt;li&gt;Split output tokens into reasoning and visible answer. If your provider reports reasoning tokens separately, the ratio is the story.&lt;/li&gt;
&lt;li&gt;Compute the floor: multiply the incumbent&amp;rsquo;s cost per task by the challenger&amp;rsquo;s worst unit-price ratio. If the challenger costs more than that, it is emitting more tokens, and you now know the minimum multiple.&lt;/li&gt;
&lt;li&gt;Re-run at each &lt;code&gt;reasoning_effort&lt;/code&gt; level before you conclude anything. A default of &lt;code&gt;max&lt;/code&gt; is a benchmark setting, not a production setting.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="designing-so-the-bill-cannot-surprise-you"&gt;Designing so the bill cannot surprise you&lt;/h2&gt;
&lt;p&gt;Set &lt;code&gt;reasoning_effort&lt;/code&gt; explicitly on every call, including the ones where you want the default, so the value is in your code rather than in the vendor&amp;rsquo;s. Cap &lt;code&gt;max_tokens&lt;/code&gt; per task class, because an uncapped verbose model has no upper bound you control. Price migrations on measured cost per task and refuse to sign off on a per-token comparison. And route by tier: long single-pass context to the flat-rate model, short reasoning-heavy turns to whichever model finishes in fewer tokens.&lt;/p&gt;
&lt;h2 id="what-i-have-now-that-i-did-not"&gt;What I have now that I did not&lt;/h2&gt;
&lt;p&gt;A floor calculation that converts two published prices and two published costs into a hard lower bound on token consumption, which I can run on any model pair the day it launches. The knowledge that &lt;code&gt;reasoning_effort&lt;/code&gt; defaults to &lt;code&gt;max&lt;/code&gt; on the model I had already put on an allowlist, which turns a rule I wrote on instinct into one I can defend. An Elo shift table that retires percentage comparisons of Arena ratings permanently. And the correction that reorganises the whole thesis: the cheap model is cheap per token and average per task, so the commoditisation of foundation models is arriving through a different door than the price list suggests.&lt;/p&gt;
&lt;p&gt;Half price bought 9.6%. Measure the task.&lt;/p&gt;</content:encoded></item><item><title>Google’s checkout form guidance ships an autofill token that does not exist</title><link>https://therezaali.com/writing/mobile-checkout-evidence/</link><guid isPermaLink="true">https://therezaali.com/writing/mobile-checkout-evidence/</guid><pubDate>Thu, 06 Aug 2026 00:00:00 +0200</pubDate><author>contact@therezaali.com (Mo RezaAli)</author><category>essay</category><description>The string address-line-1 appears five times on web.dev, four of them inside an autocomplete attribute. Also the single-column study that excluded mobile, and why device conversion splits are partly artifact.</description><content:encoded>&lt;p&gt;Google&amp;rsquo;s own guidance page for checkout forms ships an autofill token that does not exist.
On web.dev&amp;rsquo;s &lt;em&gt;Payment and address form best practices&lt;/em&gt;, the string &lt;code&gt;address-line-1&lt;/code&gt; appears
five times, four of them inside an &lt;code&gt;autocomplete&lt;/code&gt; attribute: &lt;code&gt;autocomplete=&amp;quot;address-line-1&amp;quot;&lt;/code&gt;
twice, plus &lt;code&gt;autocomplete=&amp;quot;shipping address-line-1&amp;quot;&lt;/code&gt; and &lt;code&gt;autocomplete=&amp;quot;billing address-line-1&amp;quot;&lt;/code&gt;. The HTML Standard defines the token as &lt;code&gt;address-line1&lt;/code&gt;, with no hyphen
before the digit, and contains zero occurrences of the hyphenated spelling. I checked both
pages by string search on 6 August 2026. The same web.dev code sample gets it right in one
attribute and wrong in the next: &lt;code&gt;&amp;lt;input autocomplete=&amp;quot;address-line-1&amp;quot; id=&amp;quot;address-line1&amp;quot;&amp;gt;&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;An unrecognised token does not throw and does not warn in the console. It simply fails to
match, leaving the field to the browser&amp;rsquo;s own field-name heuristics, which is where it would
have been with no attribute at all. So a developer who copies the most-linked checkout forms
page on the web gets worse mobile autofill than one who copies the specification, with no
signal that anything happened. That matters more than any layout advice on the same page,
because autofill is the only checkout change that removes the typing instead of rearranging
it.&lt;/p&gt;
&lt;h2 id="why-i-was-reading-the-specification"&gt;Why I was reading the specification&lt;/h2&gt;
&lt;p&gt;Because I stopped trusting my own reason for rebuilding a checkout.&lt;/p&gt;
&lt;p&gt;Here is what I published at the time. These are my own unaudited figures from my own work,
not anything a third party measured. 74% of my visitors were on mobile, and mobile customers
were converting at roughly half the rate of desktop. Before buying more traffic I opened the
store on my own phone and found a two-column form, shipping fields ahead of payment, and a
coupon box sitting there inviting people to go hunting for a code. I made it single column,
moved payment higher in the flow, and removed the coupon field on mobile. The edit took
around 30 minutes. I ran the variant for a week and mobile completion improved from far
behind desktop to almost matching it. The post is
&lt;a href="https://x.com/Mo_ali/status/2078624715707441344" class="external-link" rel="noopener"&gt;here&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;The change was right. The reasoning that led to it was not, and the test that confirmed it
cannot support the lesson I drew from it. Both of those are fixable, and fixing them is
worth more than the result was.&lt;/p&gt;
&lt;h2 id="the-conversion-gap-was-never-evidence-about-the-checkout"&gt;The conversion gap was never evidence about the checkout&lt;/h2&gt;
&lt;p&gt;The gap on its own carries no information about my form. It is one ratio with no comparison
group: nothing in it separates &amp;ldquo;my checkout is bad&amp;rdquo; from &amp;ldquo;a phone is a worse buying
environment than a laptop&amp;rdquo;. I looked for a free primary source publishing a current
device-split conversion benchmark and did not find one. The vendor benchmarks that get quoted
for this sit behind lead capture forms, so I will not quote a figure I could not open. What
justified the work was opening the checkout on a phone. The gap was the prompt. The
walkthrough was the evidence. I had those two backwards.&lt;/p&gt;
&lt;p&gt;The gap is worse than uninformative, though, because part of it is instrumentation.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cross-device purchases debit mobile and credit desktop.&lt;/strong&gt; GA4&amp;rsquo;s reporting identity falls
back to device ID whenever User-ID is not collected, per Google&amp;rsquo;s own documentation. A
shopper who browses on a phone and buys on a laptop adds a non-converting session to the
mobile denominator and a converting session to the desktop numerator. Most small stores never
set User-ID, so their device-split report is biased against mobile by construction.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Safari deletes the state that would let you recognise a returning visitor.&lt;/strong&gt; WebKit
announced that ITP &amp;ldquo;would cap the expiry of client-side cookies to seven days&amp;rdquo;, then extended
the same rule to everything a script can write, &amp;ldquo;deleting all of a website&amp;rsquo;s script-writable
storage after seven days of Safari use without user interaction on the site&amp;rdquo;. The affected
list includes IndexedDB, LocalStorage, SessionStorage, media keys, and Service Worker
registrations. Mobile traffic skews heavily to Safari. Returning iOS visitors are therefore
re-counted as new sessions more often than returning desktop Chrome visitors, which inflates
the mobile session denominator.&lt;/p&gt;
&lt;p&gt;Both artifacts push the same way, and neither is visible in the report they corrupt.&lt;/p&gt;
&lt;h3 id="how-to-detect-this-in-your-own-system"&gt;How to detect this in your own system&lt;/h3&gt;
&lt;p&gt;Stop comparing sessions-to-orders by device. Compare &lt;strong&gt;checkout-start to order&lt;/strong&gt; by device
instead. That ratio lives inside a single session on a single device, so cross-device
attribution cannot touch it and seven-day storage deletion cannot touch it. In GA4 terms it
is &lt;code&gt;purchase&lt;/code&gt; over &lt;code&gt;begin_checkout&lt;/code&gt;, segmented by device category.&lt;/p&gt;
&lt;p&gt;Run both numbers side by side for a week before changing anything. If session conversion
shows a wide device gap and checkout completion shows near parity, the problem is upstream of
the checkout and redesigning the form will do nothing. If checkout completion is also far
apart, you have a checkout problem and a metric clean enough to measure the fix with. Setting
User-ID removes the first artifact for logged-in traffic. Nothing removes the second.&lt;/p&gt;
&lt;h2 id="what-the-single-column-evidence-actually-says"&gt;What the single-column evidence actually says&lt;/h2&gt;
&lt;p&gt;The canonical citation is Ben Labay&amp;rsquo;s study for CXL Institute, published 17 October 2016.
CXL&amp;rsquo;s own copy now returns a 403, including to a request carrying a current Chrome user agent,
so the readable original is on Speero. The instrument was CXL&amp;rsquo;s Trust Seal survey, not a
checkout. Nobody bought anything.
The outcome measure was form completion &lt;strong&gt;time&lt;/strong&gt;, not completion rate and not revenue. The
sample was 702 people fielded in early June 2016, n = 356 linear against n = 346
multi-column, and the single-column form was completed 15.4 seconds faster, significant at
95% on a two-sample t-test. And then the sentence almost nobody who cites the study
reproduces: &amp;ldquo;We restricted the survey to desktop users.&amp;rdquo; The authors add their own caveat,
that the results are &amp;ldquo;not directly transferable to all form types and situations&amp;rdquo;.&lt;/p&gt;
&lt;p&gt;So the best-known quantitative evidence for single-column forms measured speed, on a survey,
on desktop, ten years ago, and it is routinely quoted as proof that single column &lt;em&gt;converts&lt;/em&gt;
better, in &lt;em&gt;checkouts&lt;/em&gt;, on &lt;em&gt;mobile&lt;/em&gt;.&lt;/p&gt;
&lt;p&gt;The change is still correct on mobile. It needs a mechanical argument rather than a borrowed
statistic. On a 390 CSS-pixel viewport, two columns either shrink each control below a
comfortable tap target or trigger pinch-zoom, and they break the vertical DOM order that the
on-screen keyboard&amp;rsquo;s &amp;ldquo;Next&amp;rdquo; key walks. All three are confirmable in a device emulator in
under a minute, which beats a decade-old desktop survey as a standard of proof.&lt;/p&gt;
&lt;h2 id="the-change-that-beats-layout"&gt;The change that beats layout&lt;/h2&gt;
&lt;p&gt;If friction is the problem, the strongest available answer is fewer keystrokes, and the
specification is where that lives. The HTML Standard&amp;rsquo;s autofill token list is the actual API:
&lt;code&gt;given-name&lt;/code&gt;, &lt;code&gt;family-name&lt;/code&gt;, &lt;code&gt;street-address&lt;/code&gt;, &lt;code&gt;address-line1&lt;/code&gt;, &lt;code&gt;address-line2&lt;/code&gt;,
&lt;code&gt;address-level1&lt;/code&gt;, &lt;code&gt;address-level2&lt;/code&gt;, &lt;code&gt;postal-code&lt;/code&gt;, &lt;code&gt;country&lt;/code&gt;, &lt;code&gt;country-name&lt;/code&gt;, &lt;code&gt;tel&lt;/code&gt;, &lt;code&gt;email&lt;/code&gt;,
&lt;code&gt;cc-name&lt;/code&gt;, &lt;code&gt;cc-number&lt;/code&gt;, &lt;code&gt;cc-exp&lt;/code&gt;, &lt;code&gt;cc-csc&lt;/code&gt;, and the &lt;code&gt;shipping&lt;/code&gt; and &lt;code&gt;billing&lt;/code&gt; section prefixes
for forms that collect both addresses.&lt;/p&gt;
&lt;p&gt;Three more from Google&amp;rsquo;s page, which are right even though its address token is not. Use
&lt;code&gt;type=&amp;quot;tel&amp;quot;&lt;/code&gt; for phone numbers and &lt;code&gt;type=&amp;quot;email&amp;quot;&lt;/code&gt; for email, so the phone renders the correct
keyboard. Never use &lt;code&gt;type=&amp;quot;number&amp;quot;&lt;/code&gt; for a card number: it &amp;ldquo;adds an up/down arrow to increment
numbers, which makes no sense for data such as telephone, payment card or account numbers&amp;rdquo;.
Use &lt;code&gt;type=&amp;quot;text&amp;quot;&lt;/code&gt; with &lt;code&gt;inputmode=&amp;quot;numeric&amp;quot;&lt;/code&gt; instead, in a single input rather than four split
boxes.&lt;/p&gt;
&lt;h2 id="what-the-abandonment-statistics-actually-say"&gt;What the abandonment statistics actually say&lt;/h2&gt;
&lt;p&gt;Baymard&amp;rsquo;s cart abandonment figure is 70.22%, and Baymard states plainly what it is: &amp;ldquo;an
average calculated based on 50 different studies&amp;rdquo;, spanning 2006 to 2025, ranging from 55.00%
(Forrester 2010) to 84.27% (SaleCycle 2020), most recently fed by Uptain 2025 at 71.72%. It
is a literature average of other people&amp;rsquo;s measurements across nineteen years, not a
measurement of anything.&lt;/p&gt;
&lt;p&gt;The reasons breakdown matters more here, and the circulating version is stale. On the live
page, 42% of US online shoppers abandoned because &amp;ldquo;I was just browsing / not ready to buy&amp;rdquo;,
and Baymard removes that segment before computing everything else. Of the remainder: 40%
extra costs, 20% delivery too slow, 19% did not trust the site with card details, 18% forced
account creation, and &lt;strong&gt;17% &amp;ldquo;too long / complicated checkout process&amp;rdquo;&lt;/strong&gt;. The 22% still quoted
for checkout complexity was true of this same URL two years ago. The Internet Archive&amp;rsquo;s
&lt;a href="https://web.archive.org/web/20240601015759/https://baymard.com/lists/cart-abandonment-rate" class="external-link" rel="noopener"&gt;1 June 2024 capture&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;
reads &amp;ldquo;22 % Too long / complicated checkout process&amp;rdquo; under a heading that says &amp;ldquo;2024 data&amp;rdquo;,
and the
&lt;a href="https://web.archive.org/web/20220101073841/https://baymard.com/lists/cart-abandonment-rate" class="external-link" rel="noopener"&gt;1 January 2022 capture&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;
reads 18% under &amp;ldquo;2021 data&amp;rdquo;. One URL, one question, three answers in four years: 18, then 22,
now 17. The 26% that travels alongside it is not a checkout-complexity figure at all. It is
the forced-account-creation row of that same 2024 table, &amp;ldquo;The site wanted me to create an
account&amp;rdquo;, which the live page now gives as 18%. Cite 17%, name the date you read it, and say
it is calculated
after excluding four in ten respondents who were never going to buy. Baymard publishes no
sample size or field date for that survey, so do not invent one.&lt;/p&gt;
&lt;p&gt;Two more from the same publisher, and they disagree. The abandonment page, updated 22
September 2025, says the average US checkout displays 23.48 form elements by default, 14.88
counting only fields, against an ideal of 12 to 14 elements or 7 to 8 fields. Baymard&amp;rsquo;s blog
post of 26 June 2024 says the 2024 average was 11.3 form fields, down from 11.8 in 2021 and
12.7 in 2019, against an ideal of 8. Different pages, different vintages, different
definitions of what counts. Pick one, name it with its date, and do not average them.
Baymard&amp;rsquo;s claim that better checkout UX is worth a 35% conversion increase comes from its own
usability sessions and is sold as part of a paid product, so it is a vendor estimate.&lt;/p&gt;
&lt;p&gt;The mobile speed number has degraded the same way. &lt;em&gt;Milliseconds Make Millions&lt;/em&gt;, by Google,
Deloitte Digital and Fifty-Five, reports an 8.4% conversion increase and a 9.2% average order
value increase per 0.1s, from 37 brands, retail findings resting on 20.5m sessions. Its
methodology section says &amp;ldquo;Approximately 4 weeks&amp;rsquo; worth of hourly data was collected to reach
statistical significance&amp;rdquo; and never says when those four weeks were. The document carries a
&amp;ldquo;©2020 Deloitte Ireland LLP&amp;rdquo; line, so the data is at least six years old by now, but the
window itself is undated and any date you see attached to this study is somebody&amp;rsquo;s inference.
It is a logarithmic regression over an uncontrolled panel, and nobody&amp;rsquo;s site was made faster
as an intervention. The appendix defines its own test: &amp;ldquo;In this
context, statistically significant means a direct correlation was identified.&amp;rdquo; The same
paragraph adds that where speed had &amp;ldquo;minimal or zero effect on funnel progression&amp;rdquo;, those
observations &amp;ldquo;were not included in the report&amp;rdquo;. Its lead generation section records mobile
conversion rates falling by almost 2% with faster speed, while the executive summary presents
speed as uniformly positive. Faster sessions are also newer phones on better connections in
richer postcodes, and the study cannot separate those.&lt;/p&gt;
&lt;h2 id="what-my-test-could-and-could-not-tell-me"&gt;What my test could and could not tell me&lt;/h2&gt;
&lt;p&gt;Fixing the duration at one week in advance was right for two reasons I would like to claim I
planned. It covers a full weekly cycle, so day-of-week composition is balanced, and it is a
fixed horizon rather than a stopping rule. Evan Miller&amp;rsquo;s demonstration is the one to keep in
mind: running a significance test after every observation produces a 26.1% false positive
rate against a nominal 5%, and his prescription is &amp;ldquo;Decide on a sample size in advance and
wait until the experiment is over.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;What I did not do was write down what &amp;ldquo;completion rate&amp;rdquo; was measured against. That omission
decides whether the result means anything, because required sample size moves by two orders
of magnitude depending on the denominator. Two-sided alpha of 0.05, power 0.80, standard
two-proportion test. The base rates below are illustrative, not mine: I never published my
own, and inventing them afterwards would be worse than admitting the gap.&lt;/p&gt;
&lt;p&gt;Table: Sessions required per arm to detect a given effect at 80% power, by choice of denominator&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Metric and effect&lt;/th&gt;
 &lt;th style="text-align: right"&gt;n per arm&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Total&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Mobile traffic per day for 7 days&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;Session CVR 1.5% to 2.7% (+80% rel.)&lt;/td&gt;
 &lt;td style="text-align: right"&gt;2,241&lt;/td&gt;
 &lt;td style="text-align: right"&gt;4,482&lt;/td&gt;
 &lt;td style="text-align: right"&gt;641&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Session CVR 1.5% to 1.8% (+20% rel.)&lt;/td&gt;
 &lt;td style="text-align: right"&gt;28,304&lt;/td&gt;
 &lt;td style="text-align: right"&gt;56,608&lt;/td&gt;
 &lt;td style="text-align: right"&gt;8,087&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Session CVR 1.5% to 1.65% (+10% rel.)&lt;/td&gt;
 &lt;td style="text-align: right"&gt;108,153&lt;/td&gt;
 &lt;td style="text-align: right"&gt;216,306&lt;/td&gt;
 &lt;td style="text-align: right"&gt;30,901&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Checkout completion 45% to 55% (+22% rel.)&lt;/td&gt;
 &lt;td style="text-align: right"&gt;392&lt;/td&gt;
 &lt;td style="text-align: right"&gt;784&lt;/td&gt;
 &lt;td style="text-align: right"&gt;112&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Checkout completion 45% to 50% (+11% rel.)&lt;/td&gt;
 &lt;td style="text-align: right"&gt;1,565&lt;/td&gt;
 &lt;td style="text-align: right"&gt;3,130&lt;/td&gt;
 &lt;td style="text-align: right"&gt;448&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Checkout completion 45% to 48% (+6.7% rel.)&lt;/td&gt;
 &lt;td style="text-align: right"&gt;4,338&lt;/td&gt;
 &lt;td style="text-align: right"&gt;8,676&lt;/td&gt;
 &lt;td style="text-align: right"&gt;1,240&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;If completion rate meant checkout-starts to orders, a week at a few hundred mobile checkout
starts a day is adequately powered and the finding stands. If it meant sessions to orders, a
week of small-store traffic can only detect implausibly large effects, and a result at that
volume is more likely noise than signal. I did not record which, so the honest statement is
that the power of that test is unknown.&lt;/p&gt;
&lt;p&gt;Two further limits, both structural. Three changes shipped as one variant, which is a valid
test of &amp;ldquo;ship all three or none&amp;rdquo; and an invalid basis for the lesson &amp;ldquo;single column works&amp;rdquo;.
And removing the coupon field was measured with a metric that cannot adjudicate it: taking
the discount box away mechanically raises revenue per completed order, because fewer
discounts get redeemed, while potentially suppressing orders from people holding a valid code.
Completion rate reads both of those as success. Revenue per session tells them apart.&lt;/p&gt;
&lt;h2 id="designing-so-it-cannot-recur"&gt;Designing so it cannot recur&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Instrument checkout-start to order by device permanently, and treat session-level device conversion as a directional signal at best.&lt;/li&gt;
&lt;li&gt;Write the denominator and the required sample size down before shipping the variant. If the required n exceeds what a week of traffic can deliver, do not run the test. Ship on mechanism and monitor.&lt;/li&gt;
&lt;li&gt;One change per variant, or state openly that you are testing a bundle and cannot attribute the result.&lt;/li&gt;
&lt;li&gt;Match the metric to the change. Layout goes to completion rate. Anything touching discounts goes to revenue per session.&lt;/li&gt;
&lt;li&gt;Copy autofill tokens from the WHATWG specification, not from articles about it, including this one. Check the list yourself.&lt;/li&gt;
&lt;li&gt;The strongest version of removing friction is removing the form, and it is an addition rather than a substitute. The Payment Request API hands the browser&amp;rsquo;s stored card and address to the site through a native sheet, so there is no layout to argue about. Read the banner at the top of MDN&amp;rsquo;s page for it before you plan around it: &amp;ldquo;Limited availability&amp;rdquo;, and &amp;ldquo;This feature is not Baseline because it does not work in some of the most widely-used browsers&amp;rdquo;. It also requires a secure context. So it is a progressive enhancement over a correctly tokenised form, and the form is still what a share of your mobile traffic will actually see. Building the sheet and skipping the tokens leaves the fallback path worse than where you started.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;And know your platform before quoting anyone&amp;rsquo;s timings, including mine. On Shopify,
&amp;ldquo;Checkout UI extensions for the information, shipping, and payment steps are available only
to stores on a Shopify Plus plan&amp;rdquo;, so a 30-minute checkout edit is not a number that
transfers. The platform sets the floor for the rest, too. HTTP Archive&amp;rsquo;s 2025 crawl puts
Shopify origins at 76% passing Core Web Vitals in both the desktop and mobile tables, against
WooCommerce at 33% and 35%. Before blaming your traffic, find out which of those you are
standing on.&lt;/p&gt;
&lt;h2 id="what-holds-up"&gt;What holds up&lt;/h2&gt;
&lt;p&gt;The lesson in my original post survives with one repair. Sometimes the ads are not the
problem and the funnel is. But the funnel claim needs a metric the browser and the analytics
tool have not already corrupted, and the device-split conversion rate that made me suspicious
in the first place does not qualify.&lt;/p&gt;
&lt;p&gt;The limitation I wrote then still stands and I would not soften it. A cleaner checkout does
not fix a weak offer. If people do not want the product, reducing friction only helps them
leave faster. That work paid because demand already existed and the form was in the way.&lt;/p&gt;
&lt;p&gt;What I have now that I did not have during those 30 minutes: a device metric that measures
behaviour rather than storage policy, a denominator written down before the test rather than
after it, one variable per variant, revenue per session for anything that touches a discount,
and a form specification instead of a form opinion. The result I reported is the least useful
thing I got out of that week.&lt;/p&gt;</content:encoded></item><item><title>Chain of thought is worth 0.7 points outside math and logic</title><link>https://therezaali.com/writing/three-prompting-techniques-measured/</link><guid isPermaLink="true">https://therezaali.com/writing/three-prompting-techniques-measured/</guid><pubDate>Thu, 06 Aug 2026 00:00:00 +0200</pubDate><author>contact@therezaali.com (Mo RezaAli)</author><category>essay</category><description>A meta-analysis over more than 100 papers: 12.3 points on math, 0.7 on everything else, 56.8 against 56.1. What that leaves of three techniques people still sell.</description><content:encoded>&lt;p&gt;Chain of thought buys 14.2 accuracy points on symbolic reasoning, 12.3 on math and 6.9 on
logical reasoning. On everything else it buys 0.7. That is 56.8 with it against 56.1 without,
averaged across every other category tested. In the authors&amp;rsquo; own words: &amp;ldquo;As much as 95% of the
total performance gain from CoT on MMLU is attributed to questions containing &amp;lsquo;=&amp;rsquo; in the question
or generated output.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Those figures come from
&lt;a href="https://arxiv.org/abs/2409.12183" class="external-link" rel="noopener"&gt;To CoT or not to CoT?&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; (Sprague et al., arXiv 2409.12183,
last revised 7 May 2025): &amp;ldquo;a quantitative meta-analysis covering over 100 papers using CoT&amp;rdquo; plus
the authors&amp;rsquo; own evaluations of &amp;ldquo;20 datasets across 14 models.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;I published a post recommending chain of thought for &amp;ldquo;complex reasoning, coding, math,
multi-step decision making.&amp;rdquo; Coding is not a category showing benefit in that meta-analysis, and
neither is multi-step decision making. The ratio between two items I listed as equivalent is 12.3
against 0.7, about 17.6x. Correcting that with the numbers attached is worth more than leaving
it standing.&lt;/p&gt;
&lt;h2 id="why-a-prompt-fails-in-a-way-a-better-model-cannot-fix"&gt;Why a prompt fails in a way a better model cannot fix&lt;/h2&gt;
&lt;p&gt;From 31 October 2025 to 27 April 2026 I ran a publishing bot against this domain. It made 842
&lt;code&gt;Add post:&lt;/code&gt; commits. It used &lt;code&gt;gpt-4o-mini&lt;/code&gt; with a &lt;code&gt;gpt-4o&lt;/code&gt; rewrite pass. Its entire input per
article was 450 characters of RSS &lt;code&gt;description&lt;/code&gt;, sliced in code. It had &lt;code&gt;AUTO_PR = &amp;quot;false&amp;quot;&lt;/code&gt;, so
nothing was reviewed before it hit &lt;code&gt;main&lt;/code&gt;. 180 of the 739 English posts cited a source called
&amp;ldquo;Internal Analysis&amp;rdquo; that does not exist.&lt;/p&gt;
&lt;p&gt;The prompt asked for an authoritative article with statistics and named sources, from 450
characters that contained neither. No model upgrade repairs that. The prompt specified an output
the input could not support, and the model did the only thing left: it filled the shape. The
lesson runs opposite to the folk version. Prompts matter more than model choice, but not because
a clever prompt adds capability. They matter because the prompt is where you decide what is being
asked for, and a badly specified ask fails identically at every model quality. The full account
of that pipeline is &lt;a href="https://therezaali.com/writing/the-bot-that-published-900-articles/"&gt;here&lt;/a&gt;. The commit counts come
from a private repository, so you have my word for them rather than a link.&lt;/p&gt;
&lt;h2 id="1-chain-of-thought-real-and-narrower-than-advertised"&gt;1. Chain of thought: real, and narrower than advertised&lt;/h2&gt;
&lt;p&gt;The technique has two origins and most writing conflates them.
&lt;a href="https://arxiv.org/abs/2201.11903" class="external-link" rel="noopener"&gt;Wei et al.&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; (arXiv 2201.11903, January 2022) named it, and
their method was few-shot: &amp;ldquo;a few chain of thought demonstrations are provided as exemplars in
prompting.&amp;rdquo;
&lt;a href="https://arxiv.org/abs/2205.11916" class="external-link" rel="noopener"&gt;Kojima et al.&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; (arXiv 2205.11916, May 2022) established the
zero-shot version and the phrase &amp;ldquo;Let&amp;rsquo;s think step by step.&amp;rdquo; What almost everyone actually does
is Kojima&amp;rsquo;s, under Wei&amp;rsquo;s name. The original gains are large and real.&lt;/p&gt;
&lt;p&gt;Table: Chain-of-thought accuracy on arithmetic benchmarks, as published in the two originating papers. Wei&amp;rsquo;s figures are read from Table 2 of the full text; deltas and multiples are my arithmetic from those published numbers.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Model&lt;/th&gt;
 &lt;th&gt;Benchmark&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Standard&lt;/th&gt;
 &lt;th style="text-align: right"&gt;With CoT&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Delta&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Multiple&lt;/th&gt;
 &lt;th&gt;Source&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;InstructGPT text-davinci-002&lt;/td&gt;
 &lt;td&gt;MultiArith&lt;/td&gt;
 &lt;td style="text-align: right"&gt;17.7%&lt;/td&gt;
 &lt;td style="text-align: right"&gt;78.7%&lt;/td&gt;
 &lt;td style="text-align: right"&gt;+61.0 pts&lt;/td&gt;
 &lt;td style="text-align: right"&gt;4.45x&lt;/td&gt;
 &lt;td&gt;Kojima 2022&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;InstructGPT text-davinci-002&lt;/td&gt;
 &lt;td&gt;GSM8K&lt;/td&gt;
 &lt;td style="text-align: right"&gt;10.4%&lt;/td&gt;
 &lt;td style="text-align: right"&gt;40.7%&lt;/td&gt;
 &lt;td style="text-align: right"&gt;+30.3 pts&lt;/td&gt;
 &lt;td style="text-align: right"&gt;3.91x&lt;/td&gt;
 &lt;td&gt;Kojima 2022&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;PaLM 540B&lt;/td&gt;
 &lt;td&gt;GSM8K&lt;/td&gt;
 &lt;td style="text-align: right"&gt;17.9%&lt;/td&gt;
 &lt;td style="text-align: right"&gt;56.9%&lt;/td&gt;
 &lt;td style="text-align: right"&gt;+39.0 pts&lt;/td&gt;
 &lt;td style="text-align: right"&gt;3.18x&lt;/td&gt;
 &lt;td&gt;Wei 2022&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;GPT-3 175B&lt;/td&gt;
 &lt;td&gt;GSM8K&lt;/td&gt;
 &lt;td style="text-align: right"&gt;15.6%&lt;/td&gt;
 &lt;td style="text-align: right"&gt;46.9%&lt;/td&gt;
 &lt;td style="text-align: right"&gt;+31.3 pts&lt;/td&gt;
 &lt;td style="text-align: right"&gt;3.01x&lt;/td&gt;
 &lt;td&gt;Wei 2022&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;LaMDA 137B&lt;/td&gt;
 &lt;td&gt;GSM8K&lt;/td&gt;
 &lt;td style="text-align: right"&gt;6.5%&lt;/td&gt;
 &lt;td style="text-align: right"&gt;14.3%&lt;/td&gt;
 &lt;td style="text-align: right"&gt;+7.8 pts&lt;/td&gt;
 &lt;td style="text-align: right"&gt;2.20x&lt;/td&gt;
 &lt;td&gt;Wei 2022&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Every benchmark in that table is arithmetic. Wei also stated a scale condition plainly: the
technique &amp;ldquo;does not positively impact performance for small models, and only yields performance
gains when used with models of ∼100B parameters.&amp;rdquo; That caveat has aged out, since anything a
reader touches in 2026 clears it. The task-category caveat has not.&lt;/p&gt;
&lt;p&gt;Table: What my post claimed chain of thought helps with, against what the Sprague meta-analysis measured.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Claim in the post&lt;/th&gt;
 &lt;th&gt;Measured average gain&lt;/th&gt;
 &lt;th&gt;Supported&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;Math&lt;/td&gt;
 &lt;td&gt;+12.3 pts math, +14.2 pts symbolic&lt;/td&gt;
 &lt;td&gt;Yes&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Complex reasoning&lt;/td&gt;
 &lt;td&gt;+6.9 pts logical reasoning; +0.7 pts for commonsense, knowledge and soft reasoning&lt;/td&gt;
 &lt;td&gt;Only the logical and symbolic subset&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Coding&lt;/td&gt;
 &lt;td&gt;Not a category showing benefit&lt;/td&gt;
 &lt;td&gt;No&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Multi-step decision making&lt;/td&gt;
 &lt;td&gt;Not a category showing benefit&lt;/td&gt;
 &lt;td&gt;No&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The floor is also not zero. &lt;a href="https://arxiv.org/abs/2410.21333" class="external-link" rel="noopener"&gt;Liu et al.&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; (arXiv 2410.21333)
took six tasks from cognitive psychology where deliberation degrades human performance and found
that &amp;ldquo;in three of these tasks, state-of-the-art models exhibit significant performance drop-offs
with CoT (up to 36.3% absolute accuracy for OpenAI o1-preview compared to GPT-4o).&amp;rdquo; Verbal
overshadowing has a machine analogue.&lt;/p&gt;
&lt;p&gt;Then the price. Chain of thought spends output tokens, the most expensive class on every sheet I
checked. On Anthropic&amp;rsquo;s current
&lt;a href="https://platform.claude.com/docs/en/about-claude/pricing" class="external-link" rel="noopener"&gt;price list&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;, output bills at exactly
5x input on every model in the table, from Haiku 4.5 at $1/$5 to Fable 5 at $10/$50. OpenAI&amp;rsquo;s
&lt;a href="https://developers.openai.com/api/docs/pricing" class="external-link" rel="noopener"&gt;sheet&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; does not hold to a single multiple across
its line. I am not going to quote rows off it, and the reason is the article&amp;rsquo;s own subject: that
page renders its table from JavaScript, an earlier draft of this paragraph quoted four
model-and-price pairs from it that are not on it at all, including two models it does not list, and
the error survived one round of checking. What matters for the argument holds without the
enumeration anyway. Output is the expensive class on every sheet I looked at, the multiple over
input is several times rather than a few percent, and it is not the same multiple everywhere, so
the cost of a reasoning preamble is a vendor-specific number you have to read off the current page
rather than a rule of thumb you can carry between them. A worked illustration
at Claude Sonnet 5&amp;rsquo;s introductory $10 per million output tokens: a 100-token direct answer costs
$0.0010, and the same answer behind a 900-token reasoning preamble costs $0.0100. At 100,000
calls a month, $100 against $1,000. Outside math and logic you are paying that $900 for 0.7
points. My own arithmetic from published list prices, not a bill anyone has received.&lt;/p&gt;
&lt;p&gt;Both vendors have now moved. OpenAI&amp;rsquo;s
&lt;a href="https://developers.openai.com/api/docs/guides/reasoning-best-practices" class="external-link" rel="noopener"&gt;reasoning best practices&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;
says outright: &amp;ldquo;Avoid chain-of-thought prompts: Since these models perform reasoning internally,
prompting them to &amp;rsquo;think step by step&amp;rsquo; or &amp;rsquo;explain your reasoning&amp;rsquo; is unnecessary.&amp;rdquo; Anthropic&amp;rsquo;s
&lt;a href="https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-prompting-best-practices" class="external-link" rel="noopener"&gt;prompting best practices&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;
file manual CoT under &amp;ldquo;Manual chain-of-thought (CoT) prompting as a fallback,&amp;rdquo; for when thinking
is switched off, and add that &amp;ldquo;a prompt like &amp;rsquo;think thoroughly&amp;rsquo; often produces better reasoning
than a hand-written step-by-step plan. Claude&amp;rsquo;s reasoning frequently exceeds what a human would
prescribe.&amp;rdquo; One vendor calls the instruction unnecessary, the other has demoted it to a fallback.
Neither is where the folk advice still sits.&lt;/p&gt;
&lt;h2 id="2-few-shot-three-to-five-and-they-are-not-teaching"&gt;2. Few-shot: three to five, and they are not teaching&lt;/h2&gt;
&lt;p&gt;My post said one or two examples. Anthropic&amp;rsquo;s current guidance says &amp;ldquo;Include 3–5 examples for
best results,&amp;rdquo; with examples wrapped in &lt;code&gt;&amp;lt;example&amp;gt;&lt;/code&gt; tags and required to be &amp;ldquo;Diverse: Cover edge
cases and vary enough that Claude doesn&amp;rsquo;t pick up unintended patterns.&amp;rdquo; Take the vendor&amp;rsquo;s number
over mine.&lt;/p&gt;
&lt;p&gt;Why diversity is the requirement exposes that the folk mechanism is wrong. I wrote that models
follow patterns better than vague instructions.
&lt;a href="https://arxiv.org/abs/2202.12837" class="external-link" rel="noopener"&gt;Min et al.&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; (arXiv 2202.12837) tested that and found that
randomly replacing the labels in demonstrations barely hurts performance, across 12 models
including GPT-3. What carries the effect is the label space, the distribution of the input text,
and the format of the sequence. Examples are not teaching the model correct answers. They are
declaring the shape and range of an acceptable response, which is why five lookalike examples do
less work than three that span the edges. Caveat: Min et al. tested classification and
multiple-choice, so do not carry &amp;ldquo;labels do not matter&amp;rdquo; into open-ended generation.&lt;/p&gt;
&lt;p&gt;Format is a live variable rather than a cosmetic one.
&lt;a href="https://arxiv.org/abs/2310.11324" class="external-link" rel="noopener"&gt;Sclar et al.&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; (arXiv 2310.11324) held semantic content fixed,
changed only separators, spacing and casing, and report &amp;ldquo;performance differences of up to 76
accuracy points when evaluated using LLaMA-2-13B.&amp;rdquo; Two things must travel with that number or it
misleads: LLaMA-2-13B is a 13B open-weight model from 2023, not a frontier model, and 76 points
is the maximum observed spread, not a typical one. Discounted for both, it still says separators
and casing are part of the prompt rather than packaging around it.&lt;/p&gt;
&lt;p&gt;Order is the failure people miss.
&lt;a href="https://arxiv.org/abs/2305.04388" class="external-link" rel="noopener"&gt;Turpin et al.&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; (arXiv 2305.04388) reordered multiple-choice
options in few-shot prompts so the correct answer was always &amp;ldquo;(A)&amp;rdquo;. Accuracy dropped by as much
as 36% across 13 BIG-Bench Hard tasks, tested on GPT-3.5 and Claude 1.0, and the models wrote
fluent reasoning that never mentioned the bias steering them. Your example set is an instruction
whether you intended it as one or not.&lt;/p&gt;
&lt;p&gt;OpenAI&amp;rsquo;s guidance on reasoning models, again verbatim: &amp;ldquo;Try zero shot first, then few shot if
needed.&amp;rdquo;&lt;/p&gt;
&lt;h2 id="3-role-prompting-register-not-correctness"&gt;3. Role prompting: register, not correctness&lt;/h2&gt;
&lt;p&gt;I claimed a role &amp;ldquo;changes how the model prioritizes information, explains concepts, and
structures its answers.&amp;rdquo; The first and third parts are supported. The implied fourth, that it
makes answers more correct, is not.&lt;/p&gt;
&lt;p&gt;Anthropic&amp;rsquo;s docs claim exactly the modest thing and no more: &amp;ldquo;Setting a role in the system prompt
focuses Claude&amp;rsquo;s behavior and tone for your use case.&amp;rdquo; No accuracy claim appears.
&lt;a href="https://arxiv.org/abs/2311.10054" class="external-link" rel="noopener"&gt;Zheng et al.&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; (arXiv 2311.10054) tested 162 roles across 4 LLM
families on 2,410 factual questions and found that &amp;ldquo;adding personas in system prompts does not
improve model performance across a range of questions compared to the control setting where no
persona is added,&amp;rdquo; and that where gains appear, &amp;ldquo;the effect of each persona can be largely
random.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;A second paper disagrees, and suppressing that would be dishonest.
&lt;a href="https://arxiv.org/abs/2308.07702" class="external-link" rel="noopener"&gt;Kong et al.&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; (arXiv 2308.07702) report role-play gains across
twelve reasoning benchmarks with ChatGPT, including AQuA moving 53.5% to 63.8% and Last Letter
moving 23.8% to 84.2%. Their explanation is the part that settles it for me: role-play &amp;ldquo;acts as a
more effective trigger for the CoT process.&amp;rdquo; If that is right, role prompting is an oblique way
of invoking technique one rather than a third independent technique, which predicts it adds
nothing on a model with reasoning already on. Zheng measured factual accuracy and found no
reliable gain; Kong measured reasoning tasks and found gains they attribute to reasoning
elicitation. A role reliably changes register and unreliably changes correctness, so use it for
voice and audience and never for truth.&lt;/p&gt;
&lt;p&gt;Table: What each technique costs and what it reliably buys. Prices are Anthropic list prices for Claude Sonnet 5, read 6 August 2026; the effect column is the measured average from the papers cited above.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Technique&lt;/th&gt;
 &lt;th&gt;Token class&lt;/th&gt;
 &lt;th&gt;Cacheable&lt;/th&gt;
 &lt;th&gt;What it reliably buys&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;Few-shot examples&lt;/td&gt;
 &lt;td&gt;Input, $2/MTok&lt;/td&gt;
 &lt;td&gt;Yes, cache reads bill at 0.1x base&lt;/td&gt;
 &lt;td&gt;Output format and output space&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Chain of thought&lt;/td&gt;
 &lt;td&gt;Output, $10/MTok, 5x input&lt;/td&gt;
 &lt;td&gt;No, regenerated every call&lt;/td&gt;
 &lt;td&gt;+12.3 pts on math, +0.7 elsewhere&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Role&lt;/td&gt;
 &lt;td&gt;Negligible&lt;/td&gt;
 &lt;td&gt;Yes, it sits in the system prompt&lt;/td&gt;
 &lt;td&gt;Tone, register, structure&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;They have opposite cost structures and my post presented them as interchangeable. Five 200-token
examples cost $0.0020 per call at base input price, or $0.0002 on a cache read, a 10x reduction
that repays the 1.25x write premium inside a single reuse. Anthropic&amp;rsquo;s docs reach the same
conclusion independently: &amp;ldquo;caching pays off after just one cache read for the 5-minute duration.&amp;rdquo;
The cheap technique is the one with the most reliable effect.&lt;/p&gt;
&lt;h2 id="how-to-detect-this-in-your-own-system"&gt;How to detect this in your own system&lt;/h2&gt;
&lt;p&gt;Four checks, none needing a benchmark harness.&lt;/p&gt;
&lt;p&gt;Look for an equals sign. If your task does not involve symbolic manipulation, arithmetic or
formal logic, the Sprague number for it is 0.7 points, and a &amp;ldquo;think step by step&amp;rdquo; instruction in
a copywriting, summarisation or classification prompt is spend with no documented return.&lt;/p&gt;
&lt;p&gt;Ask whether a human expert would do the task better on instinct than on deliberation. Liu et al.
found the drops clustered on tasks with exactly that property.&lt;/p&gt;
&lt;p&gt;Stop treating reasoning traces as an audit log. Turpin showed models rationalising a bias they
never named. Anthropic&amp;rsquo;s measurement of 3 April 2025 found Claude 3.7 Sonnet mentioned an
inserted hint in its reasoning 25% of the time and DeepSeek R1 39%. In a setup where models were rewarded for
following deliberately wrong hints, they took the hint &amp;ldquo;in over 99% of cases&amp;rdquo; and admitted to it
&amp;ldquo;less than 2% of the time in most of the testing scenarios.&amp;rdquo; Outcome-based reinforcement learning
on an earlier snapshot of Claude 3.7 Sonnet raised faithfulness and then plateaued, at 28% on
MMLU and 20% on GPQA. If you are logging chain of thought for compliance, you are logging a
plausible story about the computation.&lt;/p&gt;
&lt;p&gt;Last, check that the prompt&amp;rsquo;s demands are proportionate to its input. The bot failed because it
asked for sourced statistics from 450 characters that had none, and no technique on this page
would have saved it. List every fact the output requires and confirm each is present in the
input. Anything absent will be generated.&lt;/p&gt;
&lt;h2 id="what-i-have-now-that-i-did-not-have-before"&gt;What I have now that I did not have before&lt;/h2&gt;
&lt;p&gt;A decision rule instead of a habit. Chain of thought when the task contains symbols, math or
formal logic, and not otherwise. Examples in the published range, three to five, chosen for
spread rather than similarity, cached. Roles for voice and never for accuracy.&lt;/p&gt;
&lt;p&gt;Both vendor pages cited here are written against whatever models are current, so the rule needs
re-reading rather than remembering. One trap documented right now: with extended thinking
disabled, &amp;ldquo;Claude Opus 4.5 is particularly sensitive to the word &amp;rsquo;think&amp;rsquo; and its variants,&amp;rdquo; and
the docs suggest &amp;ldquo;consider,&amp;rdquo; &amp;ldquo;evaluate,&amp;rdquo; or &amp;ldquo;reason through&amp;rdquo; instead. And three techniques is not
the set. Schulhoff et al.&amp;rsquo;s &lt;a href="https://arxiv.org/abs/2406.06608" class="external-link" rel="noopener"&gt;systematic survey&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; catalogues 58
text-based prompting techniques and defines 33 vocabulary terms.&lt;/p&gt;
&lt;p&gt;The post I am correcting was not wrong about 2022. It was describing 2022. These techniques were
genuinely transformative when published, and the literature that narrowed them arrived later and
got far less attention. Reading the primary source is the only defence, and it takes an afternoon.&lt;/p&gt;</content:encoded></item><item><title>Automate up to the decision, and never the decision itself</title><link>https://therezaali.com/writing/automating-up-to-the-decision/</link><guid isPermaLink="true">https://therezaali.com/writing/automating-up-to-the-decision/</guid><pubDate>Thu, 06 Aug 2026 00:00:00 +0200</pubDate><author>contact@therezaali.com (Mo RezaAli)</author><category>essay</category><description>X’s ranker exposes nineteen prediction heads. None is an impression, four are negative, and net-negative posts leave through a different branch. Three automations I killed, eleven I kept.</description><content:encoded>&lt;p&gt;X&amp;rsquo;s open-source ranking model exposes nineteen prediction heads. None of them is an impression.
Four of them are negative feedback, and the live Rust scorer does something with those four that
is stranger than the folklore. A post whose weighted score comes out net-negative is not ranked
low on the same scale as everything else. It leaves through a different branch, gets rescaled by
the ratio of the negative weights to the total, and is then multiplied by a constant called
&lt;code&gt;NEGATIVE_SCORES_OFFSET&lt;/code&gt; that the repository imports and never defines. If that constant is
positive, every net-negative post lands in a band underneath every net-positive one.&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-rust" data-lang="rust"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;// home-mixer/scorers/ranking_scorer.rs:83
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;&lt;/span&gt;&lt;span class="kd"&gt;let&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;negative_sum&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;not_interested&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;+&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;block_author&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;+&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;mute_author&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;+&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;report&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;+&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;not_dwelled&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt;&lt;/span&gt;&lt;span class="kd"&gt;let&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;total_sum&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;positive_sum&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;+&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;negative_sum&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt;&lt;/span&gt;&lt;span class="c1"&gt;// home-mixer/scorers/ranking_scorer.rs:175
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;&lt;/span&gt;&lt;span class="k"&gt;fn&lt;/span&gt; &lt;span class="nf"&gt;offset_score&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;combined_score&lt;/span&gt;: &lt;span class="kt"&gt;f64&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;w&lt;/span&gt;: &lt;span class="kp"&gt;&amp;amp;&lt;/span&gt;&lt;span class="nc"&gt;ScoringWeights&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;-&amp;gt; &lt;span class="kt"&gt;f64&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="k"&gt;if&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;w&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;total_sum&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;==&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.0&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;combined_score&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;max&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mf"&gt;0.0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="k"&gt;if&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;combined_score&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.0&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;combined_score&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;+&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;w&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;negative_sum&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;w&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;total_sum&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="no"&gt;NEGATIVE_SCORES_OFFSET&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="k"&gt;else&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;combined_score&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;+&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="no"&gt;NEGATIVE_SCORES_OFFSET&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt;&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;I opened that file to work out whether an automated welcome DM was worth its click rate. It was
not. The click rate turned out to be the least interesting reason, and the most interesting one
had been written down years before I ran it.&lt;/p&gt;
&lt;h2 id="what-i-was-running-and-the-line-that-started-this"&gt;What I was running, and the line that started this&lt;/h2&gt;
&lt;p&gt;I run cold sequences for 1688 sourcing at Yakkyo S.p.A. One of them got a one-line reply: &amp;ldquo;Is this
a real person or a bot?&amp;rdquo;&lt;/p&gt;
&lt;p&gt;The email was technically perfect. Right name, right company, a reference to their product
category that &lt;a href="https://www.clay.com/" class="external-link" rel="noopener"&gt;Clay&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; had pulled and Claude had stitched into a clean
opening. It read like nobody was home. Reply rate had been sliding for six weeks and I had been
blaming the copy.&lt;/p&gt;
&lt;p&gt;I take that question seriously rather than laughing it off, for a specific reason. From 31 October
2025 to 27 April 2026 I ran a publishing bot against this domain. It made 842 &lt;code&gt;Add post:&lt;/code&gt; commits
and produced 2,922 machine-generated files before I deleted them; of the 739 English ones, 180
cited a source called &amp;ldquo;Internal Analysis&amp;rdquo; that does not exist. The full account is
&lt;a href="https://therezaali.com/writing/the-bot-that-published-900-articles/"&gt;here&lt;/a&gt;, and those counts come from a private
repository, so you have my word rather than a link.&lt;/p&gt;
&lt;p&gt;Same shape, different scale: an automation took over a judgment call, kept producing output that
looked fine, and nothing alerted because nothing broke. So this quarter I killed automations
instead of stacking them. Three went, eleven stayed. Every figure below for my own results is my
own unaudited number from my own work, published first
&lt;a href="https://x.com/Mo_ali/status/2082123761755541597" class="external-link" rel="noopener"&gt;on X&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="kill-1-the-fully-automated-cold-sequence"&gt;Kill 1: the fully automated cold sequence&lt;/h2&gt;
&lt;p&gt;Clay enriched the lead, an &lt;a href="https://n8n.io/" class="external-link" rel="noopener"&gt;n8n&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; workflow scored it, Claude wrote a three-email
sequence with a first line pulled from the prospect&amp;rsquo;s site, and it sent on a schedule. I forgot
about it for about two months.&lt;/p&gt;
&lt;p&gt;Reply rate fell from roughly 9 percent to under 4. What diagnosed it was not the average: positive
replies fell faster than total replies, while &amp;ldquo;unsubscribe&amp;rdquo; and &amp;ldquo;not relevant&amp;rdquo; kept arriving at the
same clip. The sequence was filtering out the people most worth talking to and retaining the ones
who were going to say no anyway.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Why it happened.&lt;/strong&gt; Automating a first line does not scale the good version of that line. It
scales the mean of every line the model would write, and that mean is competent and
interchangeable. Cold outreach does not pay on the mean. It pays on the tail, where one specific
observation lands because somebody looked. Taking the human out of that step does not lower
quality evenly. It removes variance, and the variance was the product.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;How to detect it.&lt;/strong&gt; Reply rate sums over both tails, so the drift hides inside it. Split it.
Positive and negative replies are separate series that move independently, and their divergence
leads the average by weeks.&lt;/p&gt;
&lt;p&gt;Then watch what your provider watches. Google&amp;rsquo;s sender guidelines are specific and almost everyone
quotes them wrong. The number people repeat is 0.3 percent; the guidance says &amp;ldquo;Keep spam rates
reported in Postmaster Tools below 0.10% and avoid ever reaching a spam rate of 0.30% or higher.&amp;rdquo;
Two thresholds, two jobs. 0.10 percent is the target you sit under, 0.30 percent is the line you
never touch, and treating 0.30 as your budget spends your margin before the bad week arrives.&lt;/p&gt;
&lt;p&gt;Making that measurable needs a clean opt-out channel:
&lt;a href="https://www.rfc-editor.org/rfc/rfc8058.html" class="external-link" rel="noopener"&gt;RFC 8058&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; one-click unsubscribe, the
&lt;code&gt;List-Unsubscribe-Post: List-Unsubscribe=One-Click&lt;/code&gt; header Google names directly. Without it, an
annoyed recipient&amp;rsquo;s cheapest exit is the spam button, which routes the signal into your reputation
instead of your dashboard. With it, opt-outs are a dated count you can line up against the day an
automation went live. For US recipients that channel is also statutory: 15 U.S.C.
7704(a)(4)(A)(i) makes it unlawful to send more than 10 business days after an opt-out request.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What runs now.&lt;/strong&gt; Enrichment and scoring stayed automated. That genuinely took two hours and now
takes about eleven minutes. I am leaving that unrounded and unconverted, because a percentage
computed off an &amp;ldquo;about&amp;rdquo; would carry a decimal place I never measured. Claude drafts the full
email in maybe 20 minutes and I rewrite the opening of every tier-one email by hand. Reply rate
came back to around 8 percent, and the positive-reply mix is better than it ever was on full auto.
The judgment moved back to a human. The typing stayed automated.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;A gate that did not exist when I built this.&lt;/strong&gt; Article 50 of the EU AI Act, Regulation (EU)
2024/1689, entered into force on 2 August 2026, and I work in Messina, so it reaches me. It is also
being widely misreported as a blanket duty to label AI-written email. Article 50(1) puts the
interaction-disclosure duty on &lt;em&gt;providers&lt;/em&gt; of systems intended to interact directly with natural
persons. Article 50(4)&amp;rsquo;s deployer duty on text covers material &amp;ldquo;published with the purpose of
informing the public on matters of public interest&amp;rdquo; and carves out anything that &amp;ldquo;has undergone a
process of human review or editorial control&amp;rdquo;. A cold email drafted by a model and rewritten by me
fits neither well. Read it before you panic or dismiss it: &amp;ldquo;real person or a bot?&amp;rdquo; now has a legal
surface beside the craft one.&lt;/p&gt;
&lt;h2 id="kill-2-the-automated-dm-to-every-new-follower"&gt;Kill 2: the automated DM to every new follower&lt;/h2&gt;
&lt;p&gt;Every new follower on X got an automated welcome DM with a link. A webhook fired the message. It is
the growth tactic everyone copies.&lt;/p&gt;
&lt;p&gt;The DMs converted to clicks at well under one percent, and mute and block events on the account
ticked up in the exact window I turned it on. So I killed it on the numbers. Then I read the policy
and found the numbers had been the second-best reason all along.&lt;/p&gt;
&lt;p&gt;X&amp;rsquo;s Developer Agreement and Policy, under &amp;ldquo;Spam, bots, and automation&amp;rdquo;, says that &amp;ldquo;Services that
perform write actions, including posting Posts, following accounts, or sending Direct Messages,
must follow the Automation Rules. In particular, you should:&amp;rdquo;. The first bullet under that colon
is:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Always get explicit consent before sending people automated replies or Direct Messages&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;Read the modal verbs, because they are doing two different jobs. The must points at a separate
document: that bolded sentence is a link to &lt;code&gt;help.x.com/rules-and-policies/x-automation&lt;/code&gt;, which
answers 403, with a Chrome user agent as readily as without one, so I have not read the Automation
Rules and will not quote them. The consent line is a should, and it is X writing out in its own
words what it expects compliance to look like. A follow is not explicit consent on either reading.
Nobody who clicked that button agreed to receive anything. If the automation authenticated through
the X API, it ran its whole life against the one instruction X spelled out on a page I can fetch,
independent of how it performed. The same section adds a line that reads differently after the
cold-email story above: &amp;ldquo;You should never mislead or confuse people about whether your account is
or is not a bot.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;One boundary. That section governs write actions through the X API and developer products. An
automation driving a browser falls under the X Rules on platform manipulation instead, which is
&lt;code&gt;help.x.com/rules-and-policies/platform-manipulation&lt;/code&gt; and also 403s, so the same non-quotation
applies there.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The mechanism, folklore removed.&lt;/strong&gt; Growth posts repeat a table of weights: reply-back plus 75,
mute or block minus 74, report minus 369. Those constants appear nowhere in either public
repository. What is published is a demo fusion in &lt;code&gt;phoenix/run_pipeline.py&lt;/code&gt; at lines 355-360,
containing four heads, all positive:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-python" data-lang="python"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;weighted&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt; &lt;span class="n"&gt;all_probs&lt;/span&gt;&lt;span class="p"&gt;[:,&lt;/span&gt; &lt;span class="n"&gt;IDX_FAV&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mf"&gt;1.0&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;all_probs&lt;/span&gt;&lt;span class="p"&gt;[:,&lt;/span&gt; &lt;span class="n"&gt;IDX_REPLY&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mf"&gt;0.5&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;all_probs&lt;/span&gt;&lt;span class="p"&gt;[:,&lt;/span&gt; &lt;span class="n"&gt;IDX_RT&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mf"&gt;0.3&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;all_probs&lt;/span&gt;&lt;span class="p"&gt;[:,&lt;/span&gt; &lt;span class="n"&gt;IDX_DWELL&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mf"&gt;0.2&lt;/span&gt; &lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Production weights resolve at request time from a feature-switch service that is not vendored. No
negative weight has ever been published, so nobody, me included, can honestly say what the system
punishes hardest. The same gap covers &lt;code&gt;NEGATIVE_SCORES_OFFSET&lt;/code&gt; from the opening code fence:
&lt;code&gt;ranking_scorer.rs&lt;/code&gt; pulls it in with &lt;code&gt;use crate::params::*&lt;/code&gt; and the published tree has no params
file. I pulled the whole repository listing, 244 paths, to check that. The branch is verifiable and
the magnitude is not, which is the honest state of every weight discussed in this section.&lt;/p&gt;
&lt;p&gt;What is verifiable is better than a weight, because it is structural rather than tunable.
&lt;code&gt;phoenix/runners.py&lt;/code&gt; declares nineteen heads at lines 233-252 and, at lines 266-271,
&lt;code&gt;NEGATIVE_FEEDBACK_INDICES = [14, 15, 16, 17]&lt;/code&gt;: &lt;code&gt;not_interested_score&lt;/code&gt;, &lt;code&gt;block_author_score&lt;/code&gt;,
&lt;code&gt;mute_author_score&lt;/code&gt;, &lt;code&gt;report_score&lt;/code&gt;. &amp;ldquo;Show less of this&amp;rdquo; is &lt;code&gt;not_interested_score&lt;/code&gt;. Each is a
separately modelled prediction, and &lt;code&gt;offset_score&lt;/code&gt; sends any candidate that nets out below zero
down the separate rescaled branch rather than the shifted one.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The account-level claim, corrected.&lt;/strong&gt; I used to say negative feedback drags down how the whole
account gets treated. The outcome is roughly right and the mechanism I described does not exist.
There is no account penalty ledger in the open source. &lt;code&gt;block_author_score&lt;/code&gt; and &lt;code&gt;mute_author_score&lt;/code&gt;
are per-candidate, per-viewer predictions that &lt;em&gt;this&lt;/em&gt; viewer would block or mute you over &lt;em&gt;this&lt;/em&gt;
post. The genuine account-level effect is blunter and permanent: &lt;code&gt;author_socialgraph_filter.rs&lt;/code&gt;
removes your post from a viewer&amp;rsquo;s candidate set outright once that viewer has muted or blocked you,
and it propagates through quotes and retweets. One mute and you are gone from that feed at any
score, and no amount of later quality recovers it.&lt;/p&gt;
&lt;p&gt;There is also an asymmetry worth knowing, one line in each of two files. Scoring history defaults to
&lt;code&gt;UserActionAggregationType::DenseWithNotInterestedIn&lt;/code&gt;
(&lt;code&gt;scoring_sequence_query_hydrator.rs:47&lt;/code&gt;). Retrieval history defaults to plain &lt;code&gt;Dense&lt;/code&gt;
(&lt;code&gt;retrieval_sequence_query_hydrator.rs:50&lt;/code&gt;). Explicit negative feedback is modelled by default where
posts are scored, and not by default where they are fetched.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;And the part that kills the tactic on its own terms.&lt;/strong&gt; Outbound DMs are not a ranking input at
all. The only DM among the nineteen heads is &lt;code&gt;share_via_dm_score&lt;/code&gt; at index 8, a positive that
models somebody sharing &lt;em&gt;your post&lt;/em&gt; via DM. Sending DMs earns nothing. The harm route was mute and
block plus a policy breach; the upside route did not exist.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What runs now.&lt;/strong&gt; Nothing automated. I reply to people&amp;rsquo;s posts, and when someone follows after a
good exchange I send a real DM by hand, maybe one in ten. Slower, and those conversations go
somewhere. X API access has since moved to pay-per-usage pricing, so the tier names and prices in
every guide written about this tactic are now wrong as well. Read the policy first, the pricing
page second.&lt;/p&gt;
&lt;h2 id="kill-3-the-ai-generated-weekly-content-batch"&gt;Kill 3: the AI-generated weekly content batch&lt;/h2&gt;
&lt;p&gt;Every Monday a workflow generated five posts and a thread from that week&amp;rsquo;s themes and queued them.
Five minutes of work for a week of content. Impressions held steady and replies cratered.&lt;/p&gt;
&lt;p&gt;My explanation for that was wrong, and the correct one is more useful. I said the cheap signal on X
is the impression and the expensive one is the reply. There is no impression head. The
nineteen-entry &lt;code&gt;ACTIONS&lt;/code&gt; list contains no impression at all; raw counts are telemetry, not model
inputs. Nor can I claim a reply outranks a like, because in the only fusion X has published a reply
carries 0.5 against a favorite at 1.0.&lt;/p&gt;
&lt;p&gt;What I was describing is dwell, which is modelled harder than anything else in the list:
&lt;code&gt;dwell_score&lt;/code&gt; at index 10, a discrete &lt;code&gt;dwell_time&lt;/code&gt; at index 18, and again in the continuous head
block. Then the live Rust path adds a head that is not in the nineteen-head list at all,
&lt;code&gt;not_dwelled&lt;/code&gt;, sitting inside &lt;code&gt;negative_sum&lt;/code&gt; next to block, mute and report. Being seen and scrolled
past is directly and negatively scored. That is what generated content produces. It reads fine and
it skims, and skimming is the failure the model is instrumented to catch.&lt;/p&gt;
&lt;p&gt;Corroboration sits one level up, in the content screen rather than the ranker.
&lt;code&gt;grox/classifiers/content/banger_initial_screen.py&lt;/code&gt; runs a Grok vision model over posts and returns
a result whose fields include &lt;code&gt;quality_score&lt;/code&gt; and, at line 37, &lt;code&gt;slop_score: int | None&lt;/code&gt;. Somebody at
X thought machine-generated filler deserved its own integer.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;What runs now.&lt;/strong&gt; AI writes a first draft in 20 minutes, I add my voice and the part I am genuinely
unsure about, and I post three times a week instead of five. Fewer posts, more replies. The publish
test is whether I would be slightly nervous to post it. No flinch, no reach.&lt;/p&gt;
&lt;h2 id="the-statistic-everyone-cites-about-this-and-what-it-says"&gt;The statistic everyone cites about this, and what it says&lt;/h2&gt;
&lt;p&gt;A 2025 MIT report gets quoted constantly for the claim that 95 percent of AI pilots never reach
measurable return. I have quoted it that way myself. Having now read all 26 pages, the popular
version is overstated on four axes, and the correction is worth more than the statistic was.&lt;/p&gt;
&lt;p&gt;Table: What the MIT NANDA report actually is, against how it is usually cited.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Usually cited as&lt;/th&gt;
 &lt;th&gt;What the document says&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;An MIT study&lt;/td&gt;
 &lt;td&gt;&amp;ldquo;Preliminary Findings from AI Implementation Research from Project NANDA&amp;rdquo;, July 2025. Its sole listed reviewer is also a listed author, and it states the views &amp;ldquo;do not reflect the positions of any affiliated employers&amp;rdquo;&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;A study of 300 companies&lt;/td&gt;
 &lt;td&gt;52 structured interviews, 153 senior leaders surveyed &amp;ldquo;across four major industry conferences&amp;rdquo;, and a review of 300-plus publicly disclosed initiatives&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;95 percent of AI pilots fail&lt;/td&gt;
 &lt;td&gt;5 percent success for embedded, task-specific GenAI tools. The same funnel gives 40 percent for general-purpose LLMs, and the same section reports &amp;ldquo;Generic LLM chatbots appear to show high pilot-to-implementation rates (~83%)&amp;rdquo;&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Never reaches measurable return&lt;/td&gt;
 &lt;td&gt;Success is &amp;ldquo;deployment beyond pilot phase with measurable KPIs&amp;rdquo; at six months, defined by what &amp;ldquo;users or executives have remarked&amp;rdquo;. The authors write that six months &amp;ldquo;may be insufficient&amp;hellip; potentially understating success rates&amp;rdquo;&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The 95 is one hundred minus five, taken off the branch of a two-branch funnel that excludes the
tools most people actually use. Quoting it as a verdict on AI means quoting the number that
specifically is not about ChatGPT or Copilot.&lt;/p&gt;
&lt;p&gt;One more checkable fact I have not seen mentioned: the report is no longer served from MIT. The
Internet Archive&amp;rsquo;s index for &lt;code&gt;nanda.media.mit.edu/ai_report_2025.pdf&lt;/code&gt; records four 200s on 18
August 2025, at 11:55:20, 12:50:35, 12:54:28 and 14:57:14 UTC, then a 404 at 17:28:56 UTC the same
day. Its URL now lands on a Media Lab group overview page. The most-quoted AI statistic of 2025 has
not been hosted at its published address since that afternoon.&lt;/p&gt;
&lt;p&gt;One thing to get right if you check that index yourself: the length column is not the size of the
PDF. It reads 886,480 on the first capture, 867,304, 886,477 and
865,562 on the next three, all four with the same content digest, because it measures the archive&amp;rsquo;s
own compressed record. I downloaded the archived file. It is 923,623 bytes, and it is the same file
in every capture.&lt;/p&gt;
&lt;p&gt;My reading of it is that most people automate the judgment rather than the grunt work. That is my
opinion, not a finding. Its nearest actual finding is budget misallocation, and it contradicts
itself inside one section: the takeaway line says 50 percent of GenAI budgets go to sales and
marketing while the body of that same section says &amp;ldquo;approximately 70 percent&amp;rdquo;. Cite it if you like,
but say which number you took and that the source gives two.&lt;/p&gt;
&lt;h2 id="the-rule-and-the-paper-that-named-it-in-2000"&gt;The rule, and the paper that named it in 2000&lt;/h2&gt;
&lt;p&gt;The rule I run now: automate up to the decision, and after the decision, never the decision itself.
That is a rediscovery, and the original is more precise than my phrasing.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://pubmed.ncbi.nlm.nih.gov/11760769/" class="external-link" rel="noopener"&gt;Parasuraman, Sheridan and Wickens&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; (&lt;em&gt;IEEE Transactions
on Systems, Man, and Cybernetics Part A&lt;/em&gt; 30(3):286-297, May 2000) model automation as applying to
four independent stages: information acquisition, information analysis, decision and action
selection, and action implementation. Level is chosen per stage. Their ten-point scale applies to
stage 3 specifically, running from a computer offering several options, through level 4 where it
suggests one alternative and the human retains authority, through level 6 where the human gets a
limited window to veto, up to level 10 where the computer decides and acts alone.&lt;/p&gt;
&lt;p&gt;Every automation I killed sat between levels 7 and 10 on stage 3. Every replacement sits at level 4.
Enrichment and scoring are stages 1 and 2, and I automated those harder rather than less. That is
what &amp;ldquo;buy the boring stuff, build the edge&amp;rdquo; means once you say it precisely.&lt;/p&gt;
&lt;p&gt;Running stage 3 too high has its own named failure.
&lt;a href="https://api.crossref.org/works/10.1518/001872095779064555" class="external-link" rel="noopener"&gt;Endsley and Kiris&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; (&lt;em&gt;Human Factors&lt;/em&gt;
37(2):381-394, 1995) called it the out-of-the-loop performance problem: an operator supervising an
automated decision loses the situational awareness needed to take over when it degrades. I lost six
weeks to that. I was not monitoring the sequence. I was receiving its output.&lt;/p&gt;
&lt;h2 id="the-quarterly-check-and-what-each-trip-means"&gt;The quarterly check, and what each trip means&lt;/h2&gt;
&lt;p&gt;Every automation goes through this once a quarter. Two trips and it is killed or pulled back to
level 4.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;A quality signal falling while a volume signal holds.&lt;/strong&gt; Track positive and negative replies as
separate series; their divergence leads the average.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;A negative signal rising since launch.&lt;/strong&gt; Unsubscribes, mutes, blocks, &amp;ldquo;not relevant&amp;rdquo;,
Postmaster spam rate against 0.10 percent rather than 0.30. Implement RFC 8058 so opt-outs arrive
as a dated count rather than as reputation damage.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;It is making a call a human used to make.&lt;/strong&gt; That is stage 3. Name the level it runs at; above 4
is the finding.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;I would be embarrassed if the recipient knew.&lt;/strong&gt; The bot test, now also a compliance question in
the EU and a policy question under any developer terms you have accepted.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;It could break for three days without alerting me.&lt;/strong&gt; If you would learn from a metric rather
than an alarm, you are out of the loop in the Endsley sense.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;The time saved is smaller than the trust it costs.&lt;/strong&gt; Two hours down to eleven minutes with no
quality cost, keep forever. Five minutes saved for a point of reply rate, kill it.&lt;/li&gt;
&lt;/ol&gt;
&lt;h2 id="what-i-have-now-that-i-did-not-have-in-april"&gt;What I have now that I did not have in April&lt;/h2&gt;
&lt;p&gt;Four things, and only one is time.&lt;/p&gt;
&lt;p&gt;A vocabulary with citations behind it, so &amp;ldquo;automate the boring stuff&amp;rdquo; becomes &amp;ldquo;keep stage 3 at level
4 and push stages 1, 2 and 4 as high as they go&amp;rdquo;. A compliance gate that runs before the ROI gate,
because kill 2 was on the wrong side of the policy before it was on the wrong side of the numbers,
and I checked those in the wrong order. A
detection method that watches divergence between two series rather than the level of one. And a
reading of the ranker built from files rather than from other people&amp;rsquo;s posts, which is why the
version above corrects three things I used to say.&lt;/p&gt;
&lt;p&gt;Two numbers in my original post do not reconcile, and I would rather fix that here than let it
sit. I killed three automations and kept eleven. I also said most funnels are over-automated by
about a third. Those are different claims and neither confirms the other: three of fourteen is a
count of things I run, and the one-third is a guess about other people&amp;rsquo;s funnels. Read the count as
a count and the guess as a guess.&lt;/p&gt;
&lt;p&gt;The hours I got back did not go into building more automations. They went into the manual first
lines, the real DMs, and the posts I am slightly nervous about. Those are the parts that do not
scale, which keeps turning out to be the same list as the parts that work.&lt;/p&gt;</content:encoded></item><item><title>An always-on laptop draws 3 W to 7 W, which is £7 to £16 of electricity a year</title><link>https://therezaali.com/writing/always-on-agent-costs/</link><guid isPermaLink="true">https://therezaali.com/writing/always-on-agent-costs/</guid><pubDate>Thu, 06 Aug 2026 00:00:00 +0200</pubDate><author>contact@therezaali.com (Mo RezaAli)</author><category>reference</category><description>The guess I inherited was 10 W to 30 W. The certified data puts it at 3 W to 7 W: £7 to £16 a year, four or five dollars of model spend, and $0 in licences at this scale.</description><content:encoded>&lt;p&gt;Every automation I run only runs while I am sitting in front of it. I fill the
form, the steps go green, I approve the output, I close the laptop, and the print
floor goes dark. The factory only exists while I am standing in it. So the work
that did not happen last week wasn&amp;rsquo;t blocked by money or by ideas. It was blocked
by me being asleep, at a desk, or on a train.&lt;/p&gt;
&lt;p&gt;The obvious repair is a machine that keeps working when I am not there, and the
internet is not short of guides telling you how to install one. What almost
nobody publishes is the bill. How many watts, at whose unit rate, for how many
hours. How many cents per scheduled run, at which per-token price. Which vendor
licence thresholds bite a business and leave a hobbyist alone. And, the part that
surprised me most, which spending caps actually exist rather than which ones
people assume exist.&lt;/p&gt;
&lt;p&gt;So this page is two documents: the money and the reasons to walk away, then the
build in eight parts, every command carrying the line that tells you whether it
worked. What I did was write the specification, price every component against a
primary source, and do the arithmetic before spending anything. I have not run it
at all yet. Nothing below is a measurement from my own machine, because there is
not yet a machine to measure.&lt;/p&gt;
&lt;h2 id="what-it-would-be-and-who-should-close-this-tab"&gt;What it would be, and who should close this tab&lt;/h2&gt;
&lt;p&gt;The design calls for a Windows laptop that stays home, plugged in, with sleep and
lid-close-to-sleep switched off, because a sleeping computer does nothing at
03:00. The only wallet in the system would be a prepaid
&lt;a href="https://openrouter.ai/terms" class="external-link" rel="noopener"&gt;OpenRouter&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; account funded with a single small
top-up. The interface would be a chat page served by
&lt;a href="https://docs.openwebui.com/getting-started/quick-start/" class="external-link" rel="noopener"&gt;Open WebUI&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; in a Docker
container on that laptop, reachable from a phone over
&lt;a href="https://tailscale.com/pricing" class="external-link" rel="noopener"&gt;Tailscale&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;&amp;rsquo;s private network rather than through
a hole punched in a router. And the worker would be
&lt;a href="https://hermes-agent.nousresearch.com/docs/getting-started/platform-support" class="external-link" rel="noopener"&gt;Hermes Agent&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;,
holding a list of scheduled jobs, answering a
&lt;a href="https://core.telegram.org/bots" class="external-link" rel="noopener"&gt;Telegram&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; bot allowlisted to exactly one
account, and starting itself again after a reboot. One job would fire early and
leave a market brief on a phone before its owner is awake; the bot would make the
machine reachable from anywhere, answering exactly one person; and the paid
models anyone would use in a commercial AI chat would run on a page served from a
laptop you own.&lt;/p&gt;
&lt;p&gt;Four installs, one small top-up, and no hole in the router. One inbound rule does
get added to the laptop&amp;rsquo;s own Windows firewall in Part 6, scoped to private
networks, and I would rather name that here than let this page claim nothing is
ever opened while Part 6 tells you to open something.&lt;/p&gt;
&lt;p&gt;The disqualifiers come before the instructions, because the cheapest possible
outcome of reading this is that you spend five minutes and decide against it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Your only laptop travels with you.&lt;/strong&gt; The machine has to stay home and stay on.
What still applies to you is the cost arithmetic, which holds for any hosted
model you already pay for.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;You have an Intel Mac.&lt;/strong&gt; The worker software doesn&amp;rsquo;t support it. Apple Silicon
Macs can run this, with a caveat covered further down.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The machine has less than 8 GB of RAM, or runs an older Windows.&lt;/strong&gt; Below
Windows 10 22H2 or Windows 11 23H2 it will fight you at every step.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The machine is not yours to administer.&lt;/strong&gt; A work laptop or a locked-down family
PC won&amp;rsquo;t let you install software or change power settings, and both are
required.&lt;/p&gt;
&lt;p&gt;One thing to check now rather than at step forty. Docker&amp;rsquo;s
&lt;a href="https://docs.docker.com/desktop/setup/install/windows-install/" class="external-link" rel="noopener"&gt;Windows install requirements&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;
list the Pro, Enterprise and Education editions. Press the Windows key, type
&lt;code&gt;winver&lt;/code&gt;, press Enter, and the window names your edition. If it says Home, treat
Part 2 as your test and attempt it before spending any money, which is why it is
the first install rather than the last.&lt;/p&gt;
&lt;p&gt;Table: What the build asks of you before you start. Money figures are from the sources cited in the sections below; the disk figure is this page&amp;rsquo;s working minimum rather than a vendor-published requirement.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;What it takes&lt;/th&gt;
 &lt;th&gt;The number&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;Your time&lt;/td&gt;
 &lt;td&gt;About three hours over two sittings. Two restarts during the install, one reboot test at the end.&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;One-time money&lt;/td&gt;
 &lt;td&gt;$5 of model credit. The card fee is 5.5% with a $0.80 minimum, so the charge lands around $5.80. No subscriptions.&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Monthly money&lt;/td&gt;
 &lt;td&gt;Electricity, roughly £0.60 to £1.30 at the current UK capped rate, plus model spend of roughly a third of a dollar.&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;The machine&lt;/td&gt;
 &lt;td&gt;A Windows laptop that stays home, plugged in, on your home Wi-Fi. Windows 10 (22H2) or Windows 11 (23H2 or later), 64-bit, 8 GB RAM, 20 GB free disk.&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Your phone&lt;/td&gt;
 &lt;td&gt;Android or iPhone, with Telegram and one more free app.&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Your router&lt;/td&gt;
 &lt;td&gt;Untouched. No port is forwarded from the internet, ever. One rule is added to the laptop&amp;rsquo;s own firewall in Part 6, scoped to private networks.&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;h2 id="the-money"&gt;The money&lt;/h2&gt;
&lt;h3 id="electricity-and-what-a-laptop-actually-draws"&gt;Electricity, and what a laptop actually draws&lt;/h3&gt;
&lt;p&gt;Two numbers turn a laptop into a monthly bill: what it draws, and what a
kilowatt-hour costs where you live. The second is published by regulators. The
first is where nearly every guide, including the source material this build came
from, quietly guesses, and the guess I inherited was 10 W to 30 W with the screen
off. It&amp;rsquo;s wrong, and it&amp;rsquo;s wrong in the expensive direction.&lt;/p&gt;
&lt;p&gt;ENERGY STAR publishes measured, at-the-wall figures for certified computers under
IEC 62301, and its &amp;ldquo;Long Idle&amp;rdquo; state is defined as awake with the display
backlight off, which is exactly the state this build sits in for almost the whole
day. Across the 1,174 notebook-class models in
&lt;a href="https://data.energystar.gov/d/rxdj-2c88" class="external-link" rel="noopener"&gt;that database&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; on 6 August 2026 the
median long-idle draw is 0.7 W, the 95th percentile is 2.1 W, and the highest
single entry is 7.1 W. Not one of the 1,174 reaches 10 W with the screen off. The
inherited range sits close to the ceiling
&lt;a href="https://eur-lex.europa.eu/LexUriServ/LexUriServ.do?uri=OJ:L:2013:175:0013:0033:EN:PDF" class="external-link" rel="noopener"&gt;the EU&amp;rsquo;s Ecodesign rules&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;
permit for a notebook, which is to say the figure everyone repeats is closer to
the worst computer the law allows than to a normal one.&lt;/p&gt;
&lt;p&gt;The complication has to travel with the number, because it moves it. Those
sub-1 W readings depend on the platform dropping into a low-power S0ix state when
the backlight goes off, and a resident Docker container plus a polling agent are
precisely what prevents that transition. Of the 1,148 rows typed plainly as
&amp;ldquo;Notebook&amp;rdquo;, 1,011 answer the sleep-mode field at all and 614 of those report
their long-idle state as a sleep state, so for a majority of certified machines
the published floor is a number this build will never see.&lt;/p&gt;
&lt;p&gt;That rules out the bottom of the range and leaves the top. The same database puts
short-idle, meaning awake with the screen on, at a median of 5.0 W, and the median
per-model gap between the two states is 4.1 W. That gap is not the display alone:
it is the display plus the deeper sleep the machine is no longer taking, and this
build gives up the second half of it. So the honest working envelope is the
certified screen-off band with its floor removed, roughly 3 W to 7 W, where 7 W is
within a rounding error of the highest figure any of the 1,174 models records.
Call it bounded inference from a measured ceiling rather than a measurement of
anything. One caveat travels with it in the other direction: the ENERGY STAR list
is certification data and is biased toward efficient hardware, so a gaming laptop
with discrete graphics is outside everything above.&lt;/p&gt;
&lt;p&gt;Two published rates turn watts into money. The UK figure is
&lt;a href="https://www.ofgem.gov.uk/information-consumers/energy-advice-households/energy-price-cap-unit-rates-and-standing-charges" class="external-link" rel="noopener"&gt;Ofgem&amp;rsquo;s capped average unit rate&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;
of 26.11p per kWh for 1 July to 30 September 2026. The EU figure is
&lt;a href="https://ec.europa.eu/eurostat/statistics-explained/index.php?title=Electricity_price_statistics" class="external-link" rel="noopener"&gt;Eurostat&amp;rsquo;s average household price&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;
of EUR 0.2896 per kWh for the second half of 2025. Both move, and both are linked
so you can substitute your own. The arithmetic is watts times 24 hours times 30
days, divided by 1,000.&lt;/p&gt;
&lt;p&gt;Table: Monthly electricity for a laptop left running continuously, across the idle-draw range the published measurements support. UK rate is Ofgem&amp;rsquo;s capped unit rate for 1 July to 30 September 2026; EU rate is Eurostat&amp;rsquo;s household average for the second half of 2025.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Idle draw&lt;/th&gt;
 &lt;th style="text-align: right"&gt;kWh per month&lt;/th&gt;
 &lt;th style="text-align: right"&gt;UK at 26.11p/kWh&lt;/th&gt;
 &lt;th style="text-align: right"&gt;EU at EUR 0.2896/kWh&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;3 W&lt;/td&gt;
 &lt;td style="text-align: right"&gt;2.16&lt;/td&gt;
 &lt;td style="text-align: right"&gt;£0.56&lt;/td&gt;
 &lt;td style="text-align: right"&gt;EUR 0.63&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;5 W&lt;/td&gt;
 &lt;td style="text-align: right"&gt;3.60&lt;/td&gt;
 &lt;td style="text-align: right"&gt;£0.94&lt;/td&gt;
 &lt;td style="text-align: right"&gt;EUR 1.04&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;7 W&lt;/td&gt;
 &lt;td style="text-align: right"&gt;5.04&lt;/td&gt;
 &lt;td style="text-align: right"&gt;£1.32&lt;/td&gt;
 &lt;td style="text-align: right"&gt;EUR 1.46&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;10 W&lt;/td&gt;
 &lt;td style="text-align: right"&gt;7.20&lt;/td&gt;
 &lt;td style="text-align: right"&gt;£1.88&lt;/td&gt;
 &lt;td style="text-align: right"&gt;EUR 2.09&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The last row is the inherited guess, kept for comparison. It is above the screen-off draw of every one of the 1,174 certified models, so treat it as the ceiling of the old assumption rather than a figure this build is likely to hit.&lt;/p&gt;
&lt;p&gt;Your standing charge is not in that table on purpose. You pay it whether this
laptop runs or not, so only the added kilowatt-hours belong to the machine.&lt;/p&gt;
&lt;p&gt;The envelope is roughly £0.56 to £1.32 a month in the UK, EUR 0.63 to EUR 1.46 on
the EU average. Seven to sixteen pounds a year.&lt;/p&gt;
&lt;p&gt;Which retires a claim I was going to make. At 10 W to 30 W electricity was the
largest recurring cost by a wide margin and the sentence wrote itself; at 3 W to
10 W it is merely the likely largest line at the usage level this page assumes,
the two costs are now the same order of magnitude, and which one wins depends on
how heavily you use the thing. The only real-world measurement either document
has points the other way: the author of the guide this build derives from
measured $2 to $6 a month of model spend on a heavier two-machine setup. One
person&amp;rsquo;s usage, not a forecast, and not your expected spend either. It is still
the only observed model-spend number in the file, and dropping it while asserting
the opposite would be the wrong way round.&lt;/p&gt;
&lt;p&gt;A smart plug with a power display settles the whole question in a day.&lt;/p&gt;
&lt;h3 id="what-one-scheduled-run-costs"&gt;What one scheduled run costs&lt;/h3&gt;
&lt;p&gt;The job I specified is a morning brief: at 07:30 the worker searches the web for
what moved in a chosen market overnight and sends a handful of bullets with
source links to Telegram.
&lt;a href="https://hermes-agent.nousresearch.com/docs/user-guide/features/cron" class="external-link" rel="noopener"&gt;Hermes handles the scheduling in chat&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;,
and jobs report back to wherever they were created.&lt;/p&gt;
&lt;p&gt;Start with the naive version, because it&amp;rsquo;s the one every guide publishes. Assume
one run sends about 6,000 tokens in and gets about 700 back. On
&lt;code&gt;deepseek/deepseek-v4-pro&lt;/code&gt;, which
&lt;a href="https://openrouter.ai/api/v1/models" class="external-link" rel="noopener"&gt;OpenRouter&amp;rsquo;s own model API&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; listed at
$0.435 in and $0.87 out per million tokens on 5 August 2026, that&amp;rsquo;s $0.0026 plus
$0.0006. A third of a cent.&lt;/p&gt;
&lt;p&gt;That number is wrong, and it&amp;rsquo;s wrong structurally rather than sloppily. A
search-and-summarise brief is not one request. It is a tool loop: the worker asks
the model what to do, the model calls &lt;code&gt;web_search&lt;/code&gt;, the results come back, the
model calls &lt;code&gt;web_extract&lt;/code&gt; on a page or two, the results come back again, and only
then does it write the brief. Each of those is a separate API call, and every
call re-sends the whole accumulated transcript, because that is how a stateless
chat completions API works. This page explains the same mechanism for long chats
three paragraphs from now. It applies to the scheduled job first.&lt;/p&gt;
&lt;p&gt;So take the shape rather than a false precision. Roughly N tool turns means
roughly N times the transcript re-sent, and the transcript grows as fetched text
lands in it. A three-turn brief finishing with a 6,000-token transcript sends
nearer 12,000 to 18,000 input tokens across the run.&lt;/p&gt;
&lt;p&gt;Table: One morning-brief run on deepseek/deepseek-v4-pro at OpenRouter&amp;rsquo;s listed prices on 5 August 2026, modelled as a three-turn tool loop. The token counts and the turn count are assumptions for a search-and-summarise job, not a measurement.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Direction&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Tokens&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Price per 1M tokens&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Cost&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;Input, as a single request&lt;/td&gt;
 &lt;td style="text-align: right"&gt;6,000&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$0.435&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$0.0026&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Input, three turns re-sending the transcript&lt;/td&gt;
 &lt;td style="text-align: right"&gt;about 18,000&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$0.435&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$0.0078&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Output, three turns&lt;/td&gt;
 &lt;td style="text-align: right"&gt;about 900&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$0.87&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$0.0008&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;One run&lt;/td&gt;
 &lt;td style="text-align: right"&gt;&lt;/td&gt;
 &lt;td style="text-align: right"&gt;&lt;/td&gt;
 &lt;td style="text-align: right"&gt;about $0.009&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Thirty runs&lt;/td&gt;
 &lt;td style="text-align: right"&gt;&lt;/td&gt;
 &lt;td style="text-align: right"&gt;&lt;/td&gt;
 &lt;td style="text-align: right"&gt;about $0.26&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Nearer a cent a morning than a third of one, and about a quarter a month for a
job that runs every day of it. The turn count is a guess as well, since a brief
that opens eight pages costs more than one that opens two, and I can&amp;rsquo;t tell you
which yours will be. The ground truth is
&lt;a href="https://openrouter.ai/activity" class="external-link" rel="noopener"&gt;OpenRouter&amp;rsquo;s activity page&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;, which logs every
request with its real cost, and after the first week of real runs that page
replaces this table entirely.&lt;/p&gt;
&lt;p&gt;Ad-hoc chat behaves the same way and has one trap in it. Long conversations
re-send the whole history with every message, so the twentieth turn in a thread
costs many times what the first one did. Forty ordinary standalone questions in a
month land around $0.05 to $0.15 by the same arithmetic. Forty turns of one
enormous thread do not. Add the two and the year comes to roughly $4 to $5, which
still fits inside the one-time $5 credit, but no longer with room to spare.&lt;/p&gt;
&lt;h3 id="three-models-and-the-one-that-can-make-this-expensive"&gt;Three models, and the one that can make this expensive&lt;/h3&gt;
&lt;p&gt;The chat interface is wired to exactly three model IDs rather than the hundreds
OpenRouter offers, which keeps the dropdown short and the choices deliberate.&lt;/p&gt;
&lt;p&gt;Table: The three models on the allowlist, with prices per million tokens from OpenRouter&amp;rsquo;s model API on 5 August 2026. All three carry a 1,048,576-token context window.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Model&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Input&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Output&lt;/th&gt;
 &lt;th&gt;What it is for&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;&lt;code&gt;deepseek/deepseek-v4-flash-0731&lt;/code&gt;&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$0.09&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$0.18&lt;/td&gt;
 &lt;td&gt;Bulk and mechanical work: classification, reformatting, first drafts, anything scheduled that runs often&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;&lt;code&gt;deepseek/deepseek-v4-pro&lt;/code&gt;&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$0.435&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$0.87&lt;/td&gt;
 &lt;td&gt;The default. Briefs, research, summaries, most real thinking&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;&lt;code&gt;moonshotai/kimi-k3&lt;/code&gt;&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$3.00&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$15.00&lt;/td&gt;
 &lt;td&gt;Escalation only, one task at a time&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Read the right-hand column as a ratio. Kimi K3&amp;rsquo;s output price is $15.00 against
V4 Pro&amp;rsquo;s $0.87, roughly seventeen times, on the side of the ledger you have least
control over, because you can budget your input and you cannot budget how much a
model decides to write.&lt;/p&gt;
&lt;p&gt;Which gives the escalation rule its shape. Start at V4 Pro. Drop to Flash when
the job is mechanical and frequent, because at $0.09 in you can afford to run it
hourly without thinking. Escalate to K3 only when V4 Pro has visibly failed at
the task in front of you, then go back down. A K3 habit is the only realistic
route by which this stack stops being a rounding error on your electricity bill.&lt;/p&gt;
&lt;p&gt;Context is the other half. All three models advertise a 1,048,576-token window,
and a large window is an invitation to paste enormous things into it. Feeding
500,000 tokens of context to K3 costs 500,000 divided by 1,000,000 times $3.00,
which is $1.50 before the model has written a word of reply. On V4 Pro the same
dump is about $0.22. On Flash it is $0.045. A window is what the model will
accept, not what you should send.&lt;/p&gt;
&lt;h3 id="pin-the-dated-snapshot-where-one-exists"&gt;Pin the dated snapshot where one exists&lt;/h3&gt;
&lt;p&gt;The flash entry above is written as &lt;code&gt;deepseek/deepseek-v4-flash-0731&lt;/code&gt; rather than
the unsuffixed slug, and that is a deliberate choice with money attached. Model
names come in two flavours on any hosted platform: a dated snapshot points at one
frozen release, while a bare or &lt;code&gt;-latest&lt;/code&gt; name points at whatever the provider
currently considers current, which means it can repoint under you with new
weights, different refusal behaviour, different output length and a different
price. In an interactive chat you&amp;rsquo;d notice inside a day. For a job that fires at
07:30 while you&amp;rsquo;re asleep you might not notice for a fortnight, and the first
signal would be either a strange brief or a bill that doesn&amp;rsquo;t match the
arithmetic above.&lt;/p&gt;
&lt;p&gt;So the rule is to pin the dated name in anything scheduled. It does not survive
contact with the model this build actually schedules.&lt;/p&gt;
&lt;p&gt;On 6 August 2026 I read
&lt;a href="https://openrouter.ai/api/v1/models" class="external-link" rel="noopener"&gt;OpenRouter&amp;rsquo;s model API&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; and pulled every
deepseek-v4 slug it publishes. There are four: &lt;code&gt;deepseek/deepseek-v4-flash&lt;/code&gt;,
&lt;code&gt;deepseek/deepseek-v4-flash-0731&lt;/code&gt;, &lt;code&gt;deepseek/deepseek-v4-pro&lt;/code&gt;, and
&lt;code&gt;~deepseek/deepseek-v4-flash-latest&lt;/code&gt;. Flash has a dated snapshot. Pro does not.
The default brain Part 7 sets is &lt;code&gt;deepseek/deepseek-v4-pro&lt;/code&gt;, so the rule as
stated is unfollowable for the one model that matters most here.&lt;/p&gt;
&lt;p&gt;Two ways out, and pick one on purpose. Keep Pro and accept that no pin is
available, in which case the monthly check stops being hygiene and becomes the
only control you have: read
&lt;a href="https://openrouter.ai/models" class="external-link" rel="noopener"&gt;OpenRouter&amp;rsquo;s model list&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;, confirm the slug still
exists and the price still matches the table above, and read one brief properly
instead of skimming it, because a repointed model shows up in the writing before
it shows up in the bill. Or move the schedule to
&lt;code&gt;deepseek/deepseek-v4-flash-0731&lt;/code&gt;, which is pinned, costs roughly a fifth as
much, and is what this page&amp;rsquo;s own escalation rule recommends for a mechanical job
that runs daily anyway, at the price of quality on the one output you read every
morning.&lt;/p&gt;
&lt;p&gt;Part 7 keeps Pro and pairs it with the calendar reminder. A pinned name that
starts returning &amp;ldquo;not found&amp;rdquo; is the system telling you to go look. An unpinned
one never tells you anything.&lt;/p&gt;
&lt;h2 id="the-caps-that-actually-exist"&gt;The caps that actually exist&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;OpenRouter standard accounts have no account-level spending cap.&lt;/strong&gt; That is the
finding I did not expect, and it is the sharpest thing on this page.
&lt;a href="https://openrouter.ai/docs/api-reference/limits" class="external-link" rel="noopener"&gt;Their limits documentation&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; is
explicit that what constrains you is the prepaid balance and per-key limits, and
there is no monthly-ceiling setting to switch on. Plenty of guides describe one.
It isn&amp;rsquo;t there. If you were relying on a dashboard toggle to save you from a
loop, you were relying on something that does not exist.&lt;/p&gt;
&lt;p&gt;What does exist is three layers, and they do different jobs.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The balance is the wall.&lt;/strong&gt; OpenRouter is prepaid. The minimum purchase is $5
per transaction under &lt;a href="https://openrouter.ai/terms" class="external-link" rel="noopener"&gt;section 4.1 of the terms&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;,
and the card fee is 5.5% with a $0.80 minimum per
&lt;a href="https://openrouter.ai/docs/faq" class="external-link" rel="noopener"&gt;their FAQ&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;, so the first charge lands around
$5.80. Leave auto top-up switched off and the worst possible month is the credit
you loaded. The same terms allow unused credits to expire after 365 days, so
stockpiling a large balance buys you nothing except a larger blast radius.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The per-key limit is the fuse, and a fuse only does work if it is smaller than
the wall.&lt;/strong&gt; Create one named key for this machine and set a credit limit on it in
the &lt;a href="https://openrouter.ai/docs/api-keys" class="external-link" rel="noopener"&gt;Keys interface&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;. When a key reaches its
limit, requests fail with a 402 and the jobs stop, which is the behaviour you
want: a loud, specific error on the activity page rather than a quiet drain.&lt;/p&gt;
&lt;p&gt;Part 4 departs from the guide this build came from, deliberately. That guide
loads a $5 balance and sets the key limit to $5, and because a fuse rated at
exactly the wall&amp;rsquo;s value can never trip before the wall does, what reads on the
page as two independent safety layers turns out in practice to be one layer
written down twice, which fails you precisely in the case both were there to
cover: the unattended loop at 04:00 that nobody is watching. Load $5, set the key
to $2. A loop then stops at $2 with a visible error while $3 of balance sits
untouched behind it, and you decide whether to raise the limit instead of
learning the answer from an empty account. Two keys with separate limits give you
two fuses, which
&lt;a href="https://openrouter.ai/docs/features/provisioning-api-keys" class="external-link" rel="noopener"&gt;OpenRouter&amp;rsquo;s provisioning documentation&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;
covers if you want per-job accounting.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The job list is the patrol.&lt;/strong&gt; &lt;code&gt;/cron list&lt;/code&gt; in the bot chat shows everything
scheduled. Read it weekly. Anything you no longer recognise or no longer open
gets paused or removed. One structural comfort:
&lt;a href="https://hermes-agent.nousresearch.com/docs/user-guide/features/cron" class="external-link" rel="noopener"&gt;Hermes documents that a scheduled job cannot create scheduled jobs&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;,
so a runaway job can repeat itself but cannot multiply into a family of jobs
while you sleep.&lt;/p&gt;
&lt;p&gt;A normal month on the activity page is one small entry around 07:30 each morning,
a scatter of chat entries during the day, everything in fractions of a cent.
Runaway has three shapes: the same job appearing many times an hour, which means
a schedule set wrong; entries on &lt;code&gt;moonshotai/kimi-k3&lt;/code&gt; you don&amp;rsquo;t remember asking
for, which means something defaulted to the expensive model; or a single day
costing more than a normal month, which usually means one job is being fed
something enormous every run.&lt;/p&gt;
&lt;p&gt;The response, in order, and the order is the point. Pause the job first, in plain
language or with &lt;code&gt;/cron pause&lt;/code&gt;, because stopping the spend beats diagnosing it.
Then &lt;code&gt;/cron list&lt;/code&gt; for anything else you didn&amp;rsquo;t expect. Then the Keys page to
confirm the fuse held. Diagnose last. Your damage ceiling was set when you chose
the balance and the key limit, and if those two are right the worst outcome is an
annoying afternoon rather than a bill.&lt;/p&gt;
&lt;p&gt;The absence of the morning message is also the monitoring system. If nothing
arrives at 08:00, something is wrong, and it costs nothing to notice.&lt;/p&gt;
&lt;h2 id="the-licence-lines-that-bite-a-business"&gt;The licence lines that bite a business&lt;/h2&gt;
&lt;p&gt;Two of the four installs are free for a person and not necessarily free for a
company.&lt;/p&gt;
&lt;p&gt;Docker Desktop is free for personal use and for businesses under 250 employees
and under $10 million in annual revenue.
&lt;a href="https://docs.docker.com/subscription/desktop-license/" class="external-link" rel="noopener"&gt;Docker&amp;rsquo;s licence page&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;
sets both lines, and crossing either one puts you on a paid subscription. The
cheapest paid tier on &lt;a href="https://www.docker.com/pricing/" class="external-link" rel="noopener"&gt;Docker&amp;rsquo;s pricing page&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; is
Docker Pro at $9 per user per month billed annually, but Pro is a single-user
plan, so a business licensing a team cannot buy it: that is Team at $15 or
Business at $24 per user per month. Two thresholds joined by &amp;ldquo;and&amp;rdquo; on the free
side means either one can move you, and the price you land on is not the headline
one.&lt;/p&gt;
&lt;p&gt;Tailscale reads softer, but their page resolves more of it than I first credited.
&lt;a href="https://tailscale.com/pricing" class="external-link" rel="noopener"&gt;The pricing page&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; describes the free Personal
plan as &amp;ldquo;only suitable for non-commercial use of Tailscale&amp;rdquo;, and a line earlier
as being &amp;ldquo;for individuals who want to use Tailscale at home&amp;rdquo;. It also states the
mechanical rule Tailscale applies at signup: an account created with a public
email domain lands on Personal, while a custom domain triggers a business trial.
That is Tailscale telling you how it reads your situation, and it&amp;rsquo;s why Part 6
says to sign in with a personal address. What it doesn&amp;rsquo;t resolve is the middle
case, one person at home on a personal address reaching their own laptop to run
their own business errands. Standard is listed at $8 per user per month. Read
their page and decide, or pay the $8 and stop thinking about it.&lt;/p&gt;
&lt;p&gt;Everything else is genuinely $0 a month at this scale: Open WebUI, Hermes Agent,
and the Telegram Bot Platform. The running total is a pound or two of electricity
plus pennies to a couple of dollars of model spend, on hardware you already own,
after the one-time top-up.&lt;/p&gt;
&lt;h2 id="privacy-without-hedging"&gt;Privacy, without hedging&lt;/h2&gt;
&lt;p&gt;Prompts route through OpenRouter to a model provider, and each provider handles
data under its own policy. OpenRouter
&lt;a href="https://openrouter.ai/docs/guides/privacy/data-collection" class="external-link" rel="noopener"&gt;states that it does not store prompt or completion content unless you opt in&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;,
but it always stores request metadata: token counts, timestamps, model names. Its
&lt;a href="https://openrouter.ai/docs/guides/privacy/provider-logging" class="external-link" rel="noopener"&gt;provider routing controls&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;
govern whether requests can go to providers that may train on the data, and the
&lt;a href="https://openrouter.ai/privacy" class="external-link" rel="noopener"&gt;privacy policy&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; is the document that binds.&lt;/p&gt;
&lt;p&gt;Treat the machine like any other cloud service, because that&amp;rsquo;s what it is. The
chat page runs on your laptop; the thinking does not. Market research and draft
copy are fine. Customer names, addresses, invoices, anything you wouldn&amp;rsquo;t paste
into a normal AI chat window: no. Any guide telling you a self-hosted chat page
makes your prompts private has confused the interface with the model.&lt;/p&gt;
&lt;p&gt;The chat page itself is protected by whatever password you set on its first
account, and
&lt;a href="https://docs.openwebui.com/getting-started/advanced-topics/hardening/" class="external-link" rel="noopener"&gt;Open WebUI&amp;rsquo;s hardening guide&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;
covers locking it down. That password opens a door to a machine on your home
network, so it should not be one you use anywhere else.&lt;/p&gt;
&lt;h2 id="renting-a-server-instead"&gt;Renting a server instead&lt;/h2&gt;
&lt;p&gt;The laptop is one answer to &amp;ldquo;where does this run&amp;rdquo;. The other is a small virtual
server, and it deserves a real comparison rather than a footnote.&lt;/p&gt;
&lt;p&gt;State the requirement before the prices and you stop buying compute you&amp;rsquo;ll never
use. Open WebUI publishes no minimum memory figure: I read
&lt;a href="https://docs.openwebui.com/getting-started/quick-start/" class="external-link" rel="noopener"&gt;its quick start&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; and
&lt;a href="https://github.com/open-webui/open-webui" class="external-link" rel="noopener"&gt;its README&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; on 6 August 2026 and
neither states one. So the 1 to 2 GB I work from is an estimate, and this is what
it rests on. In
&lt;a href="https://github.com/open-webui/open-webui/discussions/2583" class="external-link" rel="noopener"&gt;the project&amp;rsquo;s own discussion of memory use&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;
a maintainer reports the container idling at 1.4 GB, the person who opened the
thread measures 500 MB from a cold start and over 1 GB in use, and a contributor
gets it to roughly 200 MB by stripping optional services out; a separate thread
on
&lt;a href="https://github.com/open-webui/open-webui/discussions/736" class="external-link" rel="noopener"&gt;minimum system requirements&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;
puts an API-only install at 1 GB of RAM, 1 CPU core and 10 GB of disk. User and
maintainer reports, not a vendor requirement, and they select every row below.&lt;/p&gt;
&lt;p&gt;The agent is small, the operating system takes its share, and the model runs
somewhere else entirely, so the CPU is close to irrelevant: this thing is idle
more than 99% of the time and then spends thirty seconds waiting on someone
else&amp;rsquo;s GPU. A 2 GB instance fits. A 4 GB instance is comfortable.&lt;/p&gt;
&lt;p&gt;Table: Cheapest plan at each provider that genuinely fits Open WebUI plus the agent plus the operating system, read on 6 August 2026. Prices exclude VAT and are the on-demand monthly rate for the provider&amp;rsquo;s cheapest region. The Vultr and Linode/Akamai pages answer automated requests with HTTP 403, so those two rows were read in a browser and are cited inline below rather than in this page&amp;rsquo;s machine-checked source list.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Provider&lt;/th&gt;
 &lt;th&gt;Plan&lt;/th&gt;
 &lt;th style="text-align: right"&gt;RAM&lt;/th&gt;
 &lt;th style="text-align: right"&gt;vCPU&lt;/th&gt;
 &lt;th&gt;Disk&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Monthly&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;Vultr&lt;/td&gt;
 &lt;td&gt;Cloud Compute, Regular Performance&lt;/td&gt;
 &lt;td style="text-align: right"&gt;2 GB&lt;/td&gt;
 &lt;td style="text-align: right"&gt;1&lt;/td&gt;
 &lt;td&gt;55 GB SSD, 2 TB transfer&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$10.00&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;DigitalOcean&lt;/td&gt;
 &lt;td&gt;Basic Droplet, Regular&lt;/td&gt;
 &lt;td style="text-align: right"&gt;2 GB&lt;/td&gt;
 &lt;td style="text-align: right"&gt;1&lt;/td&gt;
 &lt;td&gt;50 GB SSD, 2,000 GiB transfer&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$12.00&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Linode / Akamai&lt;/td&gt;
 &lt;td&gt;Shared CPU, Linode 2 GB&lt;/td&gt;
 &lt;td style="text-align: right"&gt;2 GB&lt;/td&gt;
 &lt;td style="text-align: right"&gt;1&lt;/td&gt;
 &lt;td&gt;50 GB, 2 TB transfer&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$12.00&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Hetzner Cloud&lt;/td&gt;
 &lt;td&gt;CPX12, Regular Performance, AMD&lt;/td&gt;
 &lt;td style="text-align: right"&gt;2 GB&lt;/td&gt;
 &lt;td style="text-align: right"&gt;1&lt;/td&gt;
 &lt;td&gt;40 GB&lt;/td&gt;
 &lt;td style="text-align: right"&gt;EUR 11.49 plus EUR 0.50 IPv4&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Oracle Cloud&lt;/td&gt;
 &lt;td&gt;Always Free, Ampere A1&lt;/td&gt;
 &lt;td style="text-align: right"&gt;up to 12 GB&lt;/td&gt;
 &lt;td style="text-align: right"&gt;up to 2 OCPU&lt;/td&gt;
 &lt;td&gt;200 GB block storage&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$0.00&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The 4 GB step-up, same sources and same date, runs $20 at Vultr, $24 at
DigitalOcean, $24 at Linode, and EUR 19.49 for Hetzner&amp;rsquo;s CPX22.&lt;/p&gt;
&lt;p&gt;Three rows came out of awkward pages, and saying so is more useful than
pretending otherwise. &lt;a href="https://www.vultr.com/pricing/" class="external-link" rel="noopener"&gt;Vultr&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; and
&lt;a href="https://www.akamai.com/cloud/pricing/north-america" class="external-link" rel="noopener"&gt;Akamai&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;, which is where
Linode&amp;rsquo;s plan table lives now that &lt;code&gt;linode.com/pricing&lt;/code&gt; redirects, both answer
non-browser requests with HTTP 403, so their numbers were read in a browser on
6 August 2026. Hetzner&amp;rsquo;s
&lt;a href="https://www.hetzner.com/cloud/regular-performance/" class="external-link" rel="noopener"&gt;regular performance page&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;
renders prices client-side and every cell came out blank in my session, so rather
than fill the gap from memory I read each row&amp;rsquo;s product key out of the markup and
resolved it against
&lt;a href="https://www.hetzner.com/_resources/app/data/app/live_data_prices.json" class="external-link" rel="noopener"&gt;the live price feed the page itself loads&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;.
And before planning around Hetzner at all,
&lt;a href="https://www.hetzner.com/cloud/cost-optimized/" class="external-link" rel="noopener"&gt;every row on its cost-optimized line&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;
carried a &amp;ldquo;not available&amp;rdquo; label on the day I looked, including the cheap Arm plan
that usually makes Hetzner the obvious answer. Stock changes. Check it yourself.&lt;/p&gt;
&lt;p&gt;Oracle&amp;rsquo;s free tier is the genuinely interesting row.
&lt;a href="https://docs.oracle.com/en-us/iaas/Content/FreeTier/freetier_topic-Always_Free_Resources.htm" class="external-link" rel="noopener"&gt;Oracle&amp;rsquo;s own Always Free documentation&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;
states an Ampere A1 allocation of &amp;ldquo;1,500 OCPU hours per month&amp;rdquo; and &amp;ldquo;9,000 GB
hours per month&amp;rdquo;, which is 2 OCPUs and 12 GB of memory running continuously, plus
200 GB of storage and 10 TB of outbound transfer. Several times what this build
needs, at nothing a month.&lt;/p&gt;
&lt;h3 id="two-points-for-the-server-two-for-the-laptop"&gt;Two points for the server, two for the laptop&lt;/h3&gt;
&lt;p&gt;I am not stacking this. The server genuinely wins twice. It simplifies the
network, because Telegram&amp;rsquo;s API and the model API are both outbound HTTPS from
any host and a server has a public IP, so the private-network step becomes
optional rather than load-bearing.&lt;/p&gt;
&lt;p&gt;And it dissolves the Docker licensing question. A server runs Linux, so you run
Docker Engine rather than Docker Desktop.
&lt;a href="https://docs.docker.com/engine/" class="external-link" rel="noopener"&gt;Docker&amp;rsquo;s engine documentation&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; states the
Apache License, Version 2.0, and
&lt;a href="https://raw.githubusercontent.com/moby/moby/master/LICENSE" class="external-link" rel="noopener"&gt;the LICENSE file in the moby repository&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;
is Apache License, Version 2.0, January 2004. The thresholds on
&lt;a href="https://docs.docker.com/subscription/desktop-license/" class="external-link" rel="noopener"&gt;Docker Desktop&amp;rsquo;s licence page&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;
do not reach Docker Engine on Linux, so if your company is over those lines that
is $9 to $24 per user per month the server side wins outright.&lt;/p&gt;
&lt;p&gt;The laptop wins twice as well. The first is arithmetic: it&amp;rsquo;s a sunk cost, and its
marginal cost is electricity, roughly £7 to £16 a year at the draws in the table
above, against $120 to $144 a year for the cheapest fitting server, recurring
forever, in a currency that may not be yours.&lt;/p&gt;
&lt;p&gt;The second is where the disk lives. On a rented server the provider is a party to
your data at rest: the hypervisor host can read the guest&amp;rsquo;s memory and disk,
snapshots sit on their storage, and the volume is reachable by their staff, their
subpoena process and their jurisdiction. Hetzner is Germany and Finland and is
framed around GDPR; DigitalOcean, Linode/Akamai, Vultr and Oracle are
US-headquartered whichever region you pick. On a laptop in your own home, the
chat history, the Telegram credentials and the API key sit on a disk no provider
can touch. For an assistant that reads your messages, that difference is not
small.&lt;/p&gt;
&lt;h3 id="twelve-months-and-which-one-wins"&gt;Twelve months, and which one wins&lt;/h3&gt;
&lt;p&gt;Table: Twelve-month total for the same build in three places. Assumptions: the laptop is already owned and its purchase price is excluded; idle draw is taken at 7 W, the upper end of the bounded-inference range above and not a measurement; the UK rate of 26.11p/kWh (Ofgem, 1 July to 30 September 2026) and the EU average of EUR 0.2896/kWh (Eurostat, second half of 2025) are held flat for twelve months, which they will not be; model spend is identical in all three rows because the model runs remotely in all three, and at roughly $4 to $5 a year it still fits inside the one-time credit, though only just; server prices are the on-demand rates read on 6 August 2026, exclude VAT and any overage, and assume no price change; the Oracle row assumes the Always Free tier remains available and survives the capacity limits Oracle applies to it; no currency conversion is performed anywhere in this table.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Where it runs&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Recurring per month&lt;/th&gt;
 &lt;th&gt;Twelve-month total, including the one-time $5.80 of model credit&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;Laptop you already own&lt;/td&gt;
 &lt;td style="text-align: right"&gt;£1.32 (EUR 1.46) electricity&lt;/td&gt;
 &lt;td&gt;$5.80 plus £15.84 (EUR 17.52)&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Cheapest fitting VPS (Vultr 2 GB)&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$10.00 hosting&lt;/td&gt;
 &lt;td&gt;$125.80&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Free-tier VPS (Oracle Always Free A1)&lt;/td&gt;
 &lt;td style="text-align: right"&gt;$0.00&lt;/td&gt;
 &lt;td&gt;$5.80&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Two inputs moved while this page was being checked. The electricity figure came
down when the inherited idle-draw assumption was replaced with published
measurements, so the laptop row is roughly half what an earlier draft carried,
and the model-spend estimate went up, because re-costing the brief as a tool loop
rather than a single request roughly triples it. Neither change reorders the
table: the laptop still wins on money by roughly a factor of five at any
plausible exchange rate, which is why no conversion is performed here.&lt;/p&gt;
&lt;p&gt;Rent the server when uptime matters more than ownership. When the laptop would be
closed, carried, sleeping, or sitting on a residential connection that drops.
When somebody other than you depends on the thing answering. At $10 to $12 a
month you are buying availability and simplicity, not compute. Oracle&amp;rsquo;s free tier
is the middle path worth taking seriously, with three caveats: it is capacity
limited and can refuse to launch, it wants a card at signup, and it puts your
chat history on a US hyperscaler&amp;rsquo;s disk.&lt;/p&gt;
&lt;p&gt;Keep the laptop when the hardware is already paid for and already running, when
the data is personal enough that provider access is a real objection, and when a
few hours of downtime a year cost you nothing. For a single-user assistant in
your own house, that last one is usually the honest answer.&lt;/p&gt;
&lt;h2 id="four-things-i-left-out-and-why"&gt;Four things I left out, and why&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;An Apple Silicon Mac instead of a Windows laptop.&lt;/strong&gt; Docker Desktop and Open
WebUI behave the same way on macOS, and
&lt;a href="https://hermes-agent.nousresearch.com/docs/getting-started/platform-support" class="external-link" rel="noopener"&gt;Hermes supports Apple Silicon while explicitly not supporting Intel Macs&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;.
The catch is physical: a closed MacBook sleeps, so it has to live lid-open with
sleep disabled. Workable. The Windows laptop simply has fewer ways to go wrong.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;A free local model on your own GPU.&lt;/strong&gt; Tools like Ollama run open weights on
your own hardware, which is free per run, private, and works offline. On an
ordinary laptop GPU it&amp;rsquo;s also slow, slow enough that a job set for 07:30 is not
reliably finished by 08:00. And the whole nightly workload costs a few tens of
cents a month on hosted models, so local inference is a hobby lane rather than
week-one material.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;A browser-based coding environment.&lt;/strong&gt; Open WebUI&amp;rsquo;s team ships
&lt;a href="https://github.com/open-webui/computer" class="external-link" rel="noopener"&gt;a tool called Open WebUI Computer, installed as the &lt;code&gt;cptr&lt;/code&gt; package&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;,
which serves your machine&amp;rsquo;s real files, terminal, editor and git to a browser
tab. It works. It also hands full access to your machine to whoever holds the
login, and nothing in the first month of this build needs it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Running a Claude subscription through third-party tools.&lt;/strong&gt; The setup this build
derives from plugged a subscription-authenticated CLI into that environment to
reuse a plan you already pay for. I had assumed Anthropic&amp;rsquo;s documentation ruled
that out, and on re-reading, the live pages say something more interesting.
&lt;a href="https://code.claude.com/docs/en/agent-sdk/overview" class="external-link" rel="noopener"&gt;The Agent SDK overview&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;
points third-party tools at API-key authentication, but
&lt;a href="https://support.claude.com/en/articles/15036540-use-the-claude-agent-sdk-with-your-claude-plan" class="external-link" rel="noopener"&gt;the support article on using the SDK with a Claude plan&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;
opened on 6 August 2026 with a banner reading &amp;ldquo;Update June 15: We&amp;rsquo;re pausing the
changes to Claude Agent SDK usage described below. For now, nothing has changed:
Claude Agent SDK, &lt;code&gt;claude -p&lt;/code&gt;, and third-party app usage still draw from your
subscription&amp;rsquo;s usage limits.&amp;rdquo; So it works today, by the vendor&amp;rsquo;s own words, and
in the same sentence the vendor tells you it was about to stop working and may
yet. A policy explicitly described as paused is a poor foundation for a process
you leave running unattended for a year. If you pay for Claude, use Claude&amp;rsquo;s own
apps alongside this build, and watch that banner rather than this page.&lt;/p&gt;
&lt;h2 id="the-build-eight-parts"&gt;The build, eight parts&lt;/h2&gt;
&lt;p&gt;One path, in order, with no choices to make along the way. Do all eight parts
sitting at the machine that will stay home, and read this page in a browser on
that machine so you can copy commands rather than retype them.&lt;/p&gt;
&lt;p&gt;Each command block has a line above saying what it does and a line beneath saying
what a working result looks like. That second line is the point of the format:
most walkthroughs tell you what to type and leave you to guess whether it worked,
so the first silent failure takes the next hour with it. If your screen doesn&amp;rsquo;t
match the &amp;ldquo;what you should see&amp;rdquo; line, stop and read that part&amp;rsquo;s failure table.
Anywhere you substitute your own value, the text is an OBVIOUS-PLACEHOLDER in
capitals with a filled example underneath. Parts 1 to 5 make a sensible first
sitting, Parts 6 to 8 a second.&lt;/p&gt;
&lt;h3 id="part-1-make-the-machine-stay-awake-15-min"&gt;Part 1, make the machine stay awake (15 min)&lt;/h3&gt;
&lt;p&gt;Factory settings put a laptop to sleep the moment you stop touching it.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Plug it in. It stays plugged in from now on.&lt;/li&gt;
&lt;li&gt;Open Settings. Windows 11: System, Power and battery, Screen and sleep.
Windows 10: System, Power and sleep. Different names, same settings.&lt;/li&gt;
&lt;li&gt;Set &amp;ldquo;when plugged in, put my device to sleep after&amp;rdquo; to &lt;strong&gt;Never&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Set &amp;ldquo;when plugged in, turn off my screen after&amp;rdquo; to &lt;strong&gt;5 minutes&lt;/strong&gt;. A dark
screen saves power while the machine underneath keeps working.&lt;/li&gt;
&lt;li&gt;Now the lid, which is the step people skip. Type &amp;ldquo;lid&amp;rdquo; into the Settings
search box, open &lt;strong&gt;&amp;ldquo;Change what closing the lid does&amp;rdquo;&lt;/strong&gt;, set &lt;strong&gt;&amp;ldquo;When I close
the lid&amp;rdquo;&lt;/strong&gt; in the &lt;strong&gt;Plugged in&lt;/strong&gt; column to &lt;strong&gt;Do nothing&lt;/strong&gt;, and save. Without
this, shutting the lid puts the whole build to sleep.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Table: Part 1 failures and fixes.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Symptom&lt;/th&gt;
 &lt;th&gt;Fix&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;No &amp;ldquo;plugged in&amp;rdquo; column on the lid screen&lt;/td&gt;
 &lt;td&gt;The laptop is running on battery. Plug it in and reopen the screen.&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;&amp;ldquo;Change what closing the lid does&amp;rdquo; not found in search&lt;/td&gt;
 &lt;td&gt;Windows key, type &amp;ldquo;control panel&amp;rdquo;, open it, then Hardware and Sound, Power Options, &amp;ldquo;Choose what closing the lid does&amp;rdquo; in the left sidebar.&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;&lt;strong&gt;Checkpoint.&lt;/strong&gt; Stop here and you still have a laptop that stays awake with the
lid shut, which is the physical foundation of everything else. Test it: play a
long video with sound, close the lid, listen. Audio still playing means the
machine is awake.&lt;/p&gt;
&lt;h3 id="part-2-install-docker-desktop-40-min-two-restarts"&gt;Part 2, install Docker Desktop (40 min, two restarts)&lt;/h3&gt;
&lt;p&gt;Docker runs the chat interface in the background. Install it once, set it to
start with Windows, never think about it again.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Check the machine can run it. Task Manager (Ctrl+Shift+Esc), Performance tab,
CPU, bottom right: &lt;strong&gt;&amp;ldquo;Virtualization: Enabled&amp;rdquo;&lt;/strong&gt;. If Task Manager shows only a
plain list of apps, click &lt;strong&gt;More details&lt;/strong&gt;; on Windows 11, Performance is the
graph icon in the left sidebar. Disabled means read the failure table before
going any further.&lt;/li&gt;
&lt;li&gt;Open PowerShell as administrator: Windows key, type &amp;ldquo;powershell&amp;rdquo;, right-click
&lt;strong&gt;Windows PowerShell&lt;/strong&gt;, &lt;strong&gt;Run as administrator&lt;/strong&gt;, Yes.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;This installs the Windows subsystem Docker runs on:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-powershell" data-lang="powershell"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;wsl&lt;/span&gt; &lt;span class="p"&gt;-&lt;/span&gt;&lt;span class="n"&gt;-install&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;What you should see: several components download and install, then a message
asking you to restart the computer. Restart it.&lt;/p&gt;
&lt;ol start="3"&gt;
&lt;li&gt;After the restart, a black Ubuntu window may ask for a username. Docker does
not use it; close it.&lt;/li&gt;
&lt;li&gt;From &lt;a href="https://www.docker.com/products/docker-desktop/" class="external-link" rel="noopener"&gt;Docker&amp;rsquo;s download page&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;,
take the &lt;strong&gt;AMD64&lt;/strong&gt; build of Docker Desktop for Windows (correct for both Intel
and AMD laptops; ARM64 is for rare ARM machines) and run the installer. Keep
&lt;strong&gt;&amp;ldquo;Use WSL 2 instead of Hyper-V&amp;rdquo;&lt;/strong&gt; ticked if asked. Restart again if asked.&lt;/li&gt;
&lt;li&gt;Open Docker Desktop, accept the service agreement, and choose &lt;strong&gt;Continue
without signing in&lt;/strong&gt;. Wait until the whale icon in the tray stops animating
and the window says &lt;strong&gt;Engine running&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Gear icon, General, tick &lt;strong&gt;&amp;ldquo;Start Docker Desktop when you sign in to your
computer&amp;rdquo;&lt;/strong&gt;. That&amp;rsquo;s what brings the chat interface back after a restart. Do
not skip it.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Table: Part 2 failures and fixes.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Symptom&lt;/th&gt;
 &lt;th&gt;Fix&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;Virtualization: Disabled in Task Manager&lt;/td&gt;
 &lt;td&gt;It has to be switched on in the laptop&amp;rsquo;s BIOS. The no-key-tapping route: Settings, System, Recovery, Advanced startup, Restart now, then Troubleshoot, Advanced options, UEFI Firmware Settings, Restart. Fallback is tapping F2 or Delete during startup. Enable the setting named SVM Mode (AMD) or Intel VT-x / Virtualization Technology (Intel), save and exit.&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;&lt;code&gt;wsl --install&lt;/code&gt; errors about your Windows version&lt;/td&gt;
 &lt;td&gt;Windows is too old. &lt;code&gt;winver&lt;/code&gt; shows your version; you need Windows 10 22H2 or Windows 11 23H2 or later. Run Windows Update, then retry.&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Docker Desktop says WSL needs updating&lt;/td&gt;
 &lt;td&gt;In the same admin PowerShell run &lt;code&gt;wsl --update&lt;/code&gt;, then reopen Docker Desktop.&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;&lt;strong&gt;Checkpoint.&lt;/strong&gt; Stop here and you still have Docker installed, running, and set
to start with Windows.&lt;/p&gt;
&lt;h3 id="part-3-start-the-chat-interface-20-min"&gt;Part 3, start the chat interface (20 min)&lt;/h3&gt;
&lt;p&gt;Open WebUI is a chat page served from your machine: the familiar layout, on your
laptop, with an account that is yours alone. Open PowerShell, normal this time,
no administrator needed.&lt;/p&gt;
&lt;p&gt;This downloads the chat interface and starts it in the background, set to restart
itself whenever the machine reboots:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-powershell" data-lang="powershell"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;docker&lt;/span&gt; &lt;span class="n"&gt;run&lt;/span&gt; &lt;span class="n"&gt;-d&lt;/span&gt; &lt;span class="n"&gt;-p&lt;/span&gt; &lt;span class="mf"&gt;3000&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;&lt;span class="mf"&gt;8080&lt;/span&gt; &lt;span class="n"&gt;-v&lt;/span&gt; &lt;span class="nb"&gt;open-webui&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;&lt;span class="p"&gt;/&lt;/span&gt;&lt;span class="n"&gt;app&lt;/span&gt;&lt;span class="p"&gt;/&lt;/span&gt;&lt;span class="n"&gt;backend&lt;/span&gt;&lt;span class="p"&gt;/&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt; &lt;span class="p"&gt;-&lt;/span&gt;&lt;span class="n"&gt;-name&lt;/span&gt; &lt;span class="nb"&gt;open-webui&lt;/span&gt; &lt;span class="p"&gt;-&lt;/span&gt;&lt;span class="n"&gt;-restart&lt;/span&gt; &lt;span class="n"&gt;always&lt;/span&gt; &lt;span class="n"&gt;ghcr&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;io&lt;/span&gt;&lt;span class="p"&gt;/&lt;/span&gt;&lt;span class="nb"&gt;open-webui&lt;/span&gt;&lt;span class="p"&gt;/&lt;/span&gt;&lt;span class="nb"&gt;open-webui&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;&lt;span class="n"&gt;main&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;What you should see: several download progress bars, then one long string of
random letters and numbers on its own line. That string means it started. The
first download is a few gigabytes, so let it run. If the screen goes dark
mid-download, that&amp;rsquo;s Part 1&amp;rsquo;s screen-off setting doing its job. Move the mouse;
the download never stopped.&lt;/p&gt;
&lt;p&gt;This confirms it is running:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-powershell" data-lang="powershell"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;docker&lt;/span&gt; &lt;span class="nb"&gt;ps
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;What you should see: a table with one row, &lt;code&gt;open-webui&lt;/code&gt;, whose STATUS column
starts with &lt;code&gt;Up&lt;/code&gt;.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Open a browser on the machine and go to &lt;code&gt;http://localhost:3000&lt;/code&gt;, allowing a
minute on first load. Click &lt;strong&gt;Sign up&lt;/strong&gt;. Use a real email format, since
nothing is ever sent and the address is only your username, and a strong
password you use nowhere else, because this page is a door to your machine.&lt;/li&gt;
&lt;li&gt;&lt;a href="https://docs.openwebui.com/features/authentication-access/rbac/roles/" class="external-link" rel="noopener"&gt;The first account created becomes the administrator&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;,
and sign-ups switch themselves off once it exists. Confirm that lock: your
name at the bottom left, Admin Panel, Settings, General, sign-up toggle off.
The wording shifts between versions. Your own account is unaffected.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Table: Part 3 failures and fixes.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Symptom&lt;/th&gt;
 &lt;th&gt;Fix&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;&lt;code&gt;docker: command not found&lt;/code&gt;, or a pipe error&lt;/td&gt;
 &lt;td&gt;Docker Desktop is not running yet. Open it, wait for Engine running, try again.&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Error says port 3000 is already in use&lt;/td&gt;
 &lt;td&gt;Re-run the long command with &lt;code&gt;-p 3001:8080&lt;/code&gt; instead, and use 3001 everywhere this page says 3000.&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Page never loads&lt;/td&gt;
 &lt;td&gt;Wait 60 seconds after &lt;code&gt;docker ps&lt;/code&gt; shows Up. The first start is slow.&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;&lt;strong&gt;Checkpoint.&lt;/strong&gt; Stop here and you still have a private chat page running on your
machine. It has no brain yet.&lt;/p&gt;
&lt;h3 id="part-4-open-the-brain-account-15-min"&gt;Part 4, open the brain account (15 min)&lt;/h3&gt;
&lt;p&gt;OpenRouter is one account and one balance across many paid models. Load $5, set a
smaller fuse inside it, and that balance is the hard ceiling on what this machine
can ever spend.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Sign up at &lt;a href="https://openrouter.ai/terms" class="external-link" rel="noopener"&gt;openrouter.ai&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; with a login you
control long-term. Credits and Keys live under the profile icon, top right.&lt;/li&gt;
&lt;li&gt;Credits, add &lt;strong&gt;$5&lt;/strong&gt;, the minimum per their terms. The 5.5% card fee with its
$0.80 minimum makes the total about $5.80, itemised before you pay.&lt;/li&gt;
&lt;li&gt;Leave auto top-up &lt;strong&gt;off&lt;/strong&gt;. Standard accounts have no account-wide cap, so the
balance is the cap, and with auto top-up off the worst month costs $5.&lt;/li&gt;
&lt;li&gt;Keys, create a key named &lt;code&gt;night-shift&lt;/code&gt;, and put &lt;code&gt;2&lt;/code&gt; in the credit limit field.
If anything ever loops, the key shuts off at $2 with an error while $3 of your
balance is still there. This departs deliberately from the guide this build
follows, which sets the limit to the full $5: a fuse rated at the wall&amp;rsquo;s own
value can never trip first, and two layers that fire at the same number are
one layer.&lt;/li&gt;
&lt;li&gt;The key is &lt;code&gt;sk-or-v1-&lt;/code&gt; plus a long string and it is shown exactly once. Copy
it into a password manager before you close the page, because you paste it
twice, in Part 5 and Part 7. It is the wallet: never in a chat message, a
screenshot, or a shared workflow.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Table: Part 4 failures and fixes.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Symptom&lt;/th&gt;
 &lt;th&gt;Fix&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;You closed the key window without copying it&lt;/td&gt;
 &lt;td&gt;It cannot be shown again. Delete the key on the Keys page, create a new one, copy it this time.&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Checkout total looks higher than $5&lt;/td&gt;
 &lt;td&gt;That is the purchase fee, itemised at checkout. A total near six dollars is correct.&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Later, requests start failing with a 402&lt;/td&gt;
 &lt;td&gt;The &lt;code&gt;night-shift&lt;/code&gt; fuse tripped at $2, which is what it is for. Read the activity page to find what spent it, then raise the key limit on the Keys page. The balance behind it is untouched.&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;&lt;strong&gt;Checkpoint.&lt;/strong&gt; Stop here and you still have a funded model account with two real
limits: a $5 wall and a $2 fuse inside it.&lt;/p&gt;
&lt;h3 id="part-5-wire-the-brain-to-the-chat-15-min"&gt;Part 5, wire the brain to the chat (15 min)&lt;/h3&gt;
&lt;p&gt;Two paste operations connect the account to your chat page, and an allowlist
keeps the model menu at three entries instead of hundreds.
&lt;a href="https://docs.openwebui.com/getting-started/quick-start/connect-a-provider/starting-with-openai-compatible/" class="external-link" rel="noopener"&gt;Open WebUI&amp;rsquo;s own guide to OpenAI-compatible providers&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;
covers the screens if yours differ.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;On the chat page: your name at the bottom left, Admin Panel, Settings,
Connections. Under the OpenAI-compatible section, click add. Two fields
matter:
&lt;ul&gt;
&lt;li&gt;URL: &lt;code&gt;https://openrouter.ai/api/v1&lt;/code&gt;. The &lt;code&gt;/v1&lt;/code&gt; is not optional, and a
missing &lt;code&gt;/v1&lt;/code&gt; is the classic reason no models appear.&lt;/li&gt;
&lt;li&gt;Key: &lt;code&gt;sk-or-v1-YOUR-KEY-HERE&lt;/code&gt;
Filled example: &lt;code&gt;sk-or-v1-9f3ab81c0d2e4f56a7b8c9d0e1f2a3b4&lt;/code&gt; (yours is longer)&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;In the same connection&amp;rsquo;s model IDs field, add exactly these three, one at a
time so each becomes its own entry rather than one comma-separated line:
&lt;code&gt;deepseek/deepseek-v4-pro&lt;/code&gt;, &lt;code&gt;deepseek/deepseek-v4-flash-0731&lt;/code&gt;, and
&lt;code&gt;moonshotai/kimi-k3&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Save. Open a new chat, pick &lt;code&gt;deepseek/deepseek-v4-pro&lt;/code&gt; from the dropdown at
the top, and send: &amp;ldquo;Reply with the single word OK.&amp;rdquo; Any sensible reply means
the wiring works, because models are unreliable narrators and the exact words
do not matter.&lt;/li&gt;
&lt;li&gt;The real proof is elsewhere. Open
&lt;a href="https://openrouter.ai/activity" class="external-link" rel="noopener"&gt;the activity page&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; and the test message is
there as a logged request costing a fraction of a cent. That page is your
receipts drawer from now on.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Table: Part 5 failures and fixes.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Symptom&lt;/th&gt;
 &lt;th&gt;Fix&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;No models appear after saving&lt;/td&gt;
 &lt;td&gt;The URL has to be exactly &lt;code&gt;https://openrouter.ai/api/v1&lt;/code&gt;. Re-check for a missing &lt;code&gt;/v1&lt;/code&gt; or a trailing space.&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Model errors mentioning 401&lt;/td&gt;
 &lt;td&gt;The key pasted wrong. Re-copy it with no spaces around it.&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Model errors mentioning 402&lt;/td&gt;
 &lt;td&gt;The $5 top-up did not complete, or the key&amp;rsquo;s $2 limit is spent. Check the Credits and Keys pages, in that order.&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;No model IDs field in your version&lt;/td&gt;
 &lt;td&gt;Save the connection, then go to Admin Panel, Settings, Models and switch off everything except the three above.&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;&lt;strong&gt;Checkpoint.&lt;/strong&gt; Stop here and you still have a working AI chat served from your
own machine, with receipts. Good finish line for the first sitting.&lt;/p&gt;
&lt;h3 id="part-6-put-it-in-your-pocket-15-min"&gt;Part 6, put it in your pocket (15 min)&lt;/h3&gt;
&lt;p&gt;Tailscale connects your own devices to each other over an encrypted private
network, so your phone can reach the chat page from anywhere.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Never open a port on your home router to reach any of this.&lt;/strong&gt; No step in this
build touches your router.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;On the machine, install
&lt;a href="https://tailscale.com/download/windows" class="external-link" rel="noopener"&gt;Tailscale for Windows&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; and sign in
with an address ending @gmail.com, @outlook.com or similar. An address at your
own business domain makes Tailscale treat you as a business and start a paid
trial.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Click the Tailscale tray icon, behind the caret near the clock if hidden. The
machine&amp;rsquo;s private address is on the line naming this computer, in the form
&lt;code&gt;This device: 100.x.y.z&lt;/code&gt;. Write it down. This page calls it
MACHINE-TAILSCALE-IP.
Filled example: &lt;code&gt;100.101.102.103&lt;/code&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Open the port to your private networks. Windows key, type &lt;code&gt;wf.msc&lt;/code&gt;, Enter.
&lt;strong&gt;Inbound Rules&lt;/strong&gt;, &lt;strong&gt;New Rule&lt;/strong&gt;, Port, TCP, specific local ports &lt;code&gt;3000&lt;/code&gt; (3001
if you switched ports in Part 3), Allow the connection, tick &lt;strong&gt;Private&lt;/strong&gt; only
and untick Domain and Public, name it &lt;code&gt;Open WebUI&lt;/code&gt;, Finish.&lt;/p&gt;
&lt;p&gt;That rule is on the laptop&amp;rsquo;s own firewall, not on your router, and it is
scoped to networks Windows already classifies as private. No port is forwarded
from the internet, so the only thing that can reach port 3000 is a device on a
network the laptop already trusts, which after step 4 means your own tailnet.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;On the phone, install Tailscale, sign in to the &lt;strong&gt;same&lt;/strong&gt; account, flip its
switch on, and allow the VPN configuration it asks for. That&amp;rsquo;s how Tailscale
links your devices.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;The proof: turn &lt;strong&gt;off&lt;/strong&gt; the phone&amp;rsquo;s Wi-Fi so it is on mobile data, open the
phone browser, and go to &lt;code&gt;http://MACHINE-TAILSCALE-IP:3000&lt;/code&gt;.
Filled example: &lt;code&gt;http://100.101.102.103:3000&lt;/code&gt;
Log in with the account from Part 3.&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;Then the step that stops this build failing silently in six months. Tailscale
device keys expire, and
&lt;a href="https://tailscale.com/kb/1028/key-expiry" class="external-link" rel="noopener"&gt;the default expiry period is 180 days&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;;
when a device&amp;rsquo;s key lapses, connections to and from it stop working. For a
machine you are deliberately not looking at, that is an outage with a
half-year fuse already lit. In the Tailscale admin console open &lt;strong&gt;Machines&lt;/strong&gt;,
find this laptop, click the menu icon at the right of its row, and choose
&lt;strong&gt;Disable Key Expiry&lt;/strong&gt;. Tailscale&amp;rsquo;s own documentation names trusted servers
and hard-to-reach devices as exactly the case for doing this.&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Table: Part 6 failures and fixes.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Symptom&lt;/th&gt;
 &lt;th&gt;Fix&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;Page times out on the phone&lt;/td&gt;
 &lt;td&gt;Check the Tailscale switch is on on the phone, that you allowed the VPN permission (toggle the switch again to re-prompt), and that the machine&amp;rsquo;s tray icon shows connected. Then the firewall rule: double-click &lt;code&gt;Open WebUI&lt;/code&gt; in Inbound Rules, Advanced tab, tick all three profiles, because the Tailscale adapter can register as a public network. The machine is still not reachable from the internet; the rule only applies to networks the laptop is already on.&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Wrong address&lt;/td&gt;
 &lt;td&gt;The &lt;code&gt;100.x.y.z&lt;/code&gt; address comes from the Tailscale tray icon, not from &lt;code&gt;ipconfig&lt;/code&gt;. Re-copy it.&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Login page loads but rejects the password&lt;/td&gt;
 &lt;td&gt;It is the Open WebUI account from Part 3, not your Tailscale login.&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;It worked for months, then the phone stopped reaching it&lt;/td&gt;
 &lt;td&gt;The device key expired. Sign the machine in to Tailscale again, then do step 6 so it cannot recur.&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;&lt;strong&gt;Checkpoint.&lt;/strong&gt; Stop here and you still have your machine&amp;rsquo;s chat, with your capped
paid models, on your phone anywhere.&lt;/p&gt;
&lt;h3 id="part-7-hire-the-worker-25-min"&gt;Part 7, hire the worker (25 min)&lt;/h3&gt;
&lt;p&gt;Hermes Agent is a free, open-source agent that lives on the machine, uses your
OpenRouter balance as its brain, and in Part 8 answers Telegram messages and runs
scheduled jobs. Open PowerShell, normal, not administrator.&lt;/p&gt;
&lt;p&gt;This downloads and installs the worker and everything it needs:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-powershell" data-lang="powershell"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nb"&gt;iex &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;irm &lt;/span&gt;&lt;span class="n"&gt;https&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;&lt;span class="p"&gt;//&lt;/span&gt;&lt;span class="nb"&gt;hermes-agent&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="py"&gt;nousresearch&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;com&lt;/span&gt;&lt;span class="p"&gt;/&lt;/span&gt;&lt;span class="n"&gt;install&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;ps1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;What you should see: an installer walks through downloading its components and
finishes without red errors. Then close PowerShell and open a new one so the
&lt;code&gt;hermes&lt;/code&gt; command is picked up.&lt;/p&gt;
&lt;p&gt;This confirms the install and shows which tools the worker has switched on:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-powershell" data-lang="powershell"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;hermes&lt;/span&gt; &lt;span class="n"&gt;tools&lt;/span&gt; &lt;span class="p"&gt;-&lt;/span&gt;&lt;span class="n"&gt;-summary&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;What you should see: a printed summary of enabled tools. Look for &lt;code&gt;web_search&lt;/code&gt;,
because the test below depends on it.&lt;/p&gt;
&lt;p&gt;This stores your OpenRouter key where the worker looks for it:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-powershell" data-lang="powershell"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;hermes&lt;/span&gt; &lt;span class="n"&gt;config&lt;/span&gt; &lt;span class="nb"&gt;set &lt;/span&gt;&lt;span class="n"&gt;OPENROUTER_API_KEY&lt;/span&gt; &lt;span class="nb"&gt;sk-or&lt;/span&gt;&lt;span class="n"&gt;-v1-YOUR-KEY-HERE&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Filled example: &lt;code&gt;hermes config set OPENROUTER_API_KEY sk-or-v1-9f3ab81c0d2e4f56a7b8c9d0e1f2a3b4&lt;/code&gt;&lt;/p&gt;
&lt;p&gt;What you should see: a confirmation that the value was saved.&lt;/p&gt;
&lt;p&gt;This opens the model picker, where you set the worker&amp;rsquo;s default brain:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-powershell" data-lang="powershell"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;hermes&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;What you should see: an interactive picker, arrow keys to move and Enter to
confirm. Choose &lt;strong&gt;OpenRouter&lt;/strong&gt; as the provider, then select or type
&lt;code&gt;deepseek/deepseek-v4-pro&lt;/code&gt;, and confirm it as the default. The picker on your
screen is the ground truth, not this page.&lt;/p&gt;
&lt;p&gt;That slug is unpinned, and OpenRouter published no dated Pro snapshot on 6 August
2026, so put a monthly reminder in your calendar now to check the price and read
one brief closely. If you would rather have the pin than Pro&amp;rsquo;s writing, type
&lt;code&gt;deepseek/deepseek-v4-flash-0731&lt;/code&gt; at this picker instead.&lt;/p&gt;
&lt;p&gt;This is the test that matters, brain plus live web in one shot:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-powershell" data-lang="powershell"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;hermes&lt;/span&gt; &lt;span class="n"&gt;-z&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;Search the web for the current price of Bitcoin and answer in one line with the number.&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;What you should see: one line with an actual current price. That is proof the
worker can think (OpenRouter) and fetch live information (web search), the two
abilities every scheduled job depends on.&lt;/p&gt;
&lt;p&gt;Table: Part 7 failures and fixes.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Symptom&lt;/th&gt;
 &lt;th&gt;Fix&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;The install command itself shows red errors&lt;/td&gt;
 &lt;td&gt;Safe to re-run. Open a new PowerShell, paste the same line again, allow it if your antivirus asks. If it fails twice the same way, take the error text to the docs at hermes-agent.nousresearch.com.&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;&lt;code&gt;hermes&lt;/code&gt; is not recognized&lt;/td&gt;
 &lt;td&gt;You are in the old PowerShell window. Close it, open a new one.&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;The test answers but clearly did not search (no number, or it says it cannot browse)&lt;/td&gt;
 &lt;td&gt;Run &lt;code&gt;hermes tools&lt;/code&gt;, which opens an interactive menu: arrow keys move, space toggles, Enter confirms. Enable &lt;code&gt;web_search&lt;/code&gt; and &lt;code&gt;web_extract&lt;/code&gt; for the CLI, then re-run the test.&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Model or auth errors&lt;/td&gt;
 &lt;td&gt;Re-run &lt;code&gt;hermes model&lt;/code&gt; and re-check every choice, then re-run the &lt;code&gt;hermes config set&lt;/code&gt; line. A key pasted with one character missing fails exactly like this.&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;&lt;strong&gt;Checkpoint.&lt;/strong&gt; Stop here and you still have a working agent on the machine that
answers one-shot questions with live data, on your capped balance.&lt;/p&gt;
&lt;h3 id="part-8-phone-control-and-surviving-a-reboot-25-min"&gt;Part 8, phone control, and surviving a reboot (25 min)&lt;/h3&gt;
&lt;p&gt;This part gives the worker a phone number, in the form of a private Telegram bot
only you can use, and makes everything survive a restart.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;In Telegram, search for &lt;strong&gt;@BotFather&lt;/strong&gt; and open it. The real one has a blue
verification mark and that exact username; impostors exist, so if there is no
mark, search again. Send &lt;code&gt;/newbot&lt;/code&gt;. It asks for a display name, then a
username ending in &lt;code&gt;bot&lt;/code&gt;, and replies with a &lt;strong&gt;token&lt;/strong&gt; like &lt;code&gt;1234567890:AA...&lt;/code&gt;.
That token is a house key. Never screenshot it or paste it anywhere except the
wizard below.&lt;/li&gt;
&lt;li&gt;Get your own Telegram user ID, a number rather than your @username. Search for
&lt;strong&gt;@userinfobot&lt;/strong&gt;, press Start, and it replies with your numeric ID. It is a
widely used third-party utility bot and it sees only your public profile.
Placeholder used below: YOUR-TELEGRAM-ID. Filled example: &lt;code&gt;123456789&lt;/code&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;This starts the connection wizard on the machine:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-powershell" data-lang="powershell"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;hermes&lt;/span&gt; &lt;span class="n"&gt;gateway&lt;/span&gt; &lt;span class="n"&gt;setup&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;What you should see: a wizard. Choose &lt;strong&gt;Telegram&lt;/strong&gt;, paste the bot token, and when
it asks for allowed user IDs, enter YOUR-TELEGRAM-ID and nothing else. The
allowlist is the lock: with only your ID on it the bot ignores every other
Telegram account, and
&lt;a href="https://core.telegram.org/bots" class="external-link" rel="noopener"&gt;bots cannot start conversations&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; in the first
place.&lt;/p&gt;
&lt;p&gt;This registers the worker to start by itself whenever you sign in to Windows, and
needs no administrator rights:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-powershell" data-lang="powershell"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;hermes&lt;/span&gt; &lt;span class="n"&gt;gateway&lt;/span&gt; &lt;span class="n"&gt;install&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;What you should see: confirmation that an autostart entry was created. It uses
Windows Scheduled Tasks, with a Startup-folder fallback.&lt;/p&gt;
&lt;p&gt;This starts it right now:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-powershell" data-lang="powershell"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;hermes&lt;/span&gt; &lt;span class="n"&gt;gateway&lt;/span&gt; &lt;span class="nb"&gt;start
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;What you should see: confirmation the gateway started.&lt;/p&gt;
&lt;p&gt;This checks on it, and it is your go-to command whenever the bot seems quiet:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-powershell" data-lang="powershell"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;hermes&lt;/span&gt; &lt;span class="n"&gt;gateway&lt;/span&gt; &lt;span class="n"&gt;status&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;What you should see: a status readout saying it is running.&lt;/p&gt;
&lt;ol start="3"&gt;
&lt;li&gt;On your phone, open Telegram, find your bot by its username, press Start, and
send &lt;code&gt;Are you there?&lt;/code&gt; A reply arrives. You are talking to your machine.&lt;/li&gt;
&lt;li&gt;The negative test: from any other Telegram account, message the bot. Correct
behaviour is nothing at all. If that account gets a reply, stop and re-run
&lt;code&gt;hermes gateway setup&lt;/code&gt;, because the allowlist is wrong. No second account
handy? Skip it.&lt;/li&gt;
&lt;li&gt;The reboot test, which proves the whole build. Restart Windows, sign in, wait
two minutes, then message the bot and open &lt;code&gt;http://MACHINE-TAILSCALE-IP:3000&lt;/code&gt;
in the phone browser. Both work, because Docker restarts the chat interface
(Part 3&amp;rsquo;s &lt;code&gt;--restart always&lt;/code&gt; plus Part 2&amp;rsquo;s start-on-sign-in) and
&lt;a href="https://hermes-agent.nousresearch.com/docs/user-guide/windows-native" class="external-link" rel="noopener"&gt;the gateway autostarts&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;.
The chat page can take up to five minutes while Docker wakes.&lt;/li&gt;
&lt;/ol&gt;
&lt;h3 id="what-the-builds-uptime-actually-depends-on"&gt;What the build&amp;rsquo;s uptime actually depends on&lt;/h3&gt;
&lt;p&gt;The source material I worked from says that after any restart the machine has to
be signed in to once before the worker is back, and that overnight Windows
updates therefore cost you the occasional morning. Reading the vendors&amp;rsquo; own
documentation, that is both too pessimistic about the common case and silent
about the cases that really do break it.&lt;/p&gt;
&lt;p&gt;Start with the dependency chain, because every link in it is scoped to a sign-in
rather than to a boot. Docker Desktop&amp;rsquo;s autostart setting is worded &amp;ldquo;Start Docker
Desktop when you sign in to your machine&amp;rdquo;, and it is off until you tick it.
Hermes registers itself with &lt;code&gt;schtasks /Create /SC ONLOGON&lt;/code&gt;, which is a logon
trigger. Microsoft&amp;rsquo;s own reference is explicit that the alternative,
&lt;a href="https://learn.microsoft.com/en-us/windows/win32/taskschd/boottrigger" class="external-link" rel="noopener"&gt;a boot trigger&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;,
&amp;ldquo;starts a task when the system is booted&amp;rdquo; and that &amp;ldquo;only a member of the
Administrators group can create a task with a boot trigger&amp;rdquo;. Nothing in the
default build has one. The uptime of this machine is the uptime of an interactive
Windows session, and every reboot symptom in the tables above is a consequence of
that single fact.&lt;/p&gt;
&lt;p&gt;Now the part the inherited limitation gets wrong. Windows has a feature called
Automatic Restart Sign-On, and
&lt;a href="https://learn.microsoft.com/en-us/windows-server/identity/ad-ds/manage/component-updates/winlogon-automatic-restart-sign-on--arso-" class="external-link" rel="noopener"&gt;Microsoft documents&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;
that &amp;ldquo;when Windows Update initiates an automatic reboot, ARSO extracts the
currently logged in user&amp;rsquo;s derived credentials, persists it to disk, and
configures Autologon for the user&amp;rdquo;, after which &amp;ldquo;the last interactive user is
automatically logged in and the session is locked&amp;rdquo;. A locked session is still a
signed-in one, so the logon trigger fires and the gateway comes back. The policy
&amp;ldquo;is enabled by default&amp;rdquo; if you have not configured it. So the overnight Windows
update, the scenario the limitation is usually written about, mostly does not
cost you a morning.&lt;/p&gt;
&lt;p&gt;What does cost you a morning is narrower and worth knowing precisely. ARSO &amp;ldquo;only
occurs if the last interactive user didn&amp;rsquo;t sign out before the restart or
shutdown&amp;rdquo;, so signing out before bed defeats it. On a machine joined to Active
Directory or Microsoft Entra ID &amp;ldquo;this policy only applies to Windows Update
restarts&amp;rdquo;, which means a power cut or a manual restart on a work-managed laptop
gets nothing. It also fails where the account is disabled, where logon hours
apply, or where the user must change their password at next sign-in, which is a
realistic six-month event on a personal Microsoft account. When a morning goes
missing and you want to know why rather than guess, the LSA Operational log in
Event Viewer records event 322 for a failed ARSO configuration and 320 or 321
when it worked.&lt;/p&gt;
&lt;p&gt;One honesty note about that page: it says in one place that ARSO is &amp;ldquo;opted out
for Client SKUs&amp;rdquo; and in another that the policy &amp;ldquo;is enabled by default&amp;rdquo;. I have
not been able to resolve the contradiction, so treat ARSO as probable rather than
guaranteed and let the event log settle it on your own machine. If you want the
guarantee instead of the probability, the fix is to stop depending on a logon
trigger: register the gateway as a real Windows service, which the Hermes
Windows documentation itself points at, or add a second task with a boot trigger
and the administrator rights that requires. Neither removes the Docker Desktop
dependency, which is the harder half.&lt;/p&gt;
&lt;p&gt;Table: Part 8 failures and fixes.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Symptom&lt;/th&gt;
 &lt;th&gt;Fix&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;The bot never replies&lt;/td&gt;
 &lt;td&gt;Run &lt;code&gt;hermes gateway status&lt;/code&gt; on the machine. If it is not running, &lt;code&gt;hermes gateway start&lt;/code&gt;. If it is running, re-run &lt;code&gt;hermes gateway setup&lt;/code&gt; and re-paste the token carefully.&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Replies stopped after the reboot&lt;/td&gt;
 &lt;td&gt;Sign in to Windows first, then &lt;code&gt;hermes gateway status&lt;/code&gt;, and &lt;code&gt;hermes gateway install&lt;/code&gt; again if needed.&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;The negative test got a reply&lt;/td&gt;
 &lt;td&gt;The allowlist holds the wrong value. It has to be the numeric ID, not the @username. Re-run &lt;code&gt;hermes gateway setup&lt;/code&gt;.&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;&lt;strong&gt;Checkpoint.&lt;/strong&gt; The machine is complete. It thinks on a capped budget, reaches
your pocket, ignores strangers, and survives a reboot. It has no job yet.&lt;/p&gt;
&lt;h2 id="living-with-it"&gt;Living with it&lt;/h2&gt;
&lt;h3 id="the-first-job"&gt;The first job&lt;/h3&gt;
&lt;p&gt;Create the job from your phone, in the bot chat, because
&lt;a href="https://hermes-agent.nousresearch.com/docs/user-guide/features/cron" class="external-link" rel="noopener"&gt;jobs report back to wherever they were created&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;This message creates the scheduled job. Copy the whole block into Telegram and
replace [YOUR NICHE], brackets included, with your own category. Do not retype
the quotes, because phone keyboards swap them for curly ones the command may not
parse:&lt;/p&gt;
&lt;pre tabindex="0"&gt;&lt;code&gt;/cron add &amp;#34;every day at 7:30&amp;#34; &amp;#34;Search the web for what changed in the [YOUR NICHE] market in the last 24 hours: new competitor angles, price moves, trending products, and any ad platform policy news. Send 5 short bullets, each with its source link. End with one hook angle worth testing today, written as a single sentence.&amp;#34;
&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;What you should see: nothing yet, because the placeholder is still in it. Here is
the same command filled in:&lt;/p&gt;
&lt;pre tabindex="0"&gt;&lt;code&gt;/cron add &amp;#34;every day at 7:30&amp;#34; &amp;#34;Search the web for what changed in the posture corrector market in the last 24 hours: new competitor angles, price moves, trending products, and any ad platform policy news. Send 5 short bullets, each with its source link. End with one hook angle worth testing today, written as a single sentence.&amp;#34;
&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;What you should see: the bot confirms the job was created. Send &lt;code&gt;/cron list&lt;/code&gt;, and
the next run has to say tomorrow morning at 07:30 local time. A strange hour
means the machine&amp;rsquo;s clock or timezone is off. Fix it in Windows Settings, Time
and language, then remove and re-add the job.&lt;/p&gt;
&lt;p&gt;To run it once immediately, message the bot in plain language: &amp;ldquo;Run the morning
brief job now.&amp;rdquo; What you should see is an acknowledgement, then the brief itself
a minute or two later. Read the first one with a working eye. The question is not
whether it&amp;rsquo;s perfect. It&amp;rsquo;s whether one of those five bullets just handed you
tomorrow&amp;rsquo;s angle.&lt;/p&gt;
&lt;p&gt;Management happens in the same chat, in plain language or with &lt;code&gt;/cron&lt;/code&gt;:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Different time or days: &amp;ldquo;Change the morning brief to weekdays at 7:00.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;Pause it: &amp;ldquo;Pause the morning brief.&amp;rdquo; Resume the same way.&lt;/li&gt;
&lt;li&gt;See everything scheduled: &lt;code&gt;/cron list&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;Delete it: &amp;ldquo;Remove the morning brief job.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;If it cannot tell which job you mean, &lt;code&gt;/cron list&lt;/code&gt; shows the name and the id.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;One rule as you add more. &lt;strong&gt;One job, one purpose.&lt;/strong&gt; Five small pausable jobs beat
one long mush message you stop reading, and they fail one at a time instead of
all at once.&lt;/p&gt;
&lt;p&gt;Table: First-night failures and fixes.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Symptom&lt;/th&gt;
 &lt;th&gt;Fix&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;No confirmation after &lt;code&gt;/cron add&lt;/code&gt;&lt;/td&gt;
 &lt;td&gt;On the machine: &lt;code&gt;hermes gateway status&lt;/code&gt;, then &lt;code&gt;hermes gateway start&lt;/code&gt;.&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Acknowledged, but no brief after five minutes&lt;/td&gt;
 &lt;td&gt;&lt;code&gt;/cron list&lt;/code&gt; to confirm the job exists, then ask the bot what happened to its last run.&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Bullets arrive without links, or read as invented&lt;/td&gt;
 &lt;td&gt;Web tools are off. Re-run Part 7&amp;rsquo;s search test on the machine and re-enable &lt;code&gt;web_search&lt;/code&gt; with &lt;code&gt;hermes tools&lt;/code&gt;.&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Silence tomorrow at 07:30&lt;/td&gt;
 &lt;td&gt;The laptop is sitting at the Windows sign-in screen after an overnight update. Sign in and wait two minutes.&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;h3 id="weekly-and-monthly-maintenance"&gt;Weekly and monthly maintenance&lt;/h3&gt;
&lt;p&gt;The weekly pass takes five minutes. Did the morning brief arrive today, since its
absence is the monitoring system? Glance at
&lt;a href="https://openrouter.ai/activity" class="external-link" rel="noopener"&gt;the activity page&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; for shape and total, not
forensics. Run &lt;code&gt;/cron list&lt;/code&gt; and ask whether everything scheduled is still
something you read. Once a month, check free disk space is above 10 GB.&lt;/p&gt;
&lt;p&gt;Then the honest limitation about restarts. After any reboot the machine has to be
signed in to once before jobs and the bot come back, because
&lt;a href="https://hermes-agent.nousresearch.com/docs/user-guide/windows-native" class="external-link" rel="noopener"&gt;the worker&amp;rsquo;s autostart&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;
runs at sign-in and not before it. Windows updates restart at night on their own
schedule, so an occasional missed morning is normal rather than a fault. There is
a trade available and it has a real cost: Windows can be set to sign your user in
automatically, which is the setting you reach by searching for &amp;ldquo;netplwiz&amp;rdquo;, and
the price is that anyone who opens the laptop is signed in as you. Leave it alone
unless the machine is behind a door you lock.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Back up the two things you cannot recreate.&lt;/strong&gt; The uninstall section below is
honest that deleting the chat volume is permanent, and a failing disk is the same
event without the warning. The worker&amp;rsquo;s configuration and stored key live in
&lt;code&gt;%LOCALAPPDATA%\hermes&lt;/code&gt; per
&lt;a href="https://hermes-agent.nousresearch.com/docs/user-guide/windows-native" class="external-link" rel="noopener"&gt;the Hermes Windows-native page&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;,
and that is an ordinary folder: paste the path into the File Explorer address bar
and copy it elsewhere. The chat history lives in a Docker volume named
&lt;code&gt;open-webui&lt;/code&gt;, which is not a folder you can browse, so use
&lt;a href="https://docs.docker.com/engine/storage/volumes/" class="external-link" rel="noopener"&gt;Docker&amp;rsquo;s own back up, restore, or migrate procedure&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;,
which mounts the volume into a throwaway container and writes a tar file to a
directory you choose. Do both monthly and keep the copy off the laptop.&lt;/p&gt;
&lt;p&gt;The monthly pass is mostly one command sequence.
&lt;a href="https://docs.openwebui.com/getting-started/updating" class="external-link" rel="noopener"&gt;Open WebUI&amp;rsquo;s update procedure&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;
works because your chats and settings live in the storage volume rather than in
the container, so removing the container loses nothing:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-powershell" data-lang="powershell"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;docker&lt;/span&gt; &lt;span class="nb"&gt;rm &lt;/span&gt;&lt;span class="o"&gt;-f&lt;/span&gt; &lt;span class="nb"&gt;open-webui&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;docker&lt;/span&gt; &lt;span class="n"&gt;pull&lt;/span&gt; &lt;span class="n"&gt;ghcr&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;io&lt;/span&gt;&lt;span class="p"&gt;/&lt;/span&gt;&lt;span class="nb"&gt;open-webui&lt;/span&gt;&lt;span class="p"&gt;/&lt;/span&gt;&lt;span class="nb"&gt;open-webui&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;&lt;span class="n"&gt;main&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;docker&lt;/span&gt; &lt;span class="n"&gt;run&lt;/span&gt; &lt;span class="n"&gt;-d&lt;/span&gt; &lt;span class="n"&gt;-p&lt;/span&gt; &lt;span class="mf"&gt;3000&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;&lt;span class="mf"&gt;8080&lt;/span&gt; &lt;span class="n"&gt;-v&lt;/span&gt; &lt;span class="nb"&gt;open-webui&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;&lt;span class="p"&gt;/&lt;/span&gt;&lt;span class="n"&gt;app&lt;/span&gt;&lt;span class="p"&gt;/&lt;/span&gt;&lt;span class="n"&gt;backend&lt;/span&gt;&lt;span class="p"&gt;/&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt; &lt;span class="p"&gt;-&lt;/span&gt;&lt;span class="n"&gt;-name&lt;/span&gt; &lt;span class="nb"&gt;open-webui&lt;/span&gt; &lt;span class="p"&gt;-&lt;/span&gt;&lt;span class="n"&gt;-restart&lt;/span&gt; &lt;span class="n"&gt;always&lt;/span&gt; &lt;span class="n"&gt;ghcr&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;io&lt;/span&gt;&lt;span class="p"&gt;/&lt;/span&gt;&lt;span class="nb"&gt;open-webui&lt;/span&gt;&lt;span class="p"&gt;/&lt;/span&gt;&lt;span class="nb"&gt;open-webui&lt;/span&gt;&lt;span class="err"&gt;:&lt;/span&gt;&lt;span class="n"&gt;main&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;What you should see: the old container is removed, a fresh image downloads, and a
new container id prints. &lt;code&gt;http://localhost:3000&lt;/code&gt; logs you in with the same
account and the same history.&lt;/p&gt;
&lt;p&gt;This updates the worker in place:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-powershell" data-lang="powershell"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;hermes&lt;/span&gt; &lt;span class="n"&gt;update&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;What you should see: it checks for a newer version, updates if there is one, and
reports what it did.&lt;/p&gt;
&lt;p&gt;Then message the bot once to confirm it answers, and skim the model prices
against &lt;a href="https://openrouter.ai/models" class="external-link" rel="noopener"&gt;OpenRouter&amp;rsquo;s model list&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;. Providers
reprice quietly, and this is the pass where you would notice.&lt;/p&gt;
&lt;h3 id="troubleshooting"&gt;Troubleshooting&lt;/h3&gt;
&lt;p&gt;Table: Consolidated symptoms, likely causes and fixes for the running machine.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Symptom&lt;/th&gt;
 &lt;th&gt;Likely cause&lt;/th&gt;
 &lt;th&gt;Fix&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;No morning brief, bot silent&lt;/td&gt;
 &lt;td&gt;Machine asleep, powered off, or at the sign-in screen&lt;/td&gt;
 &lt;td&gt;Wake it, sign in, then &lt;code&gt;hermes gateway status&lt;/code&gt;&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;No brief, but the bot answers chat&lt;/td&gt;
 &lt;td&gt;The job is paused or errored&lt;/td&gt;
 &lt;td&gt;&lt;code&gt;/cron list&lt;/code&gt;, then &amp;ldquo;Run the morning brief job now&amp;rdquo; and read the error&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Bot silent, machine running&lt;/td&gt;
 &lt;td&gt;Gateway down&lt;/td&gt;
 &lt;td&gt;&lt;code&gt;hermes gateway status&lt;/code&gt;, then &lt;code&gt;hermes gateway start&lt;/code&gt;. Re-run &lt;code&gt;hermes gateway install&lt;/code&gt; if it did not survive a reboot&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Chat page dead from the phone&lt;/td&gt;
 &lt;td&gt;Tailscale off, an expired device key, or the firewall rule&lt;/td&gt;
 &lt;td&gt;Check both Tailscale switches and the machine&amp;rsquo;s status in the admin console, then the &lt;code&gt;Open WebUI&lt;/code&gt; inbound rule from Part 6&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Chat page dead everywhere&lt;/td&gt;
 &lt;td&gt;Container stopped&lt;/td&gt;
 &lt;td&gt;&lt;code&gt;docker ps&lt;/code&gt; on the machine. If it is empty, open Docker Desktop and check the start-on-sign-in setting from Part 2&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Errors mentioning 401&lt;/td&gt;
 &lt;td&gt;Key wrong or revoked&lt;/td&gt;
 &lt;td&gt;Re-paste the key with &lt;code&gt;hermes config set OPENROUTER_API_KEY ...&lt;/code&gt; and in the Open WebUI connection&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Errors mentioning 402&lt;/td&gt;
 &lt;td&gt;Key fuse blown at $2, or balance empty&lt;/td&gt;
 &lt;td&gt;Keys page first, then Credits. Check the activity page for what spent it before raising anything&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Answers suddenly slow or strange&lt;/td&gt;
 &lt;td&gt;Model changed upstream&lt;/td&gt;
 &lt;td&gt;Check the model name still exists, and pin a dated snapshot if the provider publishes one&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Every component here belongs to somebody else, and all of them ship changes
without asking you. So the more useful table is the one that maps a breakage to
the document that outranks this page.&lt;/p&gt;
&lt;p&gt;Table: Where the ground truth lives when something upstream changes. In every row, the linked source wins over anything written here.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;What broke&lt;/th&gt;
 &lt;th&gt;Where the ground truth lives&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;A model name starts returning &amp;ldquo;not found&amp;rdquo;&lt;/td&gt;
 &lt;td&gt;&lt;a href="https://openrouter.ai/models" class="external-link" rel="noopener"&gt;OpenRouter&amp;rsquo;s model list&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;. Find the current slug, then update &lt;code&gt;hermes model&lt;/code&gt; and the Open WebUI allowlist&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;A Hermes command or screen no longer matches this page&lt;/td&gt;
 &lt;td&gt;&lt;a href="https://hermes-agent.nousresearch.com/docs/reference/cli-commands" class="external-link" rel="noopener"&gt;The Hermes CLI reference&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; and the picker on your screen, both of which outrank this file&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Open WebUI screens moved&lt;/td&gt;
 &lt;td&gt;&lt;a href="https://docs.openwebui.com/getting-started/quick-start/" class="external-link" rel="noopener"&gt;Open WebUI&amp;rsquo;s own quick start&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Prices drift from the tables above&lt;/td&gt;
 &lt;td&gt;&lt;a href="https://openrouter.ai/api/v1/models" class="external-link" rel="noopener"&gt;OpenRouter&amp;rsquo;s model API&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; is the live source, and this page&amp;rsquo;s date stamp tells you how stale my copy is&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Docker or Tailscale licence terms move&lt;/td&gt;
 &lt;td&gt;&lt;a href="https://docs.docker.com/subscription/desktop-license/" class="external-link" rel="noopener"&gt;Docker&amp;rsquo;s licence page&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; and &lt;a href="https://tailscale.com/pricing" class="external-link" rel="noopener"&gt;Tailscale&amp;rsquo;s pricing page&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;h3 id="the-full-uninstall"&gt;The full uninstall&lt;/h3&gt;
&lt;p&gt;The exit gets the same detail as the entrance, in reverse build order.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;1. Remove the jobs.&lt;/strong&gt; In Telegram, &lt;code&gt;/cron list&lt;/code&gt;, then remove each one in plain
language: &amp;ldquo;Remove the morning brief job.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;2. Stop the worker&amp;rsquo;s autostart.&lt;/strong&gt; In PowerShell, &lt;code&gt;hermes gateway stop&lt;/code&gt;, then
&lt;code&gt;hermes gateway uninstall&lt;/code&gt;, which removes the Scheduled Task or startup entry.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;3. Delete the worker.&lt;/strong&gt; Delete the folder &lt;code&gt;%LOCALAPPDATA%\hermes&lt;/code&gt; by pasting
that into the File Explorer address bar. On native Windows that single folder is
the whole of it:
&lt;a href="https://hermes-agent.nousresearch.com/docs/user-guide/windows-native" class="external-link" rel="noopener"&gt;the Hermes Windows-native page&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;
documents it as the root holding the config, the stored credentials, the skills,
the sessions and the logs, with the install in a subfolder beneath. A &lt;code&gt;.hermes&lt;/code&gt;
folder in your user profile exists only if you repointed &lt;code&gt;HERMES_HOME&lt;/code&gt; there.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;4. Retire the bot.&lt;/strong&gt; In @BotFather, &lt;code&gt;/mybots&lt;/code&gt;, select the bot, and revoke the
token or delete the bot outright. A dead token is a dead door.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;5. Remove the chat interface and its stored chats.&lt;/strong&gt; The volume deletion is the
permanent part, so be sure before you run the second line:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-powershell" data-lang="powershell"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;docker&lt;/span&gt; &lt;span class="nb"&gt;rm &lt;/span&gt;&lt;span class="o"&gt;-f&lt;/span&gt; &lt;span class="nb"&gt;open-webui&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="n"&gt;docker&lt;/span&gt; &lt;span class="n"&gt;volume&lt;/span&gt; &lt;span class="nb"&gt;rm open-webui&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;What you should see: each command echoes the name it removed. Chats and the admin
account are gone with the volume.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;6. Uninstall Docker Desktop.&lt;/strong&gt; Settings, Apps, Installed apps, Docker Desktop,
Uninstall.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;7. Close the model account.&lt;/strong&gt; Delete the &lt;code&gt;night-shift&lt;/code&gt; key on the Keys page.
Spend any remaining credit or accept its loss; &lt;a href="https://openrouter.ai/terms" class="external-link" rel="noopener"&gt;the terms&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;
cover refunds and expiry in section 4.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;8. Remove Tailscale, then restore the power settings.&lt;/strong&gt; Uninstall the app on
the machine and the phone, and remove both devices from the machines list in your
Tailscale admin console. Then Part 1 in reverse.&lt;/p&gt;
&lt;p&gt;The laptop is now exactly the machine it was before Part 1.&lt;/p&gt;
&lt;h2 id="what-i-do-not-know-yet"&gt;What I do not know yet&lt;/h2&gt;
&lt;p&gt;Things I expect to learn only by running this. What my laptop actually draws,
which turns a bounded inference into one row of fact. What a search-and-summarise
run really costs across a real tool loop, which the activity page will answer
within a week and which may be several times the estimate above in either
direction. How often Windows restarts itself overnight. Whether Hermes ships with
web search on by default on a fresh Windows install, which Part 7 checks twice
precisely because I could not confirm the default in the documentation. Whether
Docker Desktop installs on Windows Home, which is why Part 2 comes before any
spending. And whether the morning message is something I read on day thirty or
something I stopped opening on day nine, which is the only question here that no
amount of arithmetic can answer.&lt;/p&gt;
&lt;p&gt;The costing stands on its own regardless. Roughly £7 to £16 of electricity a
year, four or five dollars of model spend across the same year, $0 in licences
below Docker&amp;rsquo;s employee and revenue lines, a one-time $5.80, and a prepaid
balance that is the hard ceiling on everything the machine can ever spend.
Against $120 to $144 a year to rent the same job a server, or $0 on a free tier
that puts the disk in someone else&amp;rsquo;s building. If those numbers make the idea
worth three hours of your evening, the eight parts are above. If they don&amp;rsquo;t,
you&amp;rsquo;ve lost the ten minutes it took to read the arithmetic, which beats finding
out in month three.&lt;/p&gt;</content:encoded></item><item><title>4,600 people sorted human text from machine text at 50 to 52%</title><link>https://therezaali.com/writing/ai-polish-and-perceived-humanness/</link><guid isPermaLink="true">https://therezaali.com/writing/ai-polish-and-perceived-humanness/</guid><pubDate>Thu, 06 Aug 2026 00:00:00 +0200</pubDate><author>contact@therezaali.com (Mo RezaAli)</author><category>essay</category><description>Polish is not a detection cue, because detection barely happens. The rougher cut of my ad did win on hook rate, and I had explained why it won incorrectly.</description><content:encoded>&lt;p&gt;Nobody detects AI content by spotting polish, because almost nobody detects it
at all. In
&lt;a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC10089155/" class="external-link" rel="noopener"&gt;Jakesch, Hancock and Naaman&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;
(PNAS, March 2023), 4,600 participants across six experiments &amp;ldquo;identified the
source of a self-presentation with only 50 to 52% accuracy.&amp;rdquo; In
&lt;a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC8872790/" class="external-link" rel="noopener"&gt;Nightingale and Farid&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;
(PNAS, February 2022), 315 participants sorting synthetic faces from real ones
scored 48.2%, &amp;ldquo;close to chance performance of 50%.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;That matters to me because I built a conclusion on the opposite assumption.
Meta defines hook rate as arithmetic, not a vibe: the metric &amp;ldquo;counts the number
of 3-second video plays and divides it by ad impressions.&amp;rdquo; I ran two cuts of the
same AI UGC ad. In the second, the creator stumbled over the opening sentence,
paused, and started again. I nearly cut that clip, left it in, and that version
produced a noticeably higher hook rate. I explained it to myself by saying
polished AI content has become easy to recognise and the imperfection defeated
recognition. The recognition half of that sentence is false. The stumble did not
make my ad harder to identify as AI, and the paper I reached for to explain it
turns out to argue against me on the one cue I used. What survives is a narrower
mechanism with better evidence behind it, and it happens to sit inside the exact
window the metric samples.&lt;/p&gt;
&lt;p&gt;This was my test, my budget, one product, a handful of creatives, a few weeks.
I published the pattern in
&lt;a href="https://x.com/Mo_ali/status/2077860368307925061" class="external-link" rel="noopener"&gt;the X post this article is built from&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;
and did not publish the hook rate figure there. I am not going to invent one
here. What I can do is take the test apart, which is worth more than the number
would have been.&lt;/p&gt;
&lt;h2 id="the-design-fault-is-in-the-first-three-seconds"&gt;The design fault is in the first three seconds&lt;/h2&gt;
&lt;p&gt;Hook rate samples the opening three seconds. The stumble, the pause and the
restart are in the opening three seconds. The variable I believed I was testing,
global polish, and the variable I actually changed, the content of the
measurement window, were the same edit. Meta&amp;rsquo;s own
&lt;a href="https://www.facebook.com/business/help/290009911394576" class="external-link" rel="noopener"&gt;A/B testing best practices&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;
page states the requirement I broke: &amp;ldquo;You&amp;rsquo;ll have more conclusive results for
your test if your ad sets are identical except for the variable that you&amp;rsquo;re
testing.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;The metric is also narrower than the argument I hung on it.&lt;/p&gt;
&lt;p&gt;Table: How Meta defines the four metrics involved, in Meta&amp;rsquo;s own words, read 6 August 2026.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Metric&lt;/th&gt;
 &lt;th&gt;Meta&amp;rsquo;s definition&lt;/th&gt;
 &lt;th&gt;Status&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;&lt;a href="https://www.facebook.com/business/help/1591938568393543" class="external-link" rel="noopener"&gt;Hook rate&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/td&gt;
 &lt;td&gt;3-second video plays divided by ad impressions&lt;/td&gt;
 &lt;td&gt;&amp;ldquo;This metric is in development&amp;rdquo;&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;&lt;a href="https://www.facebook.com/business/help/1216404939613119" class="external-link" rel="noopener"&gt;Hold rate&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/td&gt;
 &lt;td&gt;ThruPlays divided by 3-second video plays&lt;/td&gt;
 &lt;td&gt;&amp;ldquo;This metric is in development&amp;rdquo;&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;&lt;a href="https://www.facebook.com/business/help/743427195703387" class="external-link" rel="noopener"&gt;3-second video plays&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/td&gt;
 &lt;td&gt;Played at least 3 seconds, or 97% of length if shorter. &amp;ldquo;Time spent replaying the video for a single impression won&amp;rsquo;t be included.&amp;rdquo;&lt;/td&gt;
 &lt;td&gt;Stable&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;&lt;a href="https://www.facebook.com/business/help/471190536725647" class="external-link" rel="noopener"&gt;ThruPlay&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;&lt;/td&gt;
 &lt;td&gt;Played at least 15 seconds, or 97% of length if shorter than 15 seconds&lt;/td&gt;
 &lt;td&gt;Stable&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The denominator is impressions, not reach and not video plays, so a fatiguing
creative drags its own hook rate down through repeat impressions with no change
to the creative at all. And replays inside an impression are excluded, so if the
stumble makes someone rewind, the metric pays nothing for it.&lt;/p&gt;
&lt;h2 id="what-people-actually-detect-measured"&gt;What people actually detect, measured&lt;/h2&gt;
&lt;p&gt;The popular version of this research is &amp;ldquo;people cannot tell AI faces from real
ones.&amp;rdquo; That overstates it, and the correction is worth more than the statistic.
Nightingale and Farid ran a second experiment with training and trial-by-trial
feedback, and accuracy rose to 59.0% (95% CI 57.7% to 60.4%), which is above
chance and reliably so. At chance untrained, barely above chance trained. In
their third experiment, participants rated synthetic faces 4.82 for
trustworthiness against 4.48 for real faces, &amp;ldquo;only 7.7% more trustworthy&amp;rdquo; but
significant at t(222) = 14.6, P &amp;lt; 0.001, d = 0.49.&lt;/p&gt;
&lt;p&gt;Table: What three primary studies measured about human detection of synthetic content.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Study&lt;/th&gt;
 &lt;th&gt;Stimuli&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Accuracy&lt;/th&gt;
 &lt;th&gt;Direction of the error&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;Nightingale &amp;amp; Farid 2022, N = 315&lt;/td&gt;
 &lt;td&gt;StyleGAN2 still faces&lt;/td&gt;
 &lt;td style="text-align: right"&gt;48.2%&lt;/td&gt;
 &lt;td&gt;Synthetic rated more trustworthy&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Nightingale &amp;amp; Farid 2022, N = 219, trained&lt;/td&gt;
 &lt;td&gt;Same, with feedback&lt;/td&gt;
 &lt;td style="text-align: right"&gt;59.0%&lt;/td&gt;
 &lt;td&gt;Training helps a little&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;&lt;a href="https://pubmed.ncbi.nlm.nih.gov/37955384/" class="external-link" rel="noopener"&gt;Miller et al. 2023&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;, N = 124&lt;/td&gt;
 &lt;td&gt;White AI still faces&lt;/td&gt;
 &lt;td style="text-align: right"&gt;No figure published&lt;/td&gt;
 &lt;td&gt;AI faces &amp;ldquo;judged as human more often than actual human faces&amp;rdquo;&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Jakesch et al. 2023, N = 4,600&lt;/td&gt;
 &lt;td&gt;Self-presentation text&lt;/td&gt;
 &lt;td style="text-align: right"&gt;50 to 52%&lt;/td&gt;
 &lt;td&gt;Human roughness misread as machine&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Miller and colleagues called the third row &amp;ldquo;AI hyperrealism&amp;rdquo; and found that
&amp;ldquo;people who made the most errors in this task were the most confident (a
Dunning-Kruger effect).&amp;rdquo; Two ceilings travel with it. The effect was found for
White AI faces specifically, because generators are trained disproportionately
on White faces. And the distinguishing attributes do exist: they &amp;ldquo;permitted high
accuracy using machine learning.&amp;rdquo; People have the signal and misread it, which
is a fact about human heuristics rather than about how clean the output is.&lt;/p&gt;
&lt;h2 id="the-heuristic-runs-the-other-way"&gt;The heuristic runs the other way&lt;/h2&gt;
&lt;p&gt;Jakesch maps the heuristic feature by feature, and read properly the map
convicts my explanation. The paper fits two logistic regressions over the same
language features: one predicting whether participants &lt;em&gt;called&lt;/em&gt; a
self-presentation AI-generated, one predicting whether it actually &lt;em&gt;was&lt;/em&gt;. Where
the two odds ratios point the same way, the cue works. Where they point opposite
ways, people are not merely guessing, they are confidently backwards.&lt;/p&gt;
&lt;p&gt;Table: Jakesch, Hancock and Naaman 2023, Table 1. Odds ratios above 1 mean more likely. Model 1 covers 38,866 judgments, model 2 covers 4,690 self-presentations. Rows marked ns were not significant in model 2.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Feature&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Rated AI&lt;/th&gt;
 &lt;th style="text-align: right"&gt;Actually AI&lt;/th&gt;
 &lt;th&gt;Reading&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;Nonsensical content&lt;/td&gt;
 &lt;td style="text-align: right"&gt;1.105&lt;/td&gt;
 &lt;td style="text-align: right"&gt;1.233&lt;/td&gt;
 &lt;td&gt;Machine signal, read correctly&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Repetitive content&lt;/td&gt;
 &lt;td style="text-align: right"&gt;1.083&lt;/td&gt;
 &lt;td style="text-align: right"&gt;1.470&lt;/td&gt;
 &lt;td&gt;Machine signal, read correctly&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Conversational words&lt;/td&gt;
 &lt;td style="text-align: right"&gt;0.947&lt;/td&gt;
 &lt;td style="text-align: right"&gt;0.898&lt;/td&gt;
 &lt;td&gt;Human signal, read correctly&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Grammatical issues&lt;/td&gt;
 &lt;td style="text-align: right"&gt;1.048&lt;/td&gt;
 &lt;td style="text-align: right"&gt;0.851&lt;/td&gt;
 &lt;td&gt;Read as machine, actually human&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Rare bigrams&lt;/td&gt;
 &lt;td style="text-align: right"&gt;1.042&lt;/td&gt;
 &lt;td style="text-align: right"&gt;0.666&lt;/td&gt;
 &lt;td&gt;Read as machine, actually human&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Long words&lt;/td&gt;
 &lt;td style="text-align: right"&gt;1.034&lt;/td&gt;
 &lt;td style="text-align: right"&gt;0.783&lt;/td&gt;
 &lt;td&gt;Read as machine, actually human&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Contractions&lt;/td&gt;
 &lt;td style="text-align: right"&gt;0.947&lt;/td&gt;
 &lt;td style="text-align: right"&gt;1.134&lt;/td&gt;
 &lt;td&gt;Read as human, actually machine&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;First-person pronouns&lt;/td&gt;
 &lt;td style="text-align: right"&gt;0.925&lt;/td&gt;
 &lt;td style="text-align: right"&gt;0.992 ns&lt;/td&gt;
 &lt;td&gt;Read as human, diagnostic of nothing&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Authentic or spontaneous words&lt;/td&gt;
 &lt;td style="text-align: right"&gt;0.946&lt;/td&gt;
 &lt;td style="text-align: right"&gt;0.945 ns&lt;/td&gt;
 &lt;td&gt;Read as human, diagnostic of nothing&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Three features are functional cues and that is all: &amp;ldquo;Only nonsense, repetition,
and conversational words were functional cues.&amp;rdquo; The paper is blunt about the
rest. &amp;ldquo;Participants were 5% more likely to rate self-presentations with
grammatical issues as AI-generated, although grammatically flawed
self-presentations were, in fact, 15% less likely to be AI-generated.&amp;rdquo;
&amp;ldquo;Participants often rated self-presentations with long words or rare bigrams as
generated by AI, while most self-presentations with long words or rare bigrams
had been written by humans.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Read the direction twice, because it inverts the folk model. Roughness is not
what people mistake for humanity. Roughness is what people mistake for a broken
machine, and in this corpus roughness was a genuine human tell that they read
backwards. The mirror image is on the contractions row: the single most
casual-sounding feature on the list was 13% more likely to appear in the
generated text, and people took it as evidence of a person. The cues they do
read as human, first-person speech and spontaneous-sounding words, are not
significantly associated with either source. So the heuristic is not just weak.
It is wrong in a specific direction: it penalises the mess that humans actually
make and rewards the ease that machines actually produce.&lt;/p&gt;
&lt;p&gt;Then the part that should make anyone in advertising sit up. The authors trained
classifiers on participants&amp;rsquo; judgments and used them to select generated text
that hit the cues people read as human. Those optimised versions &amp;ldquo;were rated as
human more often than regular generated self-presentations (65.7% vs. 51.6%)&amp;rdquo;
and, more sharply, were &amp;ldquo;more likely to be seen as human than self-presentations
that were actually written by humans (65.7% vs. 51.7%).&amp;rdquo; Optimising for the
heuristic beats being human at seeming human.&lt;/p&gt;
&lt;p&gt;That finding is real and exploitable. It is not a licence for the edit I made.
An optimiser aimed at this heuristic would smooth the grammar, shorten the
words, drop the rare phrasing and push conversational register up. It would cut
the stumble. I had been telling myself the stumble was the video-native form of
&amp;ldquo;more human than human&amp;rdquo;, and on the only measured version of that heuristic it
is closer to the opposite.&lt;/p&gt;
&lt;p&gt;One hedge in my own favour, and it does not rescue the argument. Jakesch measured
written self-presentations, and &amp;ldquo;grammatical issues&amp;rdquo; there is a crowdworker
label on text, not a spoken false start in a video. I am not entitled to map one
onto the other. I am also not entitled to cite the paper as support when the
feature nearest to what I did carries the wrong sign, which is what I was doing.&lt;/p&gt;
&lt;h2 id="the-mechanism-that-does-fit"&gt;The mechanism that does fit&lt;/h2&gt;
&lt;p&gt;It is older and I had never considered it.
&lt;a href="https://pubmed.ncbi.nlm.nih.gov/17173887/" class="external-link" rel="noopener"&gt;Corley, MacGregor and Donaldson&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;
(Cognition, 2007) recorded ERPs during spoken sentences and found the N400
effect for unpredictable words was reduced when the target was preceded by a
hesitation marked by &amp;ldquo;er&amp;rdquo;. In a later recognition memory test, words preceded by
disfluency were more likely to be remembered. Disfluency raises attention and
improves encoding.&lt;/p&gt;
&lt;p&gt;That is a claim about the moment of hesitation, not about the speaker&amp;rsquo;s
perceived species, and it is the only mechanism I have found that joins the edit
I made to the number that moved. Hook rate asks whether someone is still there
at three seconds. A cue documented to raise attention, sitting at second one, is
a better account of that than anything about humanness.&lt;/p&gt;
&lt;p&gt;The ceilings are real and one of them is a citation ceiling. The experiment used
lab sentences rather than paid video, it measured comprehension and memory
rather than scroll behaviour, and the sample size is not reported in the PubMed
abstract I can link, with the full text behind Cognition&amp;rsquo;s paywall. A reader
following my citation cannot check how many people it ran on, so treat the size
of the effect as unverified from here.&lt;/p&gt;
&lt;h2 id="why-it-happened-and-why-sample-size-was-not-the-reason"&gt;Why it happened, and why sample size was not the reason&lt;/h2&gt;
&lt;p&gt;Three mechanisms stack, in order of size. First, the confounded variable: I
changed the measurement window and read the result as evidence about a global
property.&lt;/p&gt;
&lt;p&gt;Second, non-random impression assignment. Meta&amp;rsquo;s
&lt;a href="https://www.facebook.com/business/help/430291176997542" class="external-link" rel="noopener"&gt;ad auction page&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; lists
estimated action rates as an auction input, &amp;ldquo;an estimate of whether a particular
person engages with or converts from a particular ad.&amp;rdquo; If both cuts sit in one
ad set, delivery routes each toward the people it predicts will engage with that
specific ad, so any hook rate gap is partly the system&amp;rsquo;s own prediction of hook
rate returned as impression allocation. More spend does not fix that. It feeds
it. Meta&amp;rsquo;s
&lt;a href="https://www.facebook.com/business/help/1738164643098669" class="external-link" rel="noopener"&gt;A/B testing page&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; says
as much: &amp;ldquo;We do not recommend testing informally, such as by turning ad sets or
campaigns on and off manually. This can lead to inefficient ad delivery and
unreliable test results.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Third, learning phase instability. Ad sets &amp;ldquo;exit the learning phase as soon as
they can deliver stably. This usually occurs after about 50 results in the week
after the ad set&amp;rsquo;s last significant edit&amp;rdquo;, and during it &amp;ldquo;performance is less
stable, so your results aren&amp;rsquo;t necessarily indicative of future performance&amp;rdquo;
(&lt;a href="https://www.facebook.com/business/help/112167992830700" class="external-link" rel="noopener"&gt;learning phase&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;).&lt;/p&gt;
&lt;p&gt;I originally blamed sample size, and at impression level that is almost
certainly wrong. For a two-proportion test at 80% power and alpha 0.05
two-sided, normal approximation:&lt;/p&gt;
&lt;p&gt;&lt;code&gt;n = [1.95996*sqrt(2*p̄*q̄) + 0.84162*sqrt(p1q1 + p2q2)]^2 / (p1 - p2)^2&lt;/code&gt;&lt;/p&gt;
&lt;p&gt;At a 25% baseline and a 20% relative lift that is
&lt;code&gt;[1.23779 + 0.53061]^2 / 0.0025 = 1,251&lt;/code&gt; impressions per arm.&lt;/p&gt;
&lt;p&gt;Table: Impressions per arm needed to detect a hook rate difference at 80% power, alpha 0.05 two-sided, normal approximation. My arithmetic.&lt;/p&gt;
&lt;table&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th&gt;Baseline hook rate&lt;/th&gt;
 &lt;th style="text-align: right"&gt;+5% relative&lt;/th&gt;
 &lt;th style="text-align: right"&gt;+10% relative&lt;/th&gt;
 &lt;th style="text-align: right"&gt;+20% relative&lt;/th&gt;
 &lt;th style="text-align: right"&gt;+30% relative&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;20%&lt;/td&gt;
 &lt;td style="text-align: right"&gt;25,582&lt;/td&gt;
 &lt;td style="text-align: right"&gt;6,509&lt;/td&gt;
 &lt;td style="text-align: right"&gt;1,682&lt;/td&gt;
 &lt;td style="text-align: right"&gt;771&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;25%&lt;/td&gt;
 &lt;td style="text-align: right"&gt;19,146&lt;/td&gt;
 &lt;td style="text-align: right"&gt;4,861&lt;/td&gt;
 &lt;td style="text-align: right"&gt;1,251&lt;/td&gt;
 &lt;td style="text-align: right"&gt;570&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;30%&lt;/td&gt;
 &lt;td style="text-align: right"&gt;14,855&lt;/td&gt;
 &lt;td style="text-align: right"&gt;3,762&lt;/td&gt;
 &lt;td style="text-align: right"&gt;963&lt;/td&gt;
 &lt;td style="text-align: right"&gt;437&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;Any live Meta test clears the +20% and +30% columns in hours. Treat those
figures as a floor, since impressions are not independent Bernoulli trials and
clustering by person and session inflates the true variance. The conclusion
holds anyway: my test was almost certainly powered well enough and still could
not answer the question. The defect was assignment and confounding, not n.&lt;/p&gt;
&lt;h2 id="how-to-detect-this-in-your-own-account"&gt;How to detect this in your own account&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;Move the imperfection out of the first three seconds. Put the stumble at
second eight and re-run. If polish itself is the driver, hook rate should not
move. If it collapses back, the effect was positional.&lt;/li&gt;
&lt;li&gt;Compare impressions delivered to each cut. Materially unequal delivery inside
one ad set is the optimiser&amp;rsquo;s fingerprint on your result.&lt;/li&gt;
&lt;li&gt;Check frequency, because the denominator is impressions.&lt;/li&gt;
&lt;li&gt;Confirm both ad sets left the learning phase, roughly 50 results in a week,
before reading anything.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;To design so it cannot recur: one variable, isolated outside the metric window;
separate ad sets through Meta&amp;rsquo;s A/B test tool so audiences are split rather than
competing; seven days minimum and thirty maximum; budget for both arms to exit
learning; hypothesis and metric written down before launch. Meta&amp;rsquo;s page on
&lt;a href="https://www.facebook.com/business/help/166313650471318" class="external-link" rel="noopener"&gt;how winning campaigns are determined&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;
lists three reasons its own declared winner might still have turned out more
expensive than the alternative: &amp;ldquo;The length of the study was too short.&amp;rdquo;
&amp;ldquo;There weren&amp;rsquo;t enough results to calculate an accurate winner.&amp;rdquo; &amp;ldquo;Best practices
for study length weren&amp;rsquo;t followed or met.&amp;rdquo; Two of the three are duration.&lt;/p&gt;
&lt;h2 id="the-friction-question-is-becoming-a-compliance-question"&gt;The friction question is becoming a compliance question&lt;/h2&gt;
&lt;p&gt;Meta&amp;rsquo;s transparency post, published 3 February 2025 and updated 1 June 2026,
describes two labelling regimes, and they are not the same size. For Meta&amp;rsquo;s own
tools: &amp;ldquo;When an image or video is created or significantly edited with our
generative AI creative features in our advertiser marketing tools, a label will
appear in the three-dot menu or next to the &amp;lsquo;Sponsored&amp;rsquo; label&amp;rdquo;, and when those
in-house tools &amp;ldquo;result in the inclusion of an AI-generated photorealistic human,
the label will appear next to the Sponsored label (not behind the three-dot
menu).&amp;rdquo; For everything else, Meta &amp;ldquo;will also begin automatically detecting ads
created or edited using third-party AI tools through industry-standard signals.
When detected, we&amp;rsquo;ll apply an &amp;lsquo;AI info&amp;rsquo; label included in About this ad&amp;rdquo;, and
About this ad is the three-dot menu.&lt;/p&gt;
&lt;p&gt;So the loud label is documented for Meta&amp;rsquo;s in-house generation, and an AI UGC ad
built in a third-party tool currently earns the quiet one. I would not plan
around that gap. Meta does not say what the industry-standard signals are, and
my reading that they are provenance metadata of the C2PA kind is an inference,
not something the post states. What that class of metadata does is documented.
&lt;a href="https://spec.c2pa.org/specifications/specifications/2.4/specs/C2PA_Specification.html#_soft_bindings" class="external-link" rel="noopener"&gt;C2PA 2.4&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;
defines soft bindings, computed from the content rather than the raw bits, and
says they &amp;ldquo;enable digital content to be matched even if the underlying bits
differ&amp;rdquo;, giving as the example &amp;ldquo;an asset rendition in a different resolution or
encoding format.&amp;rdquo; A re-encode does not shake it off. Detection is moving from
perception to metadata, where the viewer&amp;rsquo;s eye does not participate.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-50" class="external-link" rel="noopener"&gt;Article 50&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; of
Regulation (EU) 2024/1689 (&lt;a href="https://artificialintelligenceact.eu/article/50/" class="external-link" rel="noopener"&gt;full text&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;)
became applicable on 2 August 2026, four days before I wrote this. Providers of
systems generating synthetic audio, image, video or text &amp;ldquo;shall ensure that the
outputs of the AI system are marked in a machine-readable format and detectable
as artificially generated or manipulated&amp;rdquo;, and deployers generating a deep fake
&amp;ldquo;shall disclose that the content has been artificially generated or
manipulated&amp;rdquo;. I am not a lawyer, read
the text yourself, but one consequence of the definitions is worth flagging. The
Commission&amp;rsquo;s
&lt;a href="https://digital-strategy.ec.europa.eu/en/faqs/transparency-obligations-under-article-50-ai-act" class="external-link" rel="noopener"&gt;FAQ&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;
gives three cumulative criteria for a deep fake under Article 3(60). The second
one decides this, and it is easy to read only half of. In full: &amp;ldquo;simulated
persons, objects, places, entities or events need to resemble someone or
something that exists, can plausibly exist or could have plausibly existed in
reality.&amp;rdquo; Stop at &amp;ldquo;exists&amp;rdquo; and a fully synthetic presenter resembling no real
person looks like it falls outside. The clause that matters is &amp;ldquo;can plausibly
exist&amp;rdquo;, and a photorealistic synthetic person built to pass as a real one is
close to the paradigm case of it. On the Commission&amp;rsquo;s own wording I read the
deployer disclosure duty as reaching an AI UGC ad, with the provider&amp;rsquo;s marking
duty binding the generation tool on top of that. You inherit a marked file you
did not choose to mark, and you probably owe the disclosure as well. The
Commission has assessed a
&lt;a href="https://digital-strategy.ec.europa.eu/en/policies/code-practice-ai-generated-content" class="external-link" rel="noopener"&gt;Code of Practice on transparency of AI-generated content&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;
as adequate, so the mechanics are being standardised now rather than eventually.&lt;/p&gt;
&lt;p&gt;Which makes the size of the label effect the question worth asking, and the two
studies that measure it are routinely misreported.
&lt;a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC11443540/" class="external-link" rel="noopener"&gt;Altay and Gilardi&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;
(PNAS Nexus, October 2024, two preregistered experiments, N = 4,976) found that
labelling headlines as AI-generated &amp;ldquo;reduced sharing and accuracy ratings by.17
points [-0.29, -0.04], P = 0.010 on the 6-point scale&amp;rdquo;, replicated at 0.11
points in Study 2, with sharing intentions alone reaching significance in
neither (P = 0.27 and P = 0.085). And the line the citations leave out: &amp;ldquo;the
effect of labeling headlines as AI-generated (2.66pp) was three times smaller
than the effect of labeling headlines as false (9.33pp).&amp;rdquo; The direction is
negative and the magnitude is small.
&lt;a href="https://dare.uva.nl/personal/pure/en/publications/disclaimer-this-content-is-aigenerated-how-aidisclosures-influence-trust-in-advertisements-and-organizations%28dc6a541b-3dac-4913-9f07-a6573b5764f0%29.html" class="external-link" rel="noopener"&gt;Koning and Voorveld&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;
(Journal of Interactive Advertising, September 2025, N = 304) found an AI
disclosure raised persuasion knowledge, which reduced trust in the ad and the
organisation, and also reports a countervailing path through attitudinal
persuasion knowledge that increased trust. Citing only the negative half
misrepresents it.&lt;/p&gt;
&lt;h2 id="what-i-have-now"&gt;What I have now&lt;/h2&gt;
&lt;p&gt;A test design that separates the variable from the metric window, and a
corrected model of my own result. The win did not come from being harder to
detect, because polish is not what people detect on. It did not come from
feeding the humanness heuristic either, because the measured version of that
heuristic reads roughness as a machine and would have deleted my stumble. What
is left is attention: a disfluency effect on encoding that has been in the
literature since 2007, landing by accident inside the three seconds the metric
samples.&lt;/p&gt;
&lt;p&gt;That changes the next build instead of confirming it. If the driver is
disfluency, the lever is not &amp;ldquo;leave humanity in&amp;rdquo;, it is placement, and a
hesitation belongs immediately before whatever the viewer most needs to
remember, not at second one where I happened to put it. Step one of the recipe
above tests that directly, and it can come back against me.&lt;/p&gt;
&lt;p&gt;My original conclusion was that the advantage lies in knowing how much humanity
to leave in. The sentence is wrong about the noun, and it is wrong in a more
useful way than I expected. Where people do carry a shortcut for humanity, it is
mapped, unreliable, and it rewards smoothness rather than friction. The friction
I left in bought me something else. Calling that attention rather than humanity
is the difference between a lever I can aim and a story I liked.&lt;/p&gt;</content:encoded></item><item><title>The scaled content abuse policy does not ban AI writing</title><link>https://therezaali.com/writing/google-scaled-content-abuse/</link><guid isPermaLink="true">https://therezaali.com/writing/google-scaled-content-abuse/</guid><pubDate>Wed, 05 Aug 2026 00:00:00 +0200</pubDate><author>contact@therezaali.com (Mo RezaAli)</author><category>reference</category><description>It bans generating many pages to manipulate rankings, “no matter how it’s created”. The policy in full, then the strongest argument that the distinction cannot be applied.</description><content:encoded>&lt;blockquote&gt;
&lt;p&gt;Scaled content abuse is when many pages are generated for the primary purpose of
manipulating search rankings and not helping users. This abusive practice is typically
focused on creating large amounts of unoriginal content that provides little to no value
to users, no matter how it&amp;rsquo;s created.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;a href="https://developers.google.com/search/docs/essentials/spam-policies" class="external-link" rel="noopener"&gt;Spam policies for Google web search&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;, Google Search Central&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;Two sentences. That is the whole definition, it is free to read, and it does not contain the
word &amp;ldquo;AI&amp;rdquo;. That is worth sitting with, given how much of the commentary about it is about
nothing else.&lt;/p&gt;
&lt;p&gt;Read the last clause again. &lt;em&gt;No matter how it&amp;rsquo;s created.&lt;/em&gt; Method is not the test.&lt;/p&gt;
&lt;p&gt;That is my claim, and I think it is right. It is also the claim that draws the sharpest
objection I have heard on the subject, which is that the distinction is a lawyer&amp;rsquo;s
distinction with no cash value. This piece states the text, then states that objection as
strongly as I can put it, then says where I think it fails.&lt;/p&gt;
&lt;p&gt;I am quoting the version of the spam policies page that carried &amp;ldquo;Last updated 2026-05-15 UTC&amp;rdquo;
in its footer on the day I read it. The same page lists five examples, and only the first
mentions generative tools:&lt;/p&gt;
&lt;blockquote&gt;
&lt;ul&gt;
&lt;li&gt;Using generative AI tools or other similar tools to generate many pages without adding
value for users&lt;/li&gt;
&lt;li&gt;Scraping feeds, search results, or other content to generate many pages (including
through automated transformations like synonymizing, translating, or other obfuscation
techniques), where little value is provided to users&lt;/li&gt;
&lt;li&gt;Stitching or combining content from different web pages without adding value&lt;/li&gt;
&lt;li&gt;Creating multiple sites with the intent of hiding the scaled nature of the content&lt;/li&gt;
&lt;li&gt;Creating many pages where the content makes little or no sense to a reader but contains
search keywords&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;em&gt;&lt;a href="https://developers.google.com/search/docs/essentials/spam-policies" class="external-link" rel="noopener"&gt;Spam policies for Google web search&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;, Google Search Central&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;Every one of the five turns on the same two conditions: &lt;em&gt;many pages&lt;/em&gt;, and &lt;em&gt;no added value&lt;/em&gt;.
Neither condition mentions a model. A policy that bans writing with a language model and a
policy that bans flooding an index are two different policies, and Google published the
second one.&lt;/p&gt;
&lt;h2 id="what-changed-in-march-2024-and-what-did-not"&gt;What changed in March 2024, and what did not&lt;/h2&gt;
&lt;p&gt;Google announced the policy on 5 March 2024, alongside the March 2024 core update, in a post
by Chris Nelson. Three spam policies launched that day: expired domain abuse, scaled
content abuse, and site reputation abuse. The post is explicit that scaled content abuse was
a rewrite rather than a new prohibition:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;This new policy builds on our previous spam policy about automatically-generated content,
ensuring that we can take action on scaled content abuse as needed, no matter whether
content is produced through automation, human efforts, or some combination of human and
automated processes.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;a href="https://developers.google.com/search/blog/2024/03/core-update-spam-policies" class="external-link" rel="noopener"&gt;What web creators should know about our March 2024 core update and new spam policies&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;, Google Search Central Blog&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;The post anticipates the misreading and answers it directly:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Our long-standing spam policy has been that use of automation, including generative AI, is
spam if the primary purpose is manipulating ranking in Search results. The updated policy
is in the same spirit of our previous policy and based on the same principle. It&amp;rsquo;s been
expanded to account for more sophisticated scaled content creation methods where it isn&amp;rsquo;t
always clear whether low quality content was created purely through automation.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;a href="https://developers.google.com/search/blog/2024/03/core-update-spam-policies" class="external-link" rel="noopener"&gt;What web creators should know about our March 2024 core update and new spam policies&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;, Google Search Central Blog&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;So the widening in 2024 went in the opposite direction from the one people assumed. The old
rule named automation and could be dodged by putting a human in the loop. The new rule drops
the method entirely and keeps only the purpose and the scale, which means a room of underpaid
writers producing the same output is covered too.&lt;/p&gt;
&lt;p&gt;Google&amp;rsquo;s position on AI writing predates all of this by a year, and sits in the FAQ at the
foot of the February 2023 post by Danny Sullivan and Chris Nelson rather than in its body.
Worth saying, because the difference matters if you go looking. On &amp;ldquo;Is AI content against
Google Search&amp;rsquo;s guidelines?&amp;rdquo;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Appropriate use of AI or automation is not against our guidelines. This means that it is
not used to generate content primarily to manipulate search rankings, which is against our
spam policies.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;a href="https://developers.google.com/search/blog/2023/02/google-search-and-ai-content" class="external-link" rel="noopener"&gt;Google Search&amp;rsquo;s guidance about AI-generated content&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;, Google Search Central Blog&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;And on whether it will rank: &amp;ldquo;Using AI doesn&amp;rsquo;t give content any special gains. It&amp;rsquo;s just
content. If it is useful, helpful, original, and satisfies aspects of E-E-A-T, it might do
well in Search. If it doesn&amp;rsquo;t, it might not.&amp;rdquo;&lt;/p&gt;
&lt;h2 id="the-objection-at-full-strength"&gt;The objection, at full strength&lt;/h2&gt;
&lt;p&gt;Here is the case against everything above, and it is better than it usually gets credit for.&lt;/p&gt;
&lt;p&gt;Google cannot observe purpose. Purpose is a fact about a publisher&amp;rsquo;s head, and no crawler
reaches it. What Google can observe is the artefact (the page, the site, the rate at which
pages appeared), so &amp;ldquo;primary purpose&amp;rdquo; has to be inferred from surface features. And the
surface features that read as &lt;em&gt;generated at scale&lt;/em&gt; are, in 2026, overwhelmingly the features
of machine-written text: uniform section structure, uniform length, hedged summary prose,
citations that name a publisher and no document.&lt;/p&gt;
&lt;p&gt;Then close the loop. Google&amp;rsquo;s November 2024 site reputation abuse post says the company does
not &amp;ldquo;simply take a site&amp;rsquo;s claims about how the content was produced at face value.&amp;rdquo; So the
publisher cannot rebut the inference by explaining the process, and a disclosure page will
not do it either.&lt;/p&gt;
&lt;p&gt;Put those together and the objection lands: a rule that says &lt;em&gt;no matter how it&amp;rsquo;s created&lt;/em&gt;,
enforced by inference from features that machine writing produces, is a rule about machine
writing in everything but wording. &amp;ldquo;Method is not the test&amp;rdquo; is then true on paper and
worthless in the hand: a distinction you cannot cash, drawn by the only party who gets to
score the exam.&lt;/p&gt;
&lt;p&gt;I do not think that argument is silly. I think it is half right, and the half that is wrong
is the half that matters.&lt;/p&gt;
&lt;h2 id="where-it-fails"&gt;Where it fails&lt;/h2&gt;
&lt;p&gt;It fails on the same document it appeals to. If the policy were secretly about detecting
machine authorship, the rater instructions would tell raters to look for machine authorship.
They tell them the opposite.&lt;/p&gt;
&lt;p&gt;Section 4.6.5 of the quality rater guidelines, in the version dated 11 September 2025:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Pages and websites made up of content created at scale with no original content or added
value for users, should be rated Lowest, no matter how they are created. Even if you are
unsure of the method of creation, e.g. whether or not the page is created using generative
AI tools, you should still use the Lowest rating when you strongly suspect scaled content
abuse after looking at several pages on the website.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;General Guidelines, 11 September 2025, section 4.6.5&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;Uncertainty about the method does not protect the page, and it is not supposed to resolve
into a guess about the method either. The rater is told to judge the output and to stop
asking.&lt;/p&gt;
&lt;p&gt;Two neighbouring instructions show the same standard cutting against human work. Section
4.6.6 applies the Lowest rating to copied or paraphrased main content &amp;ldquo;even if the page
assigns credit for the content to another source&amp;rdquo;, which catches hand-built aggregation
sites with impeccable attribution and no model anywhere near them. And section 3.2 defines
the yardstick as labour: &amp;ldquo;Effort: Consider the extent to which a human being actively worked
to create satisfying content,&amp;rdquo; with the counter-example given as &amp;ldquo;the automatic creation of
thousands of pages by running existing freely available content through existing translation
software without any oversight, manual curation, etc.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;The operative words in that counter-example are &lt;em&gt;oversight&lt;/em&gt; and &lt;em&gt;manual curation&lt;/em&gt;. Both are
human acts, both are performable on machine output, and neither is available at ten thousand
pages a week. That is the axis the policy actually runs along, and it is not the same axis as
authorship. The two come apart in both directions: a content farm staffed entirely by people
fails, and a page a model drafted, a person checked against named sources, and a person put
their name on, passes.&lt;/p&gt;
&lt;p&gt;So the objection is right that intent is read off the artefact. It is right that you cannot
talk your way out. What it gets wrong is the conclusion, because the artefact is being read
for evidence of work, not for evidence of a model. Evidence of work is something you can
put on the page on purpose. Google&amp;rsquo;s helpful-content page even tells you one of the things it
wants to see there: &amp;ldquo;Is the use of automation, including AI-generation, self-evident to
visitors through disclosures or in other ways?&amp;rdquo; A disclosure will not launder a thin page.
It is still the answer to a question the policy asks out loud.&lt;/p&gt;
&lt;h2 id="what-the-policy-prohibits-in-plain-terms"&gt;What the policy prohibits, in plain terms&lt;/h2&gt;
&lt;p&gt;Three tests, all of which have to fail before you are in trouble:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Scale.&lt;/strong&gt; One page is not scaled content abuse. The word &amp;ldquo;many&amp;rdquo; is in every version of
the definition.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Purpose.&lt;/strong&gt; The pages exist mainly to rank, not mainly to answer someone. The
helpful-content page frames this as the &amp;ldquo;Why&amp;rdquo; question, and it is the only one of the
three that is about intent rather than artefact.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Absence of added value.&lt;/strong&gt; Section 3.2&amp;rsquo;s effort standard, above, is what the output gets
measured against.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Nothing in there says do not use a model. It says do not ship volume you have not supervised.&lt;/p&gt;
&lt;p&gt;One clarification before the end, because the third policy from March 2024 keeps getting
folded into this one. Site reputation abuse was rewritten on 19 November 2024 and now reads:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Site reputation abuse is a tactic where third-party content is published on a host site
mainly because of that host&amp;rsquo;s already-established ranking signals, which it has earned
primarily from its first-party content. The goal of this tactic is for the content to rank
better than it could otherwise on its own.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;a href="https://developers.google.com/search/docs/essentials/spam-policies" class="external-link" rel="noopener"&gt;Spam policies for Google web search&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;, Google Search Central&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;The November rewrite removed the escape hatch that first-party oversight used to provide;
Google&amp;rsquo;s stated reasoning was that reviewing real cases showed &amp;ldquo;no amount of first-party
involvement alters the fundamental third-party nature of the content.&amp;rdquo; Different policy,
different failure, and a coupons subdirectory is not a scaled content problem.&lt;/p&gt;
&lt;h2 id="the-practical-reading"&gt;The practical reading&lt;/h2&gt;
&lt;p&gt;Draft a page with a model, check it against sources you can name, put your name on the
result, and this policy is not about you. Point the same model at a directory and fill it
with eight hundred pages nobody asked for, and it was written about you specifically. Since
March 2024 it has not mattered whether a person pressed the button.&lt;/p&gt;
&lt;p&gt;What I cannot tell you is how any of this is enforced against a particular site. I have no
manual action notice to show you, no message from Google, and no interest in inferring a
verdict from a traffic chart. The documentation on generative AI content is careful about the
same boundary from the other side: the rater guidelines &amp;ldquo;are not a guide to ranking first in
Google,&amp;rdquo; and raters&amp;rsquo; &amp;ldquo;ratings don&amp;rsquo;t directly influence ranking.&amp;rdquo; What the guidelines give you
is the standard, written down, by the people who wrote it. Read as a description of what a
sceptical stranger is told to look for, it is unusually clear. Read as a rule about robots,
it is unreadable, which is why so much of the anxiety about it lands on people it was never
drafted for.&lt;/p&gt;</content:encoded></item><item><title>Structured data that earns a manual action</title><link>https://therezaali.com/writing/structured-data-manual-actions/</link><guid isPermaLink="true">https://therezaali.com/writing/structured-data-manual-actions/</guid><pubDate>Wed, 05 Aug 2026 00:00:00 +0200</pubDate><author>contact@therezaali.com (Mo RezaAli)</author><category>reference</category><description>Apple publishes 4.9 and 4.89705 for the same app within one hour. Six conditions markup has to meet to stay inside the guidelines, and four public directories read on one day.</description><content:encoded>&lt;p&gt;Two of Apple&amp;rsquo;s own numbers about one app, read within the same hour on 5 August 2026:&lt;/p&gt;
&lt;pre tabindex="0"&gt;&lt;code&gt;apps.apple.com ratingValue 4.9 reviewCount 1770266
itunes.apple.com/lookup averageUserRating 4.89705 userRatingCount 1769848
&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;418 ratings apart. Both are Apple&amp;rsquo;s. Both describe the Eventbrite app. Neither is wrong. That
is what a live counter looks like when you photograph it twice.&lt;/p&gt;
&lt;p&gt;I start there because it is the cheap end of a question that gets expensive quickly: where did
the number in your markup come from? &lt;code&gt;AggregateRating&lt;/code&gt; is a machine-readable statement, made
by you, in your own markup, that many people rated a thing and this is their average. When
that statement is false there is nothing left to reinterpret afterwards. No tone, no context,
no &amp;ldquo;that isn&amp;rsquo;t what I meant.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Google&amp;rsquo;s rules for the type are mostly about provenance rather than accuracy, and they are
stricter than most people implementing them realise. Six conditions have to hold at once.
Below: the two documented rules the six come from, the list itself, and then the list run
against four public directories that all publish ratings about the same product.&lt;/p&gt;
&lt;h2 id="what-a-structured-data-manual-action-is-and-is-not"&gt;What a structured data manual action is, and is not&lt;/h2&gt;
&lt;p&gt;Google&amp;rsquo;s general structured data guidelines define the penalty precisely:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;If your page contains a structured data issue, it can result in a manual action. A
structured data manual action means that a page loses eligibility for appearance as a rich
result; it doesn&amp;rsquo;t affect how the page ranks in Google web search.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;a href="https://developers.google.com/search/docs/appearance/structured-data/sd-policies" class="external-link" rel="noopener"&gt;General structured data guidelines&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;, Google Search Central&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;So it is narrower than people fear and worse than they assume. Your rankings are untouched.
Your stars are gone, and a human at Google has looked at your page and concluded you were
being manipulative. The Search Console documentation describes what triggers it:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Google has detected that some of the markup on your pages may be using techniques that are
outside our structured data guidelines, for example: marking up content that is invisible
to users, marking up irrelevant or misleading content, or other manipulative behavior.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;a href="https://support.google.com/webmasters/answer/9044175" class="external-link" rel="noopener"&gt;Manual Actions report&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;, Google Search Console Help&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;Three triggers: invisible, irrelevant or misleading, manipulative. The first is a technical
mismatch. The second and third are accusations of bad faith.&lt;/p&gt;
&lt;h2 id="rule-one-markup-describes-what-the-reader-can-see"&gt;Rule one: markup describes what the reader can see&lt;/h2&gt;
&lt;p&gt;The quality guidelines are blunt about it:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Don&amp;rsquo;t mark up content that is not visible to readers of the page. For example, if the
JSON-LD markup describes a performer, the HTML body must describe that same performer.&lt;/p&gt;
&lt;p&gt;Don&amp;rsquo;t mark up irrelevant or misleading content, such as fake reviews or content unrelated
to the focus of a page.&lt;/p&gt;
&lt;p&gt;Don&amp;rsquo;t use structured data to deceive or mislead users. Don&amp;rsquo;t impersonate any person or
organization, or misrepresent your ownership, affiliation, or primary purpose.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;a href="https://developers.google.com/search/docs/appearance/structured-data/sd-policies" class="external-link" rel="noopener"&gt;General structured data guidelines&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;, Google Search Central&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;The same page states the relevance test as one sentence: &amp;ldquo;Your structured data must be a
true representation of the page content&amp;rdquo;. It gives two examples of failure, a live
streaming site labelling broadcasts as local events, and a woodworking site labelling
instructions as recipes.&lt;/p&gt;
&lt;p&gt;The review snippet documentation repeats the visibility rule specifically for ratings:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;If you use AggregateRating, users should be able to see that aggregate rating on the page.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;a href="https://developers.google.com/search/docs/appearance/structured-data/review-snippet" class="external-link" rel="noopener"&gt;Review snippet structured data&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;, Google Search Central&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;A number that exists only in the JSON-LD is already a violation before anyone asks whether it
is accurate.&lt;/p&gt;
&lt;h2 id="rule-two-ratings-come-from-users-not-from-you"&gt;Rule two: ratings come from users, not from you&lt;/h2&gt;
&lt;p&gt;This is the rule most often broken by businesses with no intent to deceive, because it feels
reasonable to publish your own testimonials. Google closed it in September 2019 and explained
why in the announcement:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Reviews that can be perceived as &amp;ldquo;self-serving&amp;rdquo; aren&amp;rsquo;t in the best interest of users. We
call reviews &amp;ldquo;self-serving&amp;rdquo; when a review about entity A is placed on the website of entity
A - either directly in their markup or via an embedded third-party widget.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;a href="https://developers.google.com/search/blog/2019/09/making-review-rich-results-more-helpful" class="external-link" rel="noopener"&gt;Making Review Rich Results more helpful&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;, Google Search Central Blog&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;The current documentation states the consequence and the sourcing requirement:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;If the entity that&amp;rsquo;s being reviewed controls the reviews about itself, their pages that use
LocalBusiness or any other type of Organization structured data are ineligible for star
review feature. [&amp;hellip;]&lt;/p&gt;
&lt;p&gt;Ratings must be sourced directly from users.&lt;/p&gt;
&lt;p&gt;Don&amp;rsquo;t rely on human editors to create, curate, or compile ratings information for local
businesses.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;a href="https://developers.google.com/search/docs/appearance/structured-data/review-snippet" class="external-link" rel="noopener"&gt;Review snippet structured data&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;, Google Search Central&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;Two more constraints sit alongside it. &lt;code&gt;AggregateRating&lt;/code&gt; is defined by its documentation as
markup for &amp;ldquo;an aggregate evaluation of an item by many people,&amp;rdquo; which is a statement about
where the number came from and not just what shape it has. Schema.org&amp;rsquo;s own definition of the
type is &amp;ldquo;The average rating based on multiple ratings or reviews.&amp;rdquo; And &amp;ldquo;Don&amp;rsquo;t aggregate
reviews or ratings from other websites,&amp;rdquo; which rules out the obvious workaround of scraping
somebody else&amp;rsquo;s average and marking it up as your own.&lt;/p&gt;
&lt;p&gt;The enforcement note is filed where nobody reads it, inside a paragraph about completeness:
&amp;ldquo;users prefer recipes with actual user reviews and genuine star ratings (note that reviews or
ratings not by actual users may result in manual action).&amp;rdquo;&lt;/p&gt;
&lt;h2 id="the-six-conditions"&gt;The six conditions&lt;/h2&gt;
&lt;p&gt;Put the two rules together and an &lt;code&gt;AggregateRating&lt;/code&gt; you may legitimately publish has to
satisfy all of:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;It is about an item, not a category or a list.&lt;/li&gt;
&lt;li&gt;The ratings came from users.&lt;/li&gt;
&lt;li&gt;Those users were not you.&lt;/li&gt;
&lt;li&gt;You did not copy the average from another site.&lt;/li&gt;
&lt;li&gt;A human editor did not compile it.&lt;/li&gt;
&lt;li&gt;The reader can see the same number on the page.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Conditions 2 to 5 are all provenance. Only 6 is about the page. Which means most of the work
of getting this right happens before any markup is written, in the boring question of who
supplied the numbers and how you know.&lt;/p&gt;
&lt;h2 id="four-directories-one-product-one-day"&gt;Four directories, one product, one day&lt;/h2&gt;
&lt;p&gt;Provenance is easier to see in the wild than in the abstract, so: Eventbrite, read across four
public sources on 5 August 2026.&lt;/p&gt;
&lt;p&gt;On &lt;strong&gt;the App Store&lt;/strong&gt;, Apple&amp;rsquo;s Eventbrite page carries its own aggregate rating markup: a
&lt;code&gt;ratingValue&lt;/code&gt; of &lt;strong&gt;4.9&lt;/strong&gt; from a &lt;code&gt;reviewCount&lt;/code&gt; of &lt;strong&gt;1,770,266&lt;/strong&gt;, which the page renders for the
reader as &amp;ldquo;1.8M Ratings.&amp;rdquo; Apple&amp;rsquo;s public iTunes Lookup API, queried a few minutes later,
answered &lt;code&gt;averageUserRating&lt;/code&gt; 4.89705 from &lt;code&gt;userRatingCount&lt;/code&gt; 1769848. That is the 418-rating
gap at the top of this page. The long decimals here are perishable in a way the policy
quotations above are not; query the lookup endpoint yourself and you will get a third number.
The rounded value, 4.9, is the part that holds still.&lt;/p&gt;
&lt;p&gt;On &lt;strong&gt;Google Play&lt;/strong&gt;, the store listing&amp;rsquo;s own markup reported &lt;strong&gt;4.8&lt;/strong&gt; (4.81129264831543) from
&lt;strong&gt;273,901&lt;/strong&gt; ratings, read the same day.&lt;/p&gt;
&lt;p&gt;Two software directories put the organiser-facing product lower, and much closer to each
other: Capterra&amp;rsquo;s Eventbrite reviews page and GetApp&amp;rsquo;s both showed &lt;strong&gt;4.6 from 5,804 reviews&lt;/strong&gt;,
read in a browser. Those two are flagged here rather than listed as machine-checked sources,
for two reasons. Both answer a plain command-line request with HTTP 403 (I tried again while
revising this paragraph and got 403 from each), so the build cannot re-verify them on deploy
the way it re-verifies everything else, and that pair of figures is the one claim on this page
you have to take on my word. They are also both Gartner Digital Markets properties drawing on
a shared review pool, which makes the matching number one dataset rather than two independent
confirmations.&lt;/p&gt;
&lt;p&gt;Across four directories the range is 4.6 to 4.9, and the spread is not the interesting part.
The App Store and Play figures describe a consumer app used by attendees. The directory
figures describe software bought by event organisers. Different populations, answering
different questions, and no single figure among them is entitled to be published as &lt;em&gt;the&lt;/em&gt;
rating for the product.&lt;/p&gt;
&lt;p&gt;Now notice what Apple and Google are doing on those pages. Both publish &lt;code&gt;AggregateRating&lt;/code&gt;
markup about a product they do not own. Both are within policy, because each is aggregating
ratings that real users left with them, and each shows the reader the same number it marks up.
The type is not forbidden. Condition 3 is about whether the ratings are about you, not about
whether the subject is somebody else. Provenance is what makes the assertion true or
false.&lt;/p&gt;
&lt;h2 id="the-rules-short"&gt;The rules, short&lt;/h2&gt;
&lt;ul&gt;
&lt;li&gt;Mark up what is on the page. If the reader cannot see it, do not assert it.&lt;/li&gt;
&lt;li&gt;Do not rate yourself. Self-serving reviews on &lt;code&gt;LocalBusiness&lt;/code&gt; or &lt;code&gt;Organization&lt;/code&gt; make the
page ineligible for the star feature.&lt;/li&gt;
&lt;li&gt;Do not rate other people&amp;rsquo;s products unless real users handed you the ratings. Ask who those
people were. If the answer is a spreadsheet, stop.&lt;/li&gt;
&lt;li&gt;Do not copy someone else&amp;rsquo;s average into your markup.&lt;/li&gt;
&lt;li&gt;If a number is a live counter, say so in prose, because the markup cannot.&lt;/li&gt;
&lt;li&gt;If you cannot show the reader how the number was produced, write prose instead of schema.
Prose invites an argument. &lt;code&gt;AggregateRating&lt;/code&gt; does not. It only asserts, and it asserts to
machines that have no way to ask you a follow-up question.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;This build emits no &lt;code&gt;Review&lt;/code&gt; or &lt;code&gt;AggregateRating&lt;/code&gt; markup about any third party, and the
&lt;a href="https://therezaali.com/editorial-policy/"&gt;editorial policy&lt;/a&gt; carries that as a standing rule rather than as a
one-time cleanup.&lt;/p&gt;</content:encoded></item><item><title>Seven things a comparison site publishes that you can check from outside</title><link>https://therezaali.com/writing/honest-comparison-sites/</link><guid isPermaLink="true">https://therezaali.com/writing/honest-comparison-sites/</guid><pubDate>Wed, 05 Aug 2026 00:00:00 +0200</pubDate><author>contact@therezaali.com (Mo RezaAli)</author><category>essay</category><description>You are not in the room and cannot audit what was tested. These come from the methodologies of Wirecutter, Consumer Reports, G2 and Tom’s Hardware, and I graded my own deleted benchmark against them.</description><content:encoded>&lt;p&gt;The reasonable objection to an article like this one goes: you cannot audit a
comparison site from outside. You are not in the room. You do not know what was
tested, who paid for what, or which vendor called which editor. Absent that, a
disclosure page and an affiliate notice are about all anyone can ask for, and a
reader is better off just judging the verdicts on their merits.&lt;/p&gt;
&lt;p&gt;The first half of that is true. The conclusion does not follow.&lt;/p&gt;
&lt;p&gt;You cannot audit the verdicts. What you can audit (from outside, without access,
in the time it takes to open four tabs) is the &lt;em&gt;process&lt;/em&gt;, and you can do it
wherever the publisher has committed to enough specifics that a mismatch between
what they said they do and what the pages show them doing would be visible to a
stranger. That is the whole game. A small number of organisations play it well
enough to steal a standard from.&lt;/p&gt;
&lt;p&gt;Four publish enough detail to reverse-engineer one: Wirecutter, Consumer
Reports, G2 and Tom&amp;rsquo;s Hardware. Below is the checklist I pulled out of their
documents. Each check comes with the way to run it on somebody else&amp;rsquo;s site, and
then with its result on a comparison benchmark of my own, which ranked 19 online
registration products and which I deleted in August 2026. Running them together
is deliberate. A checklist nobody has ever failed in public is a wish list.&lt;/p&gt;
&lt;p&gt;I am grading myself with someone else&amp;rsquo;s ruler. That is why the confidence label
on this page reads medium and not high: the checks are documented, the weighting
I put on them is my opinion, and I am the only source for the grades, because I
removed the property.&lt;/p&gt;
&lt;h2 id="1-a-published-methodology-that-the-sites-own-process-actually-follows"&gt;1. A published methodology that the site&amp;rsquo;s own process actually follows&lt;/h2&gt;
&lt;p&gt;The first check is not &amp;ldquo;is there a methodology page.&amp;rdquo; It is &amp;ldquo;does the
methodology page describe what the site does.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Wirecutter&amp;rsquo;s guide template is public. Every guide is supposed to contain
&amp;ldquo;discussions of why you should trust us, who the guide is for, how we picked and
tested the products, our picks, flaws (but not dealbreakers) that we
encountered, other options worth considering, and the models we tested but don&amp;rsquo;t
recommend&amp;rdquo;
(&lt;a href="https://www.nytimes.com/wirecutter/blog/anatomy-of-a-guide/" class="external-link" rel="noopener"&gt;The Anatomy of a Wirecutter Guide&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;).
Thirty seconds and three guides is enough to falsify that.&lt;/p&gt;
&lt;p&gt;The same page says why they bother, and the reasoning is the reason I keep
citing them: &amp;ldquo;The Wirecutter journalists who make picks and write guides are
dedicated experts. But it&amp;rsquo;s not enough for us to simply make that assertion and
expect readers to trust us. We have to explain why.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Run it:&lt;/strong&gt; read the methodology, open two or three actual comparisons, and try
to trace one specific verdict back through the stated process. If the trail goes
cold, the methodology is decoration.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Result on mine: failed.&lt;/strong&gt; The benchmark published a scoring methodology. The
scoring function did not implement it. Two documents, one public and one
executing, disagreeing with each other. Worse: a &lt;code&gt;Math.abs()&lt;/code&gt; call in the
comparison logic took the absolute value of a score difference, so every product
appeared to lead every rival it was set against. Every head-to-head page was a
win. A rounding error does not do that. That is a claim generator.&lt;/p&gt;
&lt;h2 id="2-inclusion-thresholds-stated-as-numbers-not-adjectives"&gt;2. Inclusion thresholds stated as numbers, not adjectives&lt;/h2&gt;
&lt;p&gt;&amp;ldquo;We only cover the leading products&amp;rdquo; commits its author to nothing at all.&lt;/p&gt;
&lt;p&gt;G2 publishes real thresholds. For a category to get a Grid report, &amp;ldquo;it must have
at least six products with 10+ reviews, and 150+ reviews overall,&amp;rdquo; and &amp;ldquo;To
qualify for inclusion, a product or service must have at least 10 reviews in the
corresponding category&amp;rdquo;
(&lt;a href="https://documentation.g2.com/docs/research-scoring-methodologies" class="external-link" rel="noopener"&gt;Research Scoring Methodologies&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;).
The same page states an update cadence you could catch them missing:
&amp;ldquo;Placement on live Grids® is updated daily.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Numbers like that are a commitment with a downside attached, which is what makes
them worth reading: they tell you which products were excluded and on what
ground, they leave the publisher nowhere to retreat to when a category on the
site plainly fails to clear the bar it set for itself, and they convert a claim
about editorial rigour into something a stranger with a browser can falsify in
an afternoon.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Run it:&lt;/strong&gt; find the stated minimum, then go hunting for a product on the site
that falls under it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Result on mine: failed.&lt;/strong&gt; The site claimed &amp;ldquo;10+ sources including G2 and
Capterra.&amp;rdquo; Its own source-coverage panel (which I built and shipped, on the same
page) showed those sources as skipped or errored. The site contradicted itself
in full view of the reader.&lt;/p&gt;
&lt;h2 id="3-named-testers-dated-tests"&gt;3. Named testers, dated tests&lt;/h2&gt;
&lt;p&gt;Consumer Reports names the size of the operation:
the Auto Test Center &amp;ldquo;requires a full-time staff of about 30 — engineers, editors, statisticians, technicians, photographers, videographers, and support staff&amp;rdquo;
(&lt;a href="https://www.consumerreports.org/cars/cars-driving/how-consumer-reports-tests-cars-auto-test-center-a3516544374/" class="external-link" rel="noopener"&gt;How Consumer Reports Tests Cars&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;).
Tom&amp;rsquo;s Hardware&amp;rsquo;s storage methodology carries a byline and a date (Chris
Ramseyer, published 14 March 2015), and the page still tells you the author &amp;ldquo;was
a senior contributing editor for Tom&amp;rsquo;s Hardware. He tested and reviewed consumer
storage&amp;rdquo;
(&lt;a href="https://www.tomshardware.com/reviews/how-we-test-storage,4058.html" class="external-link" rel="noopener"&gt;How We Test HDDs And SSDs&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;).&lt;/p&gt;
&lt;p&gt;That date cuts both ways and it is worth being fair about it. A methodology page
from 2015 is old. Tom&amp;rsquo;s Hardware does maintain separate methodology pages per
category:
&lt;a href="https://www.tomshardware.com/reviews/how-we-test-psu,4042.html" class="external-link" rel="noopener"&gt;power supplies&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;
and
&lt;a href="https://www.tomshardware.com/reviews/how-we-test-network-switches,4383.html" class="external-link" rel="noopener"&gt;network switches&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;
each have their own, so the 2015 date is not the whole picture. Even so, an old
date you can see beats a fresh-looking page with no date, because it lets you
ask the right question. An undated methodology cannot be stale, in the same way
an unlabelled jar cannot be past its expiry.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Run it:&lt;/strong&gt; look for a human name and a date on the methodology, and on the
individual comparisons. Then ask whether the date on a comparison squares with
the products inside it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Result on mine: failed.&lt;/strong&gt; It claimed hourly recomputation. The inputs that
recomputation ran against were months old. The timestamp was real and it was
measuring the wrong event: when the job ran, not when the data changed. Freshness
that cannot decay is not being measured; it is being displayed.&lt;/p&gt;
&lt;h2 id="4-conflicts-disclosed-at-the-point-of-the-conflict"&gt;4. Conflicts disclosed at the point of the conflict&lt;/h2&gt;
&lt;p&gt;A reader deciding between two products on a comparison page is not going to
detour through your About page first, so a conflict that lives only there
arrives after the decision it should have informed. The FTC&amp;rsquo;s guidance refuses to
let the burden slide somewhere convenient:
&amp;ldquo;the ultimate responsibility for clearly and conspicuously disclosing a material connection rests with the influencer and the brand – not the platform&amp;rdquo;
(&lt;a href="https://www.ftc.gov/business-guidance/resources/ftcs-endorsement-guides-what-people-are-asking" class="external-link" rel="noopener"&gt;FTC&amp;rsquo;s Endorsement Guides: What People Are Asking&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;;
the underlying rules are
&lt;a href="https://www.ecfr.gov/current/title-16/chapter-I/subchapter-B/part-255" class="external-link" rel="noopener"&gt;16 CFR Part 255&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;).&lt;/p&gt;
&lt;p&gt;G2 implements the idea at the row level. Reviewers &amp;ldquo;who have a business
relationship with a vendor (or their competitor) that could create bias (such as
a reseller) can share insight, but their reviews do not count toward scoring,&amp;rdquo;
and those reviews carry a flag
(&lt;a href="https://documentation.g2.com/help/docs/how-g2-ensures-authentic-reviews" class="external-link" rel="noopener"&gt;How G2 ensures authentic reviews&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;).
Its community guidelines add: &amp;ldquo;If a review is incentivized, G2 will clearly label
the review as incentivized&amp;rdquo;
(&lt;a href="https://legal.g2.com/community-guidelines" class="external-link" rel="noopener"&gt;Community Guidelines&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Run it:&lt;/strong&gt; find a product whose vendor has any commercial relationship with the
publisher, then see whether that relationship is named on the page where the
product is ranked.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Result on mine: failed, and this is the one that matters.&lt;/strong&gt; The benchmark
ranked Jumbula. I was Marketing Specialist at Jumbula from September 2019 to March 2026. That
relationship appeared nowhere in the benchmark property, while the same property
asserted that no vendor could pay to influence rankings. The second statement may
well have been true in the narrow sense that no money changed hands. It still
worked as misdirection, because the conflict that existed was not the conflict
being denied.&lt;/p&gt;
&lt;h2 id="5-test-units-bought-or-the-alternative-stated-plainly"&gt;5. Test units bought, or the alternative stated plainly&lt;/h2&gt;
&lt;p&gt;Consumer Reports: &amp;ldquo;we purchase every vehicle we test from a dealership, just like
you do. (Last year we spent more than $2.2 million buying cars.)&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Wirecutter&amp;rsquo;s arrangement is mixed and the page says so: it requests units or
buys them, spends &amp;ldquo;tens of thousands of dollars buying models for testing&amp;rdquo; every
month, returns or donates what it receives, and is blunt about unsolicited
gifts: &amp;ldquo;We simply do not accept these freebies&amp;rdquo;
(&lt;a href="https://www.nytimes.com/wirecutter/blog/yes-i-work-at-wirecutter-no-we-dont-get-a-bunch-of-free-stuff/" class="external-link" rel="noopener"&gt;Yes, I Work at Wirecutter. No, We Don&amp;rsquo;t Get a Bunch of Free Stuff.&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;).&lt;/p&gt;
&lt;p&gt;Those are two different policies, arrived at by two organisations with different
economics, and the honest part is not that either one is ideal but that both are
described precisely enough for a reader to weigh what each arrangement might do
to a verdict. Set either beside a sentence like &amp;ldquo;we maintain independence,&amp;rdquo;
which describes no arrangement, names no counterparty, forecloses no possibility
and could sit unchanged on the About page of a publisher taking money from every
vendor it covers. &amp;ldquo;We bought it and here is what it cost&amp;rdquo; can be wrong. That is
what makes it worth something.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Result on mine: not applicable, in a way that is itself the finding.&lt;/strong&gt; Nothing
was purchased because nothing was tested. A benchmark ranking 19 products without
running any of them is a scoring exercise over secondary data, and it should have
said exactly that in its first sentence.&lt;/p&gt;
&lt;h2 id="6-affiliate-revenue-disclosed-where-the-link-is"&gt;6. Affiliate revenue disclosed where the link is&lt;/h2&gt;
&lt;p&gt;Wirecutter puts it in the site&amp;rsquo;s standing header: &amp;ldquo;We independently review
everything we recommend. We may make money from the links on our site,&amp;rdquo; and
expands on the mechanism in
&lt;a href="https://www.nytimes.com/wirecutter/about/" class="external-link" rel="noopener"&gt;About Wirecutter&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;: &amp;ldquo;we may get paid
commissions on products purchased through our links to retailer sites.&amp;rdquo; Tom&amp;rsquo;s
Hardware carries the equivalent line on the review pages themselves: &amp;ldquo;When you
purchase through links on our site, we may earn an affiliate commission.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;The FTC guidance makes the same point from the other end: the disclosure belongs
in the content, not only beside the link.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Result on mine: passed.&lt;/strong&gt; There was no affiliate revenue, so there was nothing
to disclose. The only line that passes, and it passes for the least impressive
reason available.&lt;/p&gt;
&lt;h2 id="7-a-real-corrections-log"&gt;7. A real corrections log&lt;/h2&gt;
&lt;p&gt;Not a contact form. A dated, public, standing list of things the publisher got
wrong.
&lt;a href="https://www.nytimes.com/section/corrections" class="external-link" rel="noopener"&gt;The New York Times publishes one continuously&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;.
A comparison site carrying hundreds of verdicts and zero published corrections
is either newly launched or not looking.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Result on mine: failed.&lt;/strong&gt; There was none. Instead of a correction I deleted the
whole property, which is the crudest available remedy and destroys the evidence
along with the error. Had I kept a corrections log from the start, I would have
had somewhere to put the &lt;code&gt;Math.abs()&lt;/code&gt; bug on the day I found it, and you would be
able to check my account of it now instead of taking my word.&lt;/p&gt;
&lt;h2 id="what-the-pattern-was"&gt;What the pattern was&lt;/h2&gt;
&lt;p&gt;The checks that caught me were not the sophisticated ones. Nobody had to audit my
statistics.&lt;/p&gt;
&lt;p&gt;Three of the seven failures were visible on the page. A source panel that
contradicted the sentence above it. A freshness claim contradicted by its own
inputs. A comparison table in which nothing ever lost.&lt;/p&gt;
&lt;p&gt;Underneath all of it sits one ordinary sequence: I wrote the claims first,
because claims are the part you write when you are excited about a project and
they cost nothing; I built the system second, against a spec that had already
been published to the world as a description of something finished; and then I
never went back, not once in months, to run the boring comparison between the
two documents that would have taken an afternoon and caught every failure above.
Automation widens that gap and then hides it, because output keeps arriving and
output keeps looking like output, &lt;a href="https://therezaali.com/writing/the-bot-that-published-900-articles/"&gt;which is also how the publishing side of this
domain ran for six months&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Which reduces the checklist to a single question, asked of any claim you can see:
what would I expect to find on this site if that sentence were false? Then go
look for it. On my benchmark it took under a minute.&lt;/p&gt;</content:encoded></item><item><title>Perplexity’s own documentation says its fetcher ignores robots.txt</title><link>https://therezaali.com/writing/what-llms-read-on-your-site/</link><guid isPermaLink="true">https://therezaali.com/writing/what-llms-read-on-your-site/</guid><pubDate>Wed, 05 Aug 2026 00:00:00 +0200</pubDate><author>contact@therezaali.com (Mo RezaAli)</author><category>essay</category><description>The vendor’s words, not a critic’s. Where the line falls between training crawlers and answer-time fetchers, what each of the four big vendors documents, and the manifests I got wrong.</description><content:encoded>&lt;blockquote&gt;
&lt;p&gt;&amp;ldquo;this fetcher generally ignores robots.txt rules&amp;rdquo;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;That is Perplexity, describing &lt;code&gt;Perplexity-User&lt;/code&gt; on its own crawler documentation page. Not a
critic&amp;rsquo;s characterisation. The vendor&amp;rsquo;s.&lt;/p&gt;
&lt;p&gt;It draws the line most writing about AI crawlers smudges. Crawlers that gather training data
are governed by robots.txt and their operators say so. Fetchers that run because a human has
just typed a question and is waiting for an answer largely are not, by design, and their
operators say that too. Same documents, a few paragraphs apart. Carry that distinction
and most advice about &amp;ldquo;blocking AI&amp;rdquo; sorts itself into the part that works and the part that
cannot.&lt;/p&gt;
&lt;p&gt;Short piece. Three things are documented, one is not, one is proposed.&lt;/p&gt;
&lt;h2 id="documented-who-fetches-what"&gt;Documented: who fetches what&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;OpenAI&lt;/strong&gt; documents four agents on
&lt;a href="https://developers.openai.com/api/docs/bots" class="external-link" rel="noopener"&gt;Overview of OpenAI Crawlers&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;. &lt;code&gt;GPTBot&lt;/code&gt; is &amp;ldquo;used
to make our generative AI foundation models more useful and safe.&amp;rdquo; &lt;code&gt;OAI-SearchBot&lt;/code&gt; is &amp;ldquo;used to
surface websites in search results in ChatGPT&amp;rsquo;s search features.&amp;rdquo; Of the third, the page says
OpenAI &amp;ldquo;also uses ChatGPT-User for certain user actions in ChatGPT and Custom GPTs,&amp;rdquo; and that
it &amp;ldquo;is not used for crawling the web in an automatic fashion.&amp;rdquo; &lt;code&gt;OAI-AdsBot&lt;/code&gt; validates the
safety of pages submitted as ads. One operational figure: &amp;ldquo;it can take ~24 hours from a site&amp;rsquo;s
robots.txt update for our systems to adjust.&amp;rdquo; IP ranges are published as JSON at
&lt;code&gt;openai.com/gptbot.json&lt;/code&gt;, &lt;code&gt;openai.com/searchbot.json&lt;/code&gt; and &lt;code&gt;openai.com/chatgpt-user.json&lt;/code&gt;, all
three of which returned 200 when I fetched them on 5 August 2026.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Anthropic&lt;/strong&gt; documents three: &lt;code&gt;ClaudeBot&lt;/code&gt; for training data, &lt;code&gt;Claude-User&lt;/code&gt; for user-initiated
fetches, and &lt;code&gt;Claude-SearchBot&lt;/code&gt; for search indexing. The support page states that &amp;ldquo;Anthropic&amp;rsquo;s
Bots respect &amp;lsquo;do not crawl&amp;rsquo; signals by honoring industry standard directives in robots.txt,&amp;rdquo;
gives a &lt;code&gt;Crawl-delay: 1&lt;/code&gt; example, and links &lt;code&gt;claude.com/crawling/bots.json&lt;/code&gt; for source-IP
verification
(&lt;a href="https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler" class="external-link" rel="noopener"&gt;Does Anthropic crawl data from the web?&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Google-Extended&lt;/strong&gt; is the one people misread, and the misreading has a consequence. It is not
a crawler. Google&amp;rsquo;s own documentation: &amp;ldquo;Google-Extended doesn&amp;rsquo;t have a separate HTTP request
user agent string. Crawling is done with existing Google user agent strings; the robots.txt
user-agent token is used in a control capacity&amp;rdquo;
(&lt;a href="https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers" class="external-link" rel="noopener"&gt;List of Google&amp;rsquo;s common crawlers&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;,
last updated 14 July 2026). You will never find Google-Extended in your access logs, because
there is nothing there to find. It is a switch, not a visitor. Block it, then check your logs
for a change. You will see none, and conclude, wrongly, that it was ignored.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Perplexity&lt;/strong&gt; documents &lt;code&gt;PerplexityBot&lt;/code&gt;, which is &amp;ldquo;designed to surface and link websites in
search results on Perplexity. It is not used to crawl content for AI foundation models,&amp;rdquo;
alongside the &lt;code&gt;Perplexity-User&lt;/code&gt; line at the top of this page
(&lt;a href="https://docs.perplexity.ai/guides/bots" class="external-link" rel="noopener"&gt;Perplexity Crawlers&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;).&lt;/p&gt;
&lt;h2 id="documented-what-robotstxt-is-not"&gt;Documented: what robots.txt is not&lt;/h2&gt;
&lt;p&gt;The Robots Exclusion Protocol is a real standard:
&lt;a href="https://www.rfc-editor.org/rfc/rfc9309.html" class="external-link" rel="noopener"&gt;RFC 9309&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;, Koster, Illyes, Zeller and Sassman,
September 2022. Three provisions are worth knowing before you lean on it.&lt;/p&gt;
&lt;p&gt;It is not a security control. The Security Considerations section says so directly: &amp;ldquo;The
Robots Exclusion Protocol is not a substitute for valid content security measures.&amp;rdquo; Anything
reachable without authentication is reachable.&lt;/p&gt;
&lt;p&gt;It is cached. &amp;ldquo;Crawlers SHOULD NOT use the cached version for more than 24 hours,&amp;rdquo; which lines
up with OpenAI&amp;rsquo;s stated ~24-hour adjustment window. Deploying a robots.txt change and
expecting immediate effect is a category error.&lt;/p&gt;
&lt;p&gt;It is size-limited. &amp;ldquo;The parsing limit MUST be at least 500 kibibytes.&amp;rdquo; Rules past that
boundary have no guarantee of being read at all.&lt;/p&gt;
&lt;h2 id="not-documented-whether-any-of-them-run-javascript"&gt;Not documented: whether any of them run JavaScript&lt;/h2&gt;
&lt;p&gt;The question I get asked most, and none of the operators has answered it. On 5 August 2026 I
fetched the four crawler documentation pages named above and searched their rendered text for
the string &amp;ldquo;JavaScript&amp;rdquo;. Anthropic&amp;rsquo;s, Google&amp;rsquo;s and Perplexity&amp;rsquo;s contain it zero times.
OpenAI&amp;rsquo;s contains it once, in the left-hand navigation tree, as a link to an unrelated
&amp;ldquo;JavaScript Pixel&amp;rdquo; page in the advertising docs.&lt;/p&gt;
&lt;p&gt;The only party documenting rendering is Google Search, and it documents it for Googlebot
rather than for anything AI-specific: &amp;ldquo;Googlebot queues all pages with a &lt;code&gt;200&lt;/code&gt; HTTP status
code for rendering, unless a robots &lt;code&gt;meta&lt;/code&gt; tag or header tells Google not to index the page.
The page may stay on this queue for a few seconds, but it can take longer than that&amp;rdquo;
(&lt;a href="https://developers.google.com/search/docs/crawling-indexing/javascript/javascript-seo-basics" class="external-link" rel="noopener"&gt;Understand the JavaScript SEO basics&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;,
last updated 4 March 2026).&lt;/p&gt;
&lt;p&gt;So the defensible statement is a narrow one: rendering is documented for Googlebot and
undocumented everywhere else. My practical conclusion (inference, not documentation) is to
put the substance in the initial HTML response, since that is the only path every fetcher
demonstrably has.&lt;/p&gt;
&lt;h2 id="proposed-not-adopted-llmstxt"&gt;Proposed, not adopted: llms.txt&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://llmstxt.org/" class="external-link" rel="noopener"&gt;llms.txt&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; was published by Jeremy Howard on 3 September 2024, and it
describes itself accurately, which is the part people skip. The site&amp;rsquo;s own subtitle: &amp;ldquo;A
proposal to standardise on using an /llms.txt file to provide information to help LLMs use a
website at inference time.&amp;rdquo; Further down: &amp;ldquo;the &lt;code&gt;llms.txt&lt;/code&gt; specification is open for community
input.&amp;rdquo; It specifies a Markdown file with an H1 project name, an optional summary blockquote,
and H2 sections containing lists of links.&lt;/p&gt;
&lt;p&gt;A proposal. Rather than assert that no major provider has adopted it, here is the check, which
takes about two minutes: of the four crawler documentation pages cited above, Anthropic&amp;rsquo;s and
Google&amp;rsquo;s contain the string &amp;ldquo;llms.txt&amp;rdquo; zero times. OpenAI&amp;rsquo;s and Perplexity&amp;rsquo;s do contain it.
In both cases every occurrence is a link to that documentation site&amp;rsquo;s &lt;em&gt;own&lt;/em&gt; &lt;code&gt;/llms.txt&lt;/code&gt;
index, offered to agents reading the docs. &lt;code&gt;developers.openai.com/llms.txt&lt;/code&gt; and
&lt;code&gt;docs.perplexity.ai/llms.txt&lt;/code&gt; both return 200.&lt;/p&gt;
&lt;p&gt;That distinction carries the whole section. Two providers publish an llms.txt for their
documentation. Neither one&amp;rsquo;s crawler documentation says its crawlers read yours. Serving a
file and consuming a file are different behaviours, and the claim &amp;ldquo;OpenAI supports llms.txt&amp;rdquo;
collapses them.&lt;/p&gt;
&lt;p&gt;None of which is an argument against publishing one. It costs almost nothing, and a Markdown
index is genuinely useful to an agent that fetches it. The argument is against believing that
filing it does anything by itself, and against the second failure mode, which is writing
claims into a machine-readable file that your pages do not support. Two of mine did that.
&lt;code&gt;/.well-known/ai-plugin.json&lt;/code&gt; pointed at an &lt;code&gt;openapi.yaml&lt;/code&gt; of 189 bytes whose &lt;code&gt;paths&lt;/code&gt; object
was empty: a plugin manifest advertising an API with no operations in it, filed at the
well-known path where machines go looking. Searching OpenAI&amp;rsquo;s current
&lt;a href="https://developers.openai.com/plugins" class="external-link" rel="noopener"&gt;Plugins documentation&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt; on 5 August 2026, the strings
&lt;code&gt;ai-plugin.json&lt;/code&gt; and &lt;code&gt;well-known&lt;/code&gt; appear zero times. The page describes something built on
skills and the Model Context Protocol instead. A dead spec, implemented badly. The &lt;code&gt;ai.txt&lt;/code&gt;
and &lt;code&gt;llms.txt&lt;/code&gt; next to it were worse, because they were live and false: they advertised
&amp;ldquo;19 products tracked, 171 comparison pages, updated hourly&amp;rdquo; and &amp;ldquo;public review data from 10+
sources,&amp;rdquo; about a benchmark that
&lt;a href="https://therezaali.com/writing/honest-comparison-sites/"&gt;fails most of its own published promises&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;The direction of travel on access runs the opposite way from manifests. On 1 July 2025
Cloudflare announced it was
&lt;a href="https://blog.cloudflare.com/content-independence-day-no-ai-crawl-without-compensation/" class="external-link" rel="noopener"&gt;&amp;ldquo;changing the default to block AI crawlers unless they pay creators for their content&amp;rdquo;&lt;svg class="external-mark" aria-hidden="true" focusable="false" width="10" height="10" viewBox="0 0 10 10" fill="none" stroke="currentColor" stroke-width="1.4" stroke-linecap="round" stroke-linejoin="round"&gt;&lt;path d="M3 7 7 3"/&gt;&lt;path d="M3.6 3H7v3.4"/&gt;&lt;/svg&gt;&lt;/a&gt;,
citing referral ratios: &amp;ldquo;With OpenAI, it&amp;rsquo;s 750 times more difficult to get traffic than it was
with the Google of old. With Anthropic, it&amp;rsquo;s 30,000 times more difficult.&amp;rdquo; Whatever you make
of the policy, &amp;ldquo;what will read my site&amp;rdquo; is now answered partly by your CDN and not only by
your robots file.&lt;/p&gt;
&lt;h2 id="inference-what-actually-gets-you-cited"&gt;Inference: what actually gets you cited&lt;/h2&gt;
&lt;p&gt;Argued, not documented. Read it accordingly.&lt;/p&gt;
&lt;p&gt;Nobody outside the labs knows how retrieval weights or selects a page. Every provider above
documents which agent fetches what and how to block it. None documents what happens next, so
anyone offering you the ranking factors for LLM citation is describing a system whose
operators have published nothing whatsoever about its ranking factors.&lt;/p&gt;
&lt;p&gt;What can be reasoned about is narrower. To repeat a page without hedging, a model needs a
specific claim, attached to a named source and a date, that resolves to something a reader can
open. A manifest describing you favourably supplies none of that: it is self-assertion, and
self-assertion is the cheapest text there is: the exact output a language model generates
without limit at no cost, which is why it carries no weight coming from you either.&lt;/p&gt;
&lt;p&gt;So the durable move is to be the primary source of a fact somebody needs. Publish the number
you measured, the method you used, the date. A fact available only from you can be cited; a
restatement of what a hundred other pages already say cannot, and no file in your web root
changes that. I am fairly confident about that and I cannot prove it. Inference, labelled as
inference. That is the arrangement I should have used the first time.&lt;/p&gt;</content:encoded></item><item><title>Page 6 tells the raters their ratings do not move rankings</title><link>https://therezaali.com/writing/eeat-quality-rater-guidelines/</link><guid isPermaLink="true">https://therezaali.com/writing/eeat-quality-rater-guidelines/</guid><pubDate>Wed, 05 Aug 2026 00:00:00 +0200</pubDate><author>contact@therezaali.com (Mo RezaAli)</author><category>reference</category><description>The quality rater guidelines run to 182 pages, carry the date 11 September 2025, and cost nothing to download. They name Trust as the most important member of E-E-A-T. Quoted from the PDF.</description><content:encoded>&lt;p&gt;The file is 182 pages, it is called &lt;em&gt;General Guidelines&lt;/em&gt;, its header carries the date
11 September 2025, Google links to it from the helpful-content documentation, and it costs
nothing to download.&lt;/p&gt;
&lt;p&gt;Page 6 tells the raters, before they rate anything, that their work does not move rankings.
Which is roughly the opposite of what gets asserted on this document&amp;rsquo;s behalf.&lt;/p&gt;
&lt;p&gt;So: quotes rather than summary. Section numbers refer to that version, and every block quote
below is the document&amp;rsquo;s own wording.&lt;/p&gt;
&lt;h2 id="the-hierarchy-verbatim"&gt;The hierarchy, verbatim&lt;/h2&gt;
&lt;p&gt;Section 3.4 names the hierarchy before it defines a single term:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Experience, Expertise, Authoritativeness and Trust (E-E-A-T) are all important
considerations in PQ rating. The most important member at the center of the E-E-A-T family
is Trust.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;General Guidelines, 11 September 2025, section 3.4&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;Trust is defined first, and defined as a property of the page rather than of the author:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Trust: Consider the extent to which the page is accurate, honest, safe, and reliable.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;General Guidelines, section 3.4&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;The other three arrive as supporting evidence for that judgement. The document&amp;rsquo;s own framing
is that they &amp;ldquo;can support your assessment of Trust&amp;rdquo;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Experience: Consider the extent to which the content creator has the necessary first-hand
or life experience for the topic.&lt;/p&gt;
&lt;p&gt;Expertise: Consider the extent to which the content creator has the necessary knowledge or
skill for the topic.&lt;/p&gt;
&lt;p&gt;Authoritativeness: Consider the extent to which the content creator or the website is known
as a go-to source for the topic.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;General Guidelines, section 3.4&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;Then the sentence that gets left out of every summary, stated later in the same section with
an example that is hard to misread:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Trust is the most important member of the E-E-A-T family because untrustworthy pages have
low E-E-A-T no matter how Experienced, Expert, or Authoritative they may seem. For example,
a financial scam is untrustworthy, even if the content creator is a highly experienced and
expert scammer who is considered the go-to on running scams!&lt;/p&gt;
&lt;p&gt;&lt;em&gt;General Guidelines, section 3.4&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;Max out three of the four letters and the page can still score Lowest. The acronym does not
describe four boxes of equal weight, whatever the diagrams say. It is one criterion and three
kinds of supporting evidence, and Google&amp;rsquo;s public documentation says as much in a line:
&amp;ldquo;Of these aspects, trust is most important. The others contribute to trust, but content
doesn&amp;rsquo;t necessarily have to demonstrate all of them.&amp;rdquo;&lt;/p&gt;
&lt;h2 id="what-page-6-tells-the-raters-about-their-own-work"&gt;What page 6 tells the raters about their own work&lt;/h2&gt;
&lt;p&gt;Section 0.1, before any rating happens:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;No single rating can directly impact how a particular webpage, website, or result appears
in Google Search, nor can it cause specific webpages, websites, or results to move up or
down on the search results page. Using ratings to position results on the search results
page would not be feasible, as humans could never individually rate each page on the open
web.&lt;/p&gt;
&lt;p&gt;Instead, ratings are used to measure how effectively search engines are working to deliver
helpful content to people around the world.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;General Guidelines, section 0.1&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;Google repeats it in three places outside the PDF, which I checked one at a time.&lt;/p&gt;
&lt;p&gt;The helpful-content page: &amp;ldquo;Search raters have no control over how pages rank. Rater data is
not used directly in our ranking algorithms. Rather, we use them as a restaurant might get
feedback cards from diners.&amp;rdquo; The generative-AI guidance page says the guidelines &amp;ldquo;are not a
guide to ranking first in Google,&amp;rdquo; and that raters&amp;rsquo; &amp;ldquo;ratings don&amp;rsquo;t directly influence
ranking.&amp;rdquo; The December 2022 post that added the second E, on the raters: &amp;ldquo;they don&amp;rsquo;t directly
influence ranking.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;The same helpful-content page also disposes of the phrase everyone uses, in a subordinate
clause, which is presumably why nobody quotes it: &amp;ldquo;While E-E-A-T itself isn&amp;rsquo;t a specific
ranking factor, using a mix of factors that can identify content with good E-E-A-T is
useful.&amp;rdquo; Note the shape of that sentence. Not a factor. A mix of factors that can identify
the thing.&lt;/p&gt;
&lt;h2 id="four-readings-the-document-does-not-support"&gt;Four readings the document does not support&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;&amp;ldquo;E-E-A-T is a ranking factor.&amp;rdquo;&lt;/strong&gt; Covered above, so only the part that usually gets dropped:
that mix of factors is weighted more heavily on topics &amp;ldquo;that could significantly impact the
health, financial stability, or safety of people, or the welfare or well-being of society.&amp;rdquo;
Which is the definition of YMYL, given in section 2.3.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;&amp;ldquo;You need formal credentials.&amp;rdquo;&lt;/strong&gt; For a whole class of pages, section 3.4.1 says the
opposite: &amp;ldquo;Pages that share first-hand life experience on clear YMYL topics may be considered
to have high E-E-A-T as long as the content is trustworthy, safe, and consistent with
well-established expert consensus.&amp;rdquo; The December 2022 post drew the same line with tax. Filing
your return correctly is a question for &amp;ldquo;an expert in the field of accounting.&amp;rdquo; Choosing tax
software, the post suggests, might send you to &amp;ldquo;a forum discussion from people who have
experience with different services&amp;rdquo; instead.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;&amp;ldquo;Expertise is the most important letter.&amp;rdquo;&lt;/strong&gt; Section 3.4 says Trust is, twice, and gives the
expert-scammer example specifically to close this reading.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;&amp;ldquo;A small site with no press coverage looks bad.&amp;rdquo;&lt;/strong&gt; Section 3.3.5 addresses this directly:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;However, small websites may have little or no reputation information. This is not
indicative of high or low quality. [&amp;hellip;] A lack of reputation about people who post
personal content is neither a positive nor a negative sign in your assessment of the page.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;General Guidelines, section 3.3.5&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;Absence of coverage is scored as absence, not as a negative. That is a meaningfully different
instruction from the one most SEO advice implies.&lt;/p&gt;
&lt;h2 id="what-the-document-asks-of-a-small-site"&gt;What the document asks of a small site&lt;/h2&gt;
&lt;p&gt;Section 2.5.2 is titled &lt;em&gt;Finding Who is Responsible for the Website and Who Created the
Content on the Page&lt;/em&gt;, and it sets two separate requirements:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Every page belongs to a website, and it should be clear:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Who (what individuals, company, business, organization, government agency, etc.) is
responsible for the website.&lt;/li&gt;
&lt;li&gt;Who (what individuals, company, business, organization, government agency, etc.) created
the content on the page you are evaluating.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;em&gt;General Guidelines, section 2.5.2&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;Two questions, asked separately. Who runs the site, and who wrote this page. On a one-person
site both answers are the same name, which is the rare structural advantage of being one
person: nothing to navigate, nothing to trace up an org chart.&lt;/p&gt;
&lt;p&gt;Section 2.5.3 then grants personal sites an allowance that stores and banks do not get:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Occasionally, you may encounter a website or content creator with a legitimate reason for
anonymity. For example, personal websites may omit personal contact information such as an
individual&amp;rsquo;s home address or phone number.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;General Guidelines, section 2.5.3&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;No street address required, then. Identifiable and reachable, yes. Those are two different
bars, and running them together is how people end up either publishing their home address or
hiding behind a contact form with no name on it.&lt;/p&gt;
&lt;p&gt;Section 3.4 names three sources of evidence a rater is told to weigh:&lt;/p&gt;
&lt;blockquote&gt;
&lt;ul&gt;
&lt;li&gt;What the website or content creators say about themselves&lt;/li&gt;
&lt;li&gt;What others say about the website or content creators&lt;/li&gt;
&lt;li&gt;What is visible on the page, including the Main Content and sections such as reviews and
comments&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;em&gt;General Guidelines, section 3.4&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;Only the first of those is under a publisher&amp;rsquo;s control. The second belongs to other people,
and the only way to control it is to fake it, which section 3.3.5 has already declined to
reward, since it scores an absent reputation as neutral. The third is whatever is actually on
the page.&lt;/p&gt;
&lt;p&gt;The section closes on a warning that is easy to skim past:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Important: The website or content creator may not be a trustworthy source if there is a
clear conflict of interest.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;General Guidelines, section 3.4&lt;/em&gt;&lt;/p&gt;&lt;/blockquote&gt;
&lt;p&gt;Google&amp;rsquo;s example is a manufacturer reviewing its own product. Generalise the shape (someone
with a stake evaluating the thing they have a stake in) and it covers an affiliate link
inside a &amp;ldquo;best of&amp;rdquo; list, a review written by the vendor&amp;rsquo;s own agency, and an undisclosed
employment history sitting inside a software comparison. The document does not treat any of
those as a footnote. It treats them as grounds for distrusting the source outright.&lt;/p&gt;
&lt;h2 id="how-to-read-the-thing"&gt;How to read the thing&lt;/h2&gt;
&lt;p&gt;Not as a ranking manual. Page 6 rules that out.&lt;/p&gt;
&lt;p&gt;Read it instead as the most detailed public account anyone has of what a trained stranger,
paid to be sceptical, is told to check when they land on a page cold. Read that way, the
instructions for a small site come out short.&lt;/p&gt;
&lt;p&gt;Be identifiable: a real name, attached to a real role, at an organisation that exists. Be
reachable, which section 2.5.3 is careful to say does not mean publishing a home address.
Separate what you have used from what you have only read about, because the document defines
Experience and Expertise apart from each other and puts both underneath Trust. Put the
conflict on the page before the reader finds it.&lt;/p&gt;
&lt;p&gt;And if nobody has ever written about you, section 3.3.5 says that is worth nothing. Not minus
something. Nothing. Which is also not an invitation to go and manufacture some.&lt;/p&gt;</content:encoded></item><item><title>LCP is four consecutive spans of time, not one number</title><link>https://therezaali.com/writing/core-web-vitals-lcp/</link><guid isPermaLink="true">https://therezaali.com/writing/core-web-vitals-lcp/</guid><pubDate>Wed, 05 Aug 2026 00:00:00 +0200</pubDate><author>contact@therezaali.com (Mo RezaAli)</author><category>reference</category><description>Until you know which of the four is eating the budget, you are guessing at which fix to apply. Four checks for a static site, each with the command that finds the problem.</description><content:encoded>&lt;p&gt;LCP is usually discussed as though it were a single number you make smaller. It is not. It is
four consecutive spans of time laid end to end, and until you know which of the four is eating
your budget you are guessing at which fix to apply.&lt;/p&gt;
&lt;p&gt;Largest Contentful Paint reports the render time of the largest image or text block visible in
the viewport. web.dev&amp;rsquo;s LCP page gives the target: &amp;ldquo;sites should strive to have Largest
Contentful Paint of 2.5 seconds or less,&amp;rdquo; and it is explicit about where to read that from:
&amp;ldquo;a good threshold to measure is the 75th percentile of page loads, segmented across mobile and
desktop devices.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;The percentile is the harder half of that sentence. Your 75th-percentile visitor is not your
average one. Old phone, poor connection, cold cache. A comfortable median tells you very
little about the number Google is actually scoring, and the gap between the two is where most
people&amp;rsquo;s optimism lives.&lt;/p&gt;
&lt;p&gt;So the useful question is never &amp;ldquo;is my LCP good&amp;rdquo; but &amp;ldquo;which of the four spans is spending the
budget, on the quarter of visits I have not been looking at&amp;rdquo;. That distinction is the whole
point, because a 3.4-second LCP caused by a
font that blocks a headline and a 3.4-second LCP caused by an uncompressed hero on a slow
connection are the same number describing two entirely different repairs.&lt;/p&gt;
&lt;p&gt;One thing to fix before the model. Text counts. LCP candidates are &lt;code&gt;&amp;lt;img&amp;gt;&lt;/code&gt;, &lt;code&gt;&amp;lt;image&amp;gt;&lt;/code&gt; inside
&lt;code&gt;&amp;lt;svg&amp;gt;&lt;/code&gt;, &lt;code&gt;&amp;lt;video&amp;gt;&lt;/code&gt;, an element with a background image loaded via &lt;code&gt;url()&lt;/code&gt;, and block-level
elements containing text nodes or inline text children. So if your hero is a headline in a
webfont, the font file is on the LCP path and you should treat it exactly as seriously as an
image.&lt;/p&gt;
&lt;h2 id="the-four-sub-parts"&gt;The four sub-parts&lt;/h2&gt;
&lt;p&gt;web.dev&amp;rsquo;s optimization guide breaks LCP into four consecutive parts and gives a target share
for each.&lt;/p&gt;
&lt;table&gt;
 &lt;caption&gt;The four LCP sub-parts and their target share of total LCP, per web.dev's Optimize Largest Contentful Paint guide. The guide notes these are guidelines rather than strict rules.&lt;/caption&gt;
 &lt;thead&gt;
 &lt;tr&gt;
 &lt;th scope="col"&gt;Sub-part&lt;/th&gt;
 &lt;th scope="col"&gt;What it is&lt;/th&gt;
 &lt;th scope="col"&gt;Target share of LCP&lt;/th&gt;
 &lt;/tr&gt;
 &lt;/thead&gt;
 &lt;tbody&gt;
 &lt;tr&gt;
 &lt;td&gt;Time to first byte&lt;/td&gt;
 &lt;td&gt;until the first byte of the HTML response&lt;/td&gt;
 &lt;td&gt;~40%&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Resource load delay&lt;/td&gt;
 &lt;td&gt;from TTFB until the LCP resource starts loading&lt;/td&gt;
 &lt;td&gt;&amp;lt;10%&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Resource load duration&lt;/td&gt;
 &lt;td&gt;transferring the LCP resource&lt;/td&gt;
 &lt;td&gt;~40%&lt;/td&gt;
 &lt;/tr&gt;
 &lt;tr&gt;
 &lt;td&gt;Element render delay&lt;/td&gt;
 &lt;td&gt;from resource loaded until the element paints&lt;/td&gt;
 &lt;td&gt;&amp;lt;10%&lt;/td&gt;
 &lt;/tr&gt;
 &lt;/tbody&gt;
&lt;/table&gt;
&lt;p&gt;The two small numbers are the diagnostic ones. web.dev flags this itself: of the four
sub-parts, two have the word &amp;ldquo;delay&amp;rdquo; in their names, which &amp;ldquo;is a clue that you want to get
these times as close to zero as possible.&amp;rdquo; Nothing useful happens during a delay. If either
delay is large you have a discovery problem or a blocking problem, and those are usually the
cheap fixes. TTFB and load duration are real work, which is why they get 40% each.&lt;/p&gt;
&lt;p&gt;For a static site behind a CDN, hosting has already largely solved TTFB. Which leaves load
delay plus load duration. On most pages both of those are properties of exactly one
image.&lt;/p&gt;
&lt;p&gt;The guide adds a warning worth repeating, because the table invites the mistake: &amp;ldquo;Given the
2.5 second target for LCP, it may be tempting to try to convert these percentages into
absolute numbers, but that is not recommended.&amp;rdquo;&lt;/p&gt;
&lt;h2 id="the-audit-in-the-order-i-run-it"&gt;The audit, in the order I run it&lt;/h2&gt;
&lt;p&gt;Four checks, each one a command you can run against your own repository in a few seconds. The
sample outputs below come from this site&amp;rsquo;s own history at &lt;code&gt;22b7c22d^&lt;/code&gt; (the commit before the
machine-generated corpus was deleted), because that is a codebase whose lines I can quote in
full. It is private, so the file facts are mine and the reasoning from them is Google&amp;rsquo;s, linked
throughout.&lt;/p&gt;
&lt;h3 id="check-1-weigh-everything-the-templates-can-put-above-the-fold"&gt;Check 1: weigh everything the templates can put above the fold&lt;/h3&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;git ls-tree -r --long 22b7c22d^ -- static/images/rezaali-fallback.png
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# 100644 blob 5932e974… 2074882 static/images/rezaali-fallback.png&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;2,074,882 bytes. 1536 × 1024. Non-interlaced 8-bit RGB PNG. Its visible content is the word
&lt;code&gt;LOADING..&lt;/code&gt; in white on a near-black field, with the word &lt;code&gt;SIGNALS&lt;/code&gt; under it.&lt;/p&gt;
&lt;p&gt;A placeholder that never got replaced is an ordinary accident. What turns it structural is
where the path is written down, so the second half of this check is to grep your archetypes
and partials for it:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;git show 22b7c22d^:archetypes/post.md &lt;span class="p"&gt;|&lt;/span&gt; grep image:
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# image: &amp;#34;images/rezaali-fallback.png&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Every post created from that archetype inherited the line, which made two megabytes of PNG the
default cover image. A default is the one thing nobody looks at.&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;git grep -l &lt;span class="s1"&gt;&amp;#39;images/rezaali-fallback.png&amp;#39;&lt;/span&gt; 22b7c22d^ -- content/post/insights &lt;span class="se"&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="se"&gt;&lt;/span&gt; &lt;span class="p"&gt;|&lt;/span&gt; grep -vE &lt;span class="s1"&gt;&amp;#39;\.(it|ar|zh-cn)\.md$&amp;#39;&lt;/span&gt; &lt;span class="p"&gt;|&lt;/span&gt; wc -l
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# 229&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;229 English posts, each served through a template that rendered the cover as an &lt;code&gt;&amp;lt;img&amp;gt;&lt;/code&gt;
carrying &lt;code&gt;loading=&amp;quot;lazy&amp;quot;&lt;/code&gt;. On the branch that fires for a static path rather than a Hugo
page resource, it also carried no &lt;code&gt;width&lt;/code&gt; and no &lt;code&gt;height&lt;/code&gt; at all.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The repair.&lt;/strong&gt; Not &amp;ldquo;compress the placeholder.&amp;rdquo; Do not have a placeholder. Ship no image, let
text be the LCP element, and pay a few kilobytes instead of two megabytes. PNG is lossless,
and lossless spends its bytes faithfully preserving photographic detail that a flat graphic
with three colours does not have, so the format was wrong for the content and the content was
wrong for the page. When a real photograph is ready, encode it as WebP or AVIF. Can I use
maintains current support tables for both, and WebP is the safer floor if you are shipping one
format with no &lt;code&gt;&amp;lt;picture&amp;gt;&lt;/code&gt; fallback chain.&lt;/p&gt;
&lt;h3 id="check-2-find-out-whether-your-hero-is-lazy-loaded"&gt;Check 2: find out whether your hero is lazy-loaded&lt;/h3&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;git grep -n &lt;span class="s1"&gt;&amp;#39;class=&amp;#34;portrait&amp;#34;&amp;#39;&lt;/span&gt; 22b7c22d^ -- layouts/index.html
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-go-html-template" data-lang="go-html-template"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;img&lt;/span&gt; &lt;span class="na"&gt;class&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;portrait&amp;#34;&lt;/span&gt; &lt;span class="na"&gt;src&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;&lt;/span&gt;&lt;span class="cp"&gt;{{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s"&gt;&amp;#34;images/mo-rezaali.jpeg&amp;#34;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;|&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;relURL&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="cp"&gt;}}&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="na"&gt;alt&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;Portrait of Mo RezaAli&amp;#34;&lt;/span&gt; &lt;span class="na"&gt;loading&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;lazy&amp;#34;&lt;/span&gt; &lt;span class="p"&gt;/&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;A homepage hero portrait: 311,372 bytes, inside the first screen, carrying an instruction to
the browser not to hurry. No &lt;code&gt;width&lt;/code&gt;, no &lt;code&gt;height&lt;/code&gt;, no &lt;code&gt;srcset&lt;/code&gt;, no &lt;code&gt;fetchpriority&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;&lt;code&gt;loading=&amp;quot;lazy&amp;quot;&lt;/code&gt; defers an image until it approaches the viewport. Correct below the fold.
Exactly wrong at the top, because browser-level lazy loading has to wait for layout before it
can decide whether to fetch, which pushes the request behind CSS. web.dev&amp;rsquo;s optimization guide
states it without hedging: &amp;ldquo;Never lazy-load your LCP image, as that will always lead to
unnecessary resource load delay, and will have a negative impact on LCP.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Read that against the table. Lazy-loading your hero takes the sub-part with a target under 10%
and makes it the largest line in the budget.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The repair.&lt;/strong&gt; &lt;code&gt;loading=&amp;quot;lazy&amp;quot;&lt;/code&gt; on every image except the LCP candidate, which gets
&lt;code&gt;fetchpriority=&amp;quot;high&amp;quot;&lt;/code&gt; instead.&lt;/p&gt;
&lt;h3 id="check-3-confirm-you-have-any-responsive-images-at-all"&gt;Check 3: confirm you have any responsive images at all&lt;/h3&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;git grep -l srcset 22b7c22d^ -- layouts static
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# (no output)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;No output is the finding. Not a single responsive image anywhere on the site: a phone at 400
CSS pixels wide downloaded the same file as a 5K desktop. Since load duration is around 40% of
the budget, and duration is bytes over bandwidth, serving four times the necessary bytes to
the slowest connections is the most direct route there is to failing at the 75th percentile.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The repair.&lt;/strong&gt;&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-html" data-lang="html"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;img&lt;/span&gt; &lt;span class="na"&gt;src&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;/img/mo-800.webp&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="na"&gt;srcset&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;/img/mo-400.webp 400w,
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s"&gt; /img/mo-800.webp 800w,
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s"&gt; /img/mo-1200.webp 1200w&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="na"&gt;sizes&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;(max-width: 40rem) 90vw, 400px&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="na"&gt;width&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;800&amp;#34;&lt;/span&gt; &lt;span class="na"&gt;height&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;800&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="na"&gt;alt&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;Mo RezaAli&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="na"&gt;fetchpriority&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;high&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="na"&gt;decoding&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;async&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Two details are not optional. &lt;code&gt;width&lt;/code&gt; and &lt;code&gt;height&lt;/code&gt; must be there so the browser can reserve
the box before the bytes land; a missing intrinsic size is a layout shift, and CLS has its own
threshold of 0.1 at the 75th percentile. And &lt;code&gt;sizes&lt;/code&gt; has to describe your real CSS layout. If
&lt;code&gt;sizes&lt;/code&gt; lies, the browser picks the wrong candidate and you have made the page worse than a
single fixed image would have been.&lt;/p&gt;
&lt;p&gt;&lt;code&gt;fetchpriority&lt;/code&gt; raises or lowers a resource&amp;rsquo;s priority within its class. web.dev&amp;rsquo;s Fetch
Priority guide covers both directions: &lt;code&gt;high&lt;/code&gt; for the hero, &lt;code&gt;low&lt;/code&gt; for images you know are
decorative.&lt;/p&gt;
&lt;h4 id="when-the-image-is-not-in-the-html"&gt;When the image is not in the HTML&lt;/h4&gt;
&lt;p&gt;&lt;code&gt;fetchpriority&lt;/code&gt; only helps if the preload scanner can see the image. A CSS background or a
JavaScript-injected hero cannot be discovered until the stylesheet or the script has been
parsed, and that wait is the definition of a large resource load delay. Preload it,
responsively, so you do not undo the &lt;code&gt;srcset&lt;/code&gt; work:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-html" data-lang="html"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;link&lt;/span&gt; &lt;span class="na"&gt;rel&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;preload&amp;#34;&lt;/span&gt; &lt;span class="na"&gt;as&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;image&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="na"&gt;href&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;/img/mo-800.webp&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="na"&gt;imagesrcset&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;/img/mo-400.webp 400w,
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s"&gt; /img/mo-800.webp 800w,
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s"&gt; /img/mo-1200.webp 1200w&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="na"&gt;imagesizes&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;(max-width: 40rem) 90vw, 400px&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="na"&gt;fetchpriority&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;high&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;&lt;code&gt;imagesrcset&lt;/code&gt; and &lt;code&gt;imagesizes&lt;/code&gt; mirror the &lt;code&gt;&amp;lt;img&amp;gt;&lt;/code&gt; attributes; web.dev&amp;rsquo;s guide to preloading
responsive images covers the pairing. A plain &lt;code&gt;href&lt;/code&gt; preload sitting next to a responsive
&lt;code&gt;&amp;lt;img&amp;gt;&lt;/code&gt; is a classic own goal. It fetches a fifth copy at the wrong size.&lt;/p&gt;
&lt;p&gt;Where you can get it, the better answer is to put the hero in an &lt;code&gt;&amp;lt;img&amp;gt;&lt;/code&gt; tag in the HTML and
delete the preload entirely. Preload is a workaround for resources the parser cannot find.&lt;/p&gt;
&lt;h3 id="check-4-read-the-first-three-lines-of-your-head-partial"&gt;Check 4: read the first three lines of your head partial&lt;/h3&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;git show 22b7c22d^:layouts/partials/mo/head.html &lt;span class="p"&gt;|&lt;/span&gt; head -3
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-go-html-template" data-lang="go-html-template"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="cp"&gt;{{-&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;$cssBust&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;:=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;now&lt;/span&gt;&lt;span class="na"&gt;.Unix&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="cp"&gt;-}}&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;link&lt;/span&gt; &lt;span class="na"&gt;rel&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;stylesheet&amp;#34;&lt;/span&gt; &lt;span class="na"&gt;href&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;&lt;/span&gt;&lt;span class="cp"&gt;{{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s"&gt;&amp;#34;css/new-face-mo.css&amp;#34;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;|&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;relURL&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="cp"&gt;}}&lt;/span&gt;&lt;span class="s"&gt;?v=&lt;/span&gt;&lt;span class="cp"&gt;{{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;$cssBust&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="cp"&gt;}}&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;link&lt;/span&gt; &lt;span class="na"&gt;rel&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;stylesheet&amp;#34;&lt;/span&gt; &lt;span class="na"&gt;href&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;&lt;/span&gt;&lt;span class="cp"&gt;{{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s"&gt;&amp;#34;css/mo-site.css&amp;#34;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;|&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;relURL&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="cp"&gt;}}&lt;/span&gt;&lt;span class="s"&gt;?v=&lt;/span&gt;&lt;span class="cp"&gt;{{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;$cssBust&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="cp"&gt;}}&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;&lt;code&gt;now.Unix&lt;/code&gt; is the current Unix time. As a cache key it has one disqualifying property: it is
not derived from the file&amp;rsquo;s contents. Two consequences follow.&lt;/p&gt;
&lt;p&gt;Every deploy changes the query string on every stylesheet whether or not a byte of CSS has
changed, so returning visitors re-download identical files forever. Cache-busting that fires
unconditionally is cache-disabling with extra steps.&lt;/p&gt;
&lt;p&gt;Worse, the value is computed while each template renders rather than once per build. Pages
rendered on either side of a second boundary get different values, so one deploy can serve
three distinct URLs for one physical file. Three URLs, three cache entries, and a visitor
moving from page to page cannot reuse the stylesheet they fetched a moment earlier. Cache hits
are not unlikely here. They are structurally impossible.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The repair&lt;/strong&gt; is a content hash. Hugo&amp;rsquo;s &lt;code&gt;resources.Fingerprint&lt;/code&gt; (&amp;ldquo;Cryptographically hashes
the content of the given resource.&amp;rdquo;) rewrites the filename to include the hash, defaulting to
SHA-256:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-go-html-template" data-lang="go-html-template"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="cp"&gt;{{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="k"&gt;with&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;resources&lt;/span&gt;&lt;span class="na"&gt;.Get&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s"&gt;&amp;#34;css/site.css&amp;#34;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="cp"&gt;}}&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="cp"&gt;{{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;$css&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;:=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="na"&gt;.&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;|&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;minify&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;|&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;fingerprint&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s"&gt;&amp;#34;sha256&amp;#34;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="cp"&gt;}}&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;link&lt;/span&gt; &lt;span class="na"&gt;rel&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;stylesheet&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="na"&gt;href&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;&lt;/span&gt;&lt;span class="cp"&gt;{{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;$css&lt;/span&gt;&lt;span class="na"&gt;.RelPermalink&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="cp"&gt;}}&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="na"&gt;integrity&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;&lt;/span&gt;&lt;span class="cp"&gt;{{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;$css&lt;/span&gt;&lt;span class="na"&gt;.Data.Integrity&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="cp"&gt;}}&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="na"&gt;crossorigin&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;anonymous&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="cp"&gt;{{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="k"&gt;end&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="cp"&gt;}}&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The URL now changes if and only if the CSS changes, and every page in a build gets the same
URL because the input is the same. Because the URL is content-addressed, the file can then be
cached permanently. On Cloudflare Pages, via &lt;code&gt;static/_headers&lt;/code&gt;:&lt;/p&gt;
&lt;pre tabindex="0"&gt;&lt;code&gt;/css/*
 Cache-Control: public, max-age=31536000, immutable
&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;&lt;code&gt;immutable&lt;/code&gt; tells the browser not to revalidate during the freshness lifetime, which is safe
precisely because a changed file gets a different name. MDN&amp;rsquo;s &lt;code&gt;Cache-Control&lt;/code&gt; reference
documents the directive. Do not put this on your HTML. HTML has to stay revalidatable or
nobody ever sees a new deploy.&lt;/p&gt;
&lt;h2 id="fonts-the-text-lcp-path"&gt;Fonts: the text-LCP path&lt;/h2&gt;
&lt;p&gt;If your LCP element is a headline rather than an image, the webfont is on the critical path
and the failure mode is invisible text.&lt;/p&gt;
&lt;p&gt;&lt;code&gt;font-display: swap&lt;/code&gt; renders the fallback face immediately and swaps when the webfont arrives;
MDN&amp;rsquo;s &lt;code&gt;font-display&lt;/code&gt; reference documents the timeline values. &lt;code&gt;swap&lt;/code&gt; buys you text that is
never invisible at the cost of a flash of fallback. That is the right trade when the text &lt;em&gt;is&lt;/em&gt;
the LCP element, because under &lt;code&gt;block&lt;/code&gt; the browser hides the text during its block period and
your LCP sits waiting on a font file.&lt;/p&gt;
&lt;p&gt;The cost is layout shift. Fallback and webfont almost never share metrics. Three mitigations,
ordered by how much I trust them:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Self-host the font and &lt;code&gt;&amp;lt;link rel=&amp;quot;preload&amp;quot; as=&amp;quot;font&amp;quot; type=&amp;quot;font/woff2&amp;quot; crossorigin&amp;gt;&lt;/code&gt; the
specific faces used above the fold. The &lt;code&gt;crossorigin&lt;/code&gt; attribute is required even
same-origin. Omit it and the browser fetches the file twice.&lt;/li&gt;
&lt;li&gt;Ship only the faces you actually use. Declaring a family you do not serve guarantees a
fallback render for no benefit.&lt;/li&gt;
&lt;li&gt;Tune the fallback with &lt;code&gt;size-adjust&lt;/code&gt;, &lt;code&gt;ascent-override&lt;/code&gt; and &lt;code&gt;descent-override&lt;/code&gt; in a
&lt;code&gt;@font-face&lt;/code&gt; block for the local fallback family. This genuinely reduces shift, but the
numbers have to be measured per font pair, I have not measured mine, and copying someone
else&amp;rsquo;s values is worse than skipping the step.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;web.dev&amp;rsquo;s font best-practices guide goes deeper on loading strategy.&lt;/p&gt;
&lt;h2 id="what-this-walkthrough-does-not-give-you"&gt;What this walkthrough does not give you&lt;/h2&gt;
&lt;p&gt;A number. There is no field data here: this site has never had the traffic for a Chrome UX
Report entry, and I removed its analytics, so there is no measured before-and-after on this
page and I am not going to manufacture one.&lt;/p&gt;
&lt;p&gt;What the four checks give you is a list of file facts (a byte weight, a &lt;code&gt;loading&lt;/code&gt; attribute,
an absent &lt;code&gt;srcset&lt;/code&gt;, a timestamp in a template), each with the command that produces it. Those
are cheap to gather and hard to argue with, which makes them a good place to start. They are
not a substitute for measurement. If you want a measured number for your own site, use CrUX
for field data and Lighthouse for lab data, and when the two disagree, believe the field.&lt;/p&gt;</content:encoded></item><item><title>I removed Google Analytics rather than implement consent mode, and the tag is still live</title><link>https://therezaali.com/writing/consent-mode-v2-eu-sites/</link><guid isPermaLink="true">https://therezaali.com/writing/consent-mode-v2-eu-sites/</guid><pubDate>Wed, 05 Aug 2026 00:00:00 +0200</pubDate><author>contact@therezaali.com (Mo RezaAli)</author><category>reference</category><description>What consent mode v2 requires of a one-person EU site, what running it properly costs, and the least flattering part of my own reasoning.</description><content:encoded>&lt;p&gt;I took the Google Analytics tag out of this build rather than implement consent
mode on it. What follows is the whole of that reasoning, in the order I worked
through it, ending with the least flattering part: as I write, the tag is still
live on the deployed site.&lt;/p&gt;
&lt;p&gt;Start from the consequence, because the consequence is the part nobody argues
about. Send Google no consent signal for a visitor in the EEA and that visitor
stops appearing in the audiences your linked advertising products use. Not fined.
Subtracted.&lt;/p&gt;
&lt;p&gt;That has been the arrangement since early March 2024, Google documents it in
plain words, and it changes what kind of question this is. Consent mode gets
written about as a compliance feature. It behaves like a data-loss mechanism with
a legal trigger bolted on, which is why the decision a small site actually faces
(whether to run the tag at all) is not the same decision as configuring it
correctly.&lt;/p&gt;
&lt;h2 id="the-legal-shape-in-two-sentences"&gt;The legal shape, in two sentences&lt;/h2&gt;
&lt;p&gt;Two instruments apply and people routinely conflate them.&lt;/p&gt;
&lt;p&gt;The ePrivacy Directive governs putting things onto, or reading things off, a visitor&amp;rsquo;s device. Article 5(3) of the consolidated text says that storing information, or gaining access to information already stored, in a user&amp;rsquo;s terminal equipment &amp;ldquo;is only allowed on condition that the subscriber or user concerned has given his or her consent&amp;rdquo;, with a carve-out for whatever is strictly necessary to deliver the service the user asked for.&lt;/p&gt;
&lt;p&gt;Analytics is not strictly necessary to deliver a blog post.&lt;/p&gt;
&lt;p&gt;The EDPB&amp;rsquo;s Guidelines 2/2023 on the technical scope of Article 5(3), version 2.0, adopted 16 October 2024, exist largely because people read &amp;ldquo;cookie law&amp;rdquo; as &amp;ldquo;law about cookies.&amp;rdquo; The article is about storage and access on a device. The mechanism is irrelevant to it.&lt;/p&gt;
&lt;p&gt;The GDPR then governs what counts as consent, and what you may do with the personal data afterwards. Article 4(11) defines consent as a &amp;ldquo;freely given, specific, informed and unambiguous indication of the data subject&amp;rsquo;s wishes by which he or she, by a statement or by a clear affirmative action, signifies agreement.&amp;rdquo; Article 7(3) adds the right to withdraw at any time, and requires that the person be told about that right before they give consent in the first place. For what all of that means in practice, the reference text is the EDPB&amp;rsquo;s Guidelines 05/2020, adopted 4 May 2020.&lt;/p&gt;
&lt;p&gt;So: consent before the tag stores anything, and a real withdrawal path afterwards. Consent mode is Google&amp;rsquo;s mechanism for the first half.&lt;/p&gt;
&lt;h2 id="what-consent-mode-v2-actually-is"&gt;What consent mode v2 actually is&lt;/h2&gt;
&lt;p&gt;Consent mode is an API in &lt;code&gt;gtag.js&lt;/code&gt; and Google Tag Manager. Your consent banner tells Google&amp;rsquo;s tags what the visitor agreed to, and the tags change behaviour accordingly. Google&amp;rsquo;s tag platform documentation lists seven parameters. Four matter for advertising and measurement:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;ad_storage&lt;/code&gt;: &amp;ldquo;Enables storage, such as cookies (web) or device identifiers (apps), related to advertising.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;analytics_storage&lt;/code&gt;: the same, &amp;ldquo;related to analytics, for example, visit duration.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;ad_user_data&lt;/code&gt;: &amp;ldquo;Sets consent for sending user data to Google for online advertising purposes.&amp;rdquo;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;ad_personalization&lt;/code&gt;: &amp;ldquo;Sets consent for personalized advertising.&amp;rdquo;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The last two are the &amp;ldquo;v2&amp;rdquo; part. Google&amp;rsquo;s Ads Help page on EEA traffic states the requirement plainly: to keep using the tags and SDKs for measurement, ad personalization and remarketing, you must collect consent from end users based in the EEA and share those signals with Google. The Analytics-side page is blunter about the consequence of doing nothing. As of early March 2024, if you take no action, only end users &lt;em&gt;outside&lt;/em&gt; the EEA are included in audiences used by your linked advertising products.&lt;/p&gt;
&lt;p&gt;That is the mechanism behind the sentence this article opened with, quoted from the party that operates it.&lt;/p&gt;
&lt;h2 id="basic-versus-advanced"&gt;Basic versus advanced&lt;/h2&gt;
&lt;p&gt;Google&amp;rsquo;s documentation distinguishes two implementations, and the difference is the entire compliance question.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Basic&lt;/strong&gt; mode prevents Google tags from loading at all until the visitor interacts with the banner. Nothing is transmitted to Google before that interaction. Google describes the modelling you get back as the &amp;ldquo;general model,&amp;rdquo; which is the less detailed one.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Advanced&lt;/strong&gt; mode loads Google tags when the page opens, sets default consent states of &lt;code&gt;denied&lt;/code&gt;, and sends cookieless pings until and unless the visitor grants consent. Google says this &amp;ldquo;enables improved modeling compared to the Basic one as it provides an advertiser-specific model.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Advanced mode is what most agencies install, because it recovers more data. It is also the mode where a tag runs before consent. Whether the cookieless ping is lawful under Article 5(3) is a genuinely contested question, and the honest answer for a one-person site is that you should not be the person testing it. Basic mode is the conservative reading and it is the one I would pick if I ran Google tags at all.&lt;/p&gt;
&lt;h2 id="the-correct-minimal-implementation"&gt;The correct minimal implementation&lt;/h2&gt;
&lt;p&gt;If you do run it, the ordering matters more than the syntax. The &lt;code&gt;default&lt;/code&gt; call has to execute before the Google tag loads, or the tag has already acted.&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-html" data-lang="html"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c"&gt;&amp;lt;!-- 1. Defaults. Must run BEFORE gtag.js is requested. --&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;script&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nb"&gt;window&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;dataLayer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;window&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;dataLayer&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="p"&gt;[];&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nx"&gt;gtag&lt;/span&gt;&lt;span class="p"&gt;(){&lt;/span&gt;&lt;span class="nx"&gt;dataLayer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;push&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;arguments&lt;/span&gt;&lt;span class="p"&gt;);}&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nx"&gt;gtag&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;consent&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;&amp;#39;default&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="s1"&gt;&amp;#39;ad_storage&amp;#39;&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;&amp;#39;denied&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="s1"&gt;&amp;#39;ad_user_data&amp;#39;&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;&amp;#39;denied&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="s1"&gt;&amp;#39;ad_personalization&amp;#39;&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;denied&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="s1"&gt;&amp;#39;analytics_storage&amp;#39;&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;&amp;#39;denied&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="s1"&gt;&amp;#39;wait_for_update&amp;#39;&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;500&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="p"&gt;});&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="nt"&gt;script&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c"&gt;&amp;lt;!-- 2. The Google tag. --&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;script&lt;/span&gt; &lt;span class="na"&gt;async&lt;/span&gt; &lt;span class="na"&gt;src&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;https://www.googletagmanager.com/gtag/js?id=TAG_ID&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&amp;lt;/&lt;/span&gt;&lt;span class="nt"&gt;script&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;script&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nx"&gt;gtag&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;js&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nb"&gt;Date&lt;/span&gt;&lt;span class="p"&gt;());&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nx"&gt;gtag&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;config&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;&amp;#39;TAG_ID&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="nt"&gt;script&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;&lt;code&gt;wait_for_update&lt;/code&gt; gives an asynchronous banner a window, in milliseconds, to answer before the tag proceeds on the defaults.&lt;/p&gt;
&lt;p&gt;Then, and only from inside the banner&amp;rsquo;s accept handler:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-js" data-lang="js"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nx"&gt;gtag&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;consent&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="s1"&gt;&amp;#39;update&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="s1"&gt;&amp;#39;ad_storage&amp;#39;&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;&amp;#39;granted&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="s1"&gt;&amp;#39;ad_user_data&amp;#39;&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;&amp;#39;granted&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="s1"&gt;&amp;#39;ad_personalization&amp;#39;&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;granted&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="s1"&gt;&amp;#39;analytics_storage&amp;#39;&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="s1"&gt;&amp;#39;granted&amp;#39;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The &lt;code&gt;default&lt;/code&gt; call also takes a &lt;code&gt;region&lt;/code&gt; array of ISO 3166-2 codes, so you can deny by default in the EEA and behave differently elsewhere. I would not bother. One default, denied everywhere, is fewer branches to get wrong.&lt;/p&gt;
&lt;h2 id="the-part-that-is-not-code"&gt;The part that is not code&lt;/h2&gt;
&lt;p&gt;Here is why I stopped. Google&amp;rsquo;s own EU user consent policy (the policy your account is bound by, separate from the tag documentation) requires that you obtain consent for &amp;ldquo;the use of cookies or other local storage where legally required&amp;rdquo; and for &amp;ldquo;the collection, sharing, and use of personal data for personalization of ads.&amp;rdquo; It then requires you to &amp;ldquo;retain records of consent given by end users,&amp;rdquo; to &amp;ldquo;provide end users with clear instructions for revocation of consent,&amp;rdquo; and to &amp;ldquo;clearly identify each party that may collect, receive, or use end users&amp;rsquo; personal data.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Read that as an engineering backlog for a personal site:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;A banner with a reject button no harder to press than accept.&lt;/li&gt;
&lt;li&gt;A consent record store: what was consented to, when, under which banner version, retained and retrievable.&lt;/li&gt;
&lt;li&gt;A withdrawal UI reachable from every page, plus the logic to actually revoke and delete.&lt;/li&gt;
&lt;li&gt;A privacy policy naming Google as a recipient, listing the purposes, and stating retention.&lt;/li&gt;
&lt;li&gt;Ongoing maintenance of all four as the policies change.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Every item is real work, and every item is a place a solo operator ships a bug that nobody reviews. The value I was getting in exchange was a pageview count. That trade is bad, and it stays bad no matter how well I implement it.&lt;/p&gt;
&lt;h2 id="cookieless-alternatives-as-documented-by-their-vendors"&gt;Cookieless alternatives, as documented by their vendors&lt;/h2&gt;
&lt;p&gt;I checked these against the vendors&amp;rsquo; own pages rather than comparison posts.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cloudflare Web Analytics.&lt;/strong&gt; Two documents, five and a half years apart. The launch post of 29 September 2020 says &amp;ldquo;We don&amp;rsquo;t use any client-side state, like cookies or localStorage, for the purposes of tracking users,&amp;rdquo; and, immediately after, &amp;ldquo;we don&amp;rsquo;t &amp;lsquo;fingerprint&amp;rsquo; individuals via their IP address, User Agent string, or any other data for the purpose of displaying analytics.&amp;rdquo; The current documentation for the beacon that actually executes in your visitor&amp;rsquo;s browser, last updated 16 April 2026, is narrower and therefore more useful to check against: &amp;ldquo;The RUM beacon script does not store any data in the browser or access any storage data, such as cookies, localStorage, sessionStorage, IP address, or IndexedDB.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Two things to know before you reach for it. Cloudflare&amp;rsquo;s &amp;ldquo;visit&amp;rdquo; is defined as &amp;ldquo;a successful page view that has an HTTP referer that doesn&amp;rsquo;t match the hostname of the request,&amp;rdquo; which is not a session and should not be read as one. And on the free plan the beacon is switched on for you with Europe carved out. The docs say &amp;ldquo;Free customers have RUM enabled automatically, with EU traffic excluded, and can switch it off if they prefer.&amp;rdquo; For a site whose readers are mostly in the EEA, that is most of the audience missing by default.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Plausible.&lt;/strong&gt; The data policy states &amp;ldquo;We don&amp;rsquo;t use cookies, we don&amp;rsquo;t generate persistent identifiers and we don&amp;rsquo;t collect or store personal data that can be used to identify individuals,&amp;rdquo; and &amp;ldquo;Raw IP addresses and User-Agent data are never stored.&amp;rdquo; Daily uniques are counted with &lt;code&gt;hash(daily_salt + website_domain + ip_address + user_agent)&lt;/code&gt;, with the salt rotating every 24 hours. Plausible states you do not need cookie banners for analytics with it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Fathom.&lt;/strong&gt; The data page states &amp;ldquo;Our technology doesn&amp;rsquo;t use cookies, so you won&amp;rsquo;t need an annoying cookie consent banner,&amp;rdquo; and that raw IP addresses are stored only for security purposes and are not part of customer data exports. Fathom describes EU Isolation, launched in late 2021 following the Schrems II ruling, under which EU visitors&amp;rsquo; IP addresses do not touch US-controlled infrastructure.&lt;/p&gt;
&lt;p&gt;Some caveats I am not going to paper over. Dropping cookies removes the Article 5(3) storage-and-access trigger, and that is all it does; whether processing an IP address to derive a country needs its own lawful basis under the GDPR is a separate question that none of the paragraphs above answer. Everything in them is also a company describing itself, quoted by me, which is evidence of what a vendor has committed to in public and is not a legal opinion about whether the commitment is sufficient.&lt;/p&gt;
&lt;p&gt;And none of the three gives you Google Ads conversion tracking. If you run paid acquisition, consent mode is not optional for you and the conclusion of this article does not transfer.&lt;/p&gt;
&lt;h2 id="the-decision-and-what-has-not-happened-yet"&gt;The decision, and what has not happened yet&lt;/h2&gt;
&lt;p&gt;I can be specific about what getting this wrong costs because I got it wrong
here, at length, and (at the moment I am writing this) am still getting it
wrong.&lt;/p&gt;
&lt;p&gt;This domain runs a Google Analytics 4 tag. Property &lt;code&gt;G-XMEW7KSTQD&lt;/code&gt;, declared in
&lt;code&gt;config.yaml&lt;/code&gt;, served from a host in Italy to visitors who are mostly in Europe,
and it has been running for months with no consent banner, no consent mode
implementation and no privacy policy behind it. Checking took one command:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;curl -s https://therezaali.com/ &lt;span class="p"&gt;|&lt;/span&gt; grep -c -E &lt;span class="s1"&gt;&amp;#39;consent|denied&amp;#39;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# 0&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Zero. Not a misconfigured banner, and not a banner defaulting to granted.
Nothing at all. The tag loads, sets its cookie and reports, and has done every
day for months. I have spent working hours advising other people on measurement
setups while my own site fails at the first line of the rule.&lt;/p&gt;
&lt;p&gt;So I took the tag out of the build. &lt;code&gt;G-XMEW7KSTQD&lt;/code&gt; is gone from &lt;code&gt;config.yaml&lt;/code&gt; and from every template in this rebuild, which ships no analytics of any kind: nothing setting a cookie, nothing writing to &lt;code&gt;localStorage&lt;/code&gt;, no third-party script at all. The privacy page states what is collected and by whom.&lt;/p&gt;
&lt;p&gt;Now the part it would be easy to leave out.&lt;/p&gt;
&lt;p&gt;This article is written on a branch that has not been deployed. At the time of writing, the live &lt;code&gt;therezaali.com&lt;/code&gt; still serves the GA4 tag, still with no consent banner, exactly as described in the section above. So do not take my word for either half of that:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;curl -s https://therezaali.com/ &lt;span class="p"&gt;|&lt;/span&gt; grep -c -iE &lt;span class="s1"&gt;&amp;#39;gtag|googletagmanager&amp;#39;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Running it while writing this paragraph printed &lt;code&gt;4&lt;/code&gt;. &lt;code&gt;grep -c -E 'consent|denied'&lt;/code&gt; against the same page still printed &lt;code&gt;0&lt;/code&gt;. When this rebuild ships, the first command starts printing &lt;code&gt;0&lt;/code&gt;; if it ever goes back to printing anything else, something has regressed and I would like to be told. Either way the date lands on the &lt;a href="https://therezaali.com/corrections/"&gt;corrections page&lt;/a&gt;, which is where the changes to this domain are logged rather than narrated.&lt;/p&gt;
&lt;p&gt;Writing &amp;ldquo;I removed Google Analytics&amp;rdquo; in the present tense, on a branch, about a site still serving the tag, is the specific kind of small forward-dated lie that a reader can catch with one command. On a page arguing that claims should be checkable, getting caught on that would be deserved.&lt;/p&gt;
&lt;p&gt;The reason to publish the failure instead of quietly fixing it is that the failure is the only part of this I can currently prove.&lt;/p&gt;</content:encoded></item><item><title>hreflang for sites with three pages and no translation team</title><link>https://therezaali.com/writing/hreflang-small-multilingual-sites/</link><guid isPermaLink="true">https://therezaali.com/writing/hreflang-small-multilingual-sites/</guid><pubDate>Wed, 05 Aug 2026 00:00:00 +0200</pubDate><author>contact@therezaali.com (Mo RezaAli)</author><category>reference</category><description>The return-link rule, x-default, and the code formats that fail silently. Then the measured state of the 2,183 machine-translated pages I had pointed them at.</description><content:encoded>&lt;p&gt;Two versions of one Italian headline. This is the one that got published:&lt;/p&gt;
&lt;pre tabindex="0"&gt;&lt;code&gt;Come Ottimizzare Le Campagne Google Ads
&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;This is the one an Italian would have written:&lt;/p&gt;
&lt;pre tabindex="0"&gt;&lt;code&gt;Come ottimizzare le campagne Google Ads
&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;The difference is &lt;code&gt;Le&lt;/code&gt;. It is a definite article, capitalised in the middle of a sentence, and Treccani&amp;rsquo;s grammar entry on capital usage rules it out without qualification: with the titles of a book, a work of art, a film, a song, &amp;ldquo;la maiuscola si limita alla prima parola del titolo&amp;rdquo; (the capital is limited to the first word of the title). The Accademia della Crusca&amp;rsquo;s standing answer on capitals, from Luca Serianni and Giovanni Nencioni in &lt;em&gt;La Crusca per voi&lt;/em&gt; n. 2, April 1991, treats the wider question as genuinely unsettled in places, mostly around the boundary between proper and common nouns. There is no boundary here to be near. Nothing in Italian licenses a capital &lt;code&gt;Le&lt;/code&gt; mid-title, and an Italian reader clocks it before finishing the line.&lt;/p&gt;
&lt;p&gt;I know how many there were, because I counted them.&lt;/p&gt;
&lt;p&gt;That headline came out of a machine-translation pass on this domain,
&lt;a href="https://therezaali.com/writing/the-bot-that-published-900-articles/"&gt;since deleted&lt;/a&gt;. Of the 728 Italian files it produced, 228 carried at least one Italian article, preposition or conjunction capitalised somewhere other than the first word. Thirty-one per cent. Here is how to get that number, against &lt;code&gt;$B&lt;/code&gt;, the commit immediately before the deletion:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nv"&gt;B&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;22b7c22d^
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;git archive &lt;span class="nv"&gt;$B&lt;/span&gt; content/post/insights/ &lt;span class="p"&gt;|&lt;/span&gt; tar -x -C /tmp/corpus
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nb"&gt;cd&lt;/span&gt; /tmp/corpus/content/post/insights
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;for&lt;/span&gt; f in *.it.md&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; awk &lt;span class="s1"&gt;&amp;#39;/^title:/{sub(/^title: */,&amp;#34;&amp;#34;); gsub(/^&amp;#34;|&amp;#34;$/,&amp;#34;&amp;#34;); print; exit}&amp;#39;&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;&lt;/span&gt;&lt;span class="nv"&gt;$f&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;done&lt;/span&gt; &amp;gt; /tmp/it-titles.txt
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;wc -l &amp;lt; /tmp/it-titles.txt &lt;span class="c1"&gt;# 728&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;python3 -c &lt;span class="s1"&gt;&amp;#39;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s1"&gt;import re, sys
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s1"&gt;fw = set(&amp;#34;&amp;#34;&amp;#34;il lo la i gli le un uno una di a da in con su per tra fra
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s1"&gt;del dello della dei degli delle al allo alla ai agli alle dal dalla dallo dai dagli dalle
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s1"&gt;nel nello nella nei negli nelle sul sullo sulla sui sugli sulle col coi
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s1"&gt;e ed o od che ma se non ne si come quando mentre&amp;#34;&amp;#34;&amp;#34;.split())
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s1"&gt;t = [l.strip() for l in sys.stdin if l.strip()]
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s1"&gt;print(sum(1 for x in t if any(w[0].isupper() and w.lower() in fw
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s1"&gt; for w in re.findall(r&amp;#34;[A-Za-zÀ-ÿ&amp;#39;&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;&amp;#39;&amp;#34;&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;]+&amp;#34;, x)[1:])), &amp;#34;of&amp;#34;, len(t))
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s1"&gt;&amp;#39;&lt;/span&gt; &amp;lt; /tmp/it-titles.txt &lt;span class="c1"&gt;# 228 of 728&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;The test is deliberately narrow. It fires only on closed-class words: articles, prepositions, conjunctions, the words whose capitalisation is not a matter of taste in Italian. Nothing a proper noun or a brand name does can trip it, and 228 is therefore a floor rather than an estimate, which is the direction I would rather be wrong in when the number is about my own work.&lt;/p&gt;
&lt;p&gt;An earlier draft of this page said 57%. Trying to reproduce it, I ran four different definitions of Title Case over the same 728 titles (every non-initial word capitalised, most of them capitalised, at least two of them capitalised, and the function-word test above), and not one landed near 57. So the 57 is gone. What stands is the number that has a command attached to it.&lt;/p&gt;
&lt;p&gt;Onto all 2,183 of those translated pages (728 Italian, 728 Arabic, 727 Chinese) I shipped hreflang, correctly.&lt;/p&gt;
&lt;h2 id="what-correct-looked-like"&gt;What correct looked like&lt;/h2&gt;
&lt;p&gt;The &lt;code&gt;&amp;lt;link rel=&amp;quot;alternate&amp;quot;&amp;gt;&lt;/code&gt; sets came out of the template, which is why they were clean. Every version listed every other version and itself. The codes were valid. &lt;code&gt;x-default&lt;/code&gt; was present. The return links resolved. Audit the markup and it passes.&lt;/p&gt;
&lt;p&gt;The markup was never the problem. The markup worked, and working is what did the damage: hreflang exists to route a reader in a given language to the URL in that language, so it took the headline above and delivered it, efficiently, to the people best equipped to see what was wrong with it.&lt;/p&gt;
&lt;h2 id="the-mechanics-from-googles-documentation"&gt;The mechanics, from Google&amp;rsquo;s documentation&lt;/h2&gt;
&lt;p&gt;Google&amp;rsquo;s page on telling it about localized versions is the primary source and it is short. Five rules carry most of the weight.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Return links are mandatory.&lt;/strong&gt; &amp;ldquo;If two pages don&amp;rsquo;t both point to each other, the tags will be ignored. This is so that someone on another site can&amp;rsquo;t arbitrarily create a tag naming itself as an alternative version of one of your pages.&amp;rdquo; One-directional annotations are discarded outright. They do not degrade to a hint.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Every version lists itself.&lt;/strong&gt; &amp;ldquo;Each language version must list itself as well as all other language versions.&amp;rdquo; The set of links is identical on every variant of the page, which is the property that makes the whole thing template-generatable.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;&lt;code&gt;x-default&lt;/code&gt; is the fallback.&lt;/strong&gt; &amp;ldquo;The reserved &lt;code&gt;x-default&lt;/code&gt; value is used when no other language/region matches the user&amp;rsquo;s browser setting.&amp;rdquo; Google recommends it for language selectors and auto-redirecting home pages.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The code is language first, region optional.&lt;/strong&gt; First code, the language, ISO 639-1. Optional second code after a hyphen, the region, ISO 3166-1 Alpha 2. Only codes listed in those two standards work; Google&amp;rsquo;s page names &lt;code&gt;es-419&lt;/code&gt; as one that does not. Then the warning that catches almost everyone: &amp;ldquo;You can&amp;rsquo;t specify the country code by itself. The first code stands for the language and Google doesn&amp;rsquo;t automatically derive the language from a country code.&amp;rdquo; &lt;code&gt;be&lt;/code&gt; is Belarusian. Belgium is &lt;code&gt;de-be&lt;/code&gt;, &lt;code&gt;nl-be&lt;/code&gt;, &lt;code&gt;fr-be&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Three delivery methods, pick one.&lt;/strong&gt; Head tags, an HTTP &lt;code&gt;Link:&lt;/code&gt; header, or a sitemap.&lt;/p&gt;
&lt;h3 id="head-tags"&gt;Head tags&lt;/h3&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-html" data-lang="html"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;head&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;title&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;Widgets, Inc&lt;span class="p"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="nt"&gt;title&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;link&lt;/span&gt; &lt;span class="na"&gt;rel&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;alternate&amp;#34;&lt;/span&gt; &lt;span class="na"&gt;hreflang&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;en-gb&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="na"&gt;href&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;https://en-gb.example.com/page.html&amp;#34;&lt;/span&gt; &lt;span class="p"&gt;/&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;link&lt;/span&gt; &lt;span class="na"&gt;rel&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;alternate&amp;#34;&lt;/span&gt; &lt;span class="na"&gt;hreflang&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;en-us&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="na"&gt;href&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;https://en-us.example.com/page.html&amp;#34;&lt;/span&gt; &lt;span class="p"&gt;/&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;link&lt;/span&gt; &lt;span class="na"&gt;rel&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;alternate&amp;#34;&lt;/span&gt; &lt;span class="na"&gt;hreflang&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;en&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="na"&gt;href&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;https://en.example.com/page.html&amp;#34;&lt;/span&gt; &lt;span class="p"&gt;/&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;link&lt;/span&gt; &lt;span class="na"&gt;rel&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;alternate&amp;#34;&lt;/span&gt; &lt;span class="na"&gt;hreflang&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;de&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="na"&gt;href&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;https://de.example.com/page.html&amp;#34;&lt;/span&gt; &lt;span class="p"&gt;/&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="p"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nt"&gt;link&lt;/span&gt; &lt;span class="na"&gt;rel&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;alternate&amp;#34;&lt;/span&gt; &lt;span class="na"&gt;hreflang&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;x-default&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="na"&gt;href&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;https://www.example.com/&amp;#34;&lt;/span&gt; &lt;span class="p"&gt;/&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;&amp;lt;/&lt;/span&gt;&lt;span class="nt"&gt;head&lt;/span&gt;&lt;span class="p"&gt;&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Cheapest to implement. Most expensive to serve. With &lt;em&gt;n&lt;/em&gt; locales you carry &lt;em&gt;n&lt;/em&gt; links in the head of every page, so across the site the head grows as the square of your locale count.&lt;/p&gt;
&lt;h3 id="http-header"&gt;HTTP header&lt;/h3&gt;
&lt;p&gt;Useful for non-HTML files. The header returned is identical for every version:&lt;/p&gt;
&lt;pre tabindex="0"&gt;&lt;code&gt;Link: &amp;lt;https://example.com/file.pdf&amp;gt;; rel=&amp;#34;alternate&amp;#34;; hreflang=&amp;#34;en&amp;#34;,
 &amp;lt;https://de-ch.example.com/file.pdf&amp;gt;; rel=&amp;#34;alternate&amp;#34;; hreflang=&amp;#34;de-ch&amp;#34;,
 &amp;lt;https://de.example.com/file.pdf&amp;gt;; rel=&amp;#34;alternate&amp;#34;; hreflang=&amp;#34;de&amp;#34;
&lt;/code&gt;&lt;/pre&gt;&lt;h3 id="sitemap"&gt;Sitemap&lt;/h3&gt;
&lt;p&gt;Moves the weight out of the pages entirely. Declare the XHTML namespace, then give every &lt;code&gt;&amp;lt;url&amp;gt;&lt;/code&gt; a full set of &lt;code&gt;&amp;lt;xhtml:link&amp;gt;&lt;/code&gt; children including itself:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-xml" data-lang="xml"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nt"&gt;&amp;lt;urlset&lt;/span&gt; &lt;span class="na"&gt;xmlns=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;http://www.sitemaps.org/schemas/sitemap/0.9&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="na"&gt;xmlns:xhtml=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;http://www.w3.org/1999/xhtml&amp;#34;&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nt"&gt;&amp;lt;url&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nt"&gt;&amp;lt;loc&amp;gt;&lt;/span&gt;https://www.example.com/english/page.html&lt;span class="nt"&gt;&amp;lt;/loc&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nt"&gt;&amp;lt;xhtml:link&lt;/span&gt; &lt;span class="na"&gt;rel=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;alternate&amp;#34;&lt;/span&gt; &lt;span class="na"&gt;hreflang=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;de&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="na"&gt;href=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;https://www.example.de/deutsch/page.html&amp;#34;&lt;/span&gt;&lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nt"&gt;&amp;lt;xhtml:link&lt;/span&gt; &lt;span class="na"&gt;rel=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;alternate&amp;#34;&lt;/span&gt; &lt;span class="na"&gt;hreflang=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;de-ch&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="na"&gt;href=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;https://www.example.de/schweiz-deutsch/page.html&amp;#34;&lt;/span&gt;&lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nt"&gt;&amp;lt;xhtml:link&lt;/span&gt; &lt;span class="na"&gt;rel=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;alternate&amp;#34;&lt;/span&gt; &lt;span class="na"&gt;hreflang=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;en&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="na"&gt;href=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;https://www.example.com/english/page.html&amp;#34;&lt;/span&gt;&lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nt"&gt;&amp;lt;/url&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nt"&gt;&amp;lt;url&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nt"&gt;&amp;lt;loc&amp;gt;&lt;/span&gt;https://www.example.de/schweiz-deutsch/page.html&lt;span class="nt"&gt;&amp;lt;/loc&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nt"&gt;&amp;lt;xhtml:link&lt;/span&gt; &lt;span class="na"&gt;rel=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;alternate&amp;#34;&lt;/span&gt; &lt;span class="na"&gt;hreflang=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;de&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="na"&gt;href=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;https://www.example.de/deutsch/page.html&amp;#34;&lt;/span&gt;&lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nt"&gt;&amp;lt;xhtml:link&lt;/span&gt; &lt;span class="na"&gt;rel=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;alternate&amp;#34;&lt;/span&gt; &lt;span class="na"&gt;hreflang=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;de-ch&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="na"&gt;href=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;https://www.example.de/schweiz-deutsch/page.html&amp;#34;&lt;/span&gt;&lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nt"&gt;&amp;lt;xhtml:link&lt;/span&gt; &lt;span class="na"&gt;rel=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;alternate&amp;#34;&lt;/span&gt; &lt;span class="na"&gt;hreflang=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;en&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="na"&gt;href=&lt;/span&gt;&lt;span class="s"&gt;&amp;#34;https://www.example.com/english/page.html&amp;#34;&lt;/span&gt;&lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nt"&gt;&amp;lt;/url&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nt"&gt;&amp;lt;/urlset&amp;gt;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Three versions, three &lt;code&gt;&amp;lt;url&amp;gt;&lt;/code&gt; entries, three identical children each. It is &lt;em&gt;n&lt;/em&gt;² either way. The sitemap simply moves the entries somewhere a static host does not re-send on every request.&lt;/p&gt;
&lt;p&gt;One more line from the documentation that people miss once a site grows past a few locales: &amp;ldquo;If it becomes difficult to maintain a complete set of bidirectional links for every language, you can omit some languages on some pages; Google will still process the ones that point to each other.&amp;rdquo; What you should not omit is the link from a newly added language back to your dominant one.&lt;/p&gt;
&lt;p&gt;Separately, and unrelated to hreflang: set &lt;code&gt;lang&lt;/code&gt; on the &lt;code&gt;&amp;lt;html&amp;gt;&lt;/code&gt; element. That is the declaration browsers and assistive technology use, W3C&amp;rsquo;s internationalization guidance covers it, and the tag syntax is RFC 5646. hreflang tells a search engine which URL to show whom. &lt;code&gt;lang&lt;/code&gt; tells a screen reader which voice to use. Different jobs, and you need both.&lt;/p&gt;
&lt;h2 id="underneath-the-markup"&gt;Underneath the markup&lt;/h2&gt;
&lt;p&gt;&lt;strong&gt;Italian.&lt;/strong&gt; The 228 titles above came from one mechanical cause: the pipeline ran an English-language title-casing function over output that was no longer English. The transform was locale-blind. It fired on every title in every language, and on Italian it produced a visible error in any title containing an article, a preposition or a conjunction past the first word.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Arabic.&lt;/strong&gt; Here I have to correct myself, and the correction is large. An earlier version of this page claimed 116 Arabic files shipped a verb cut in half with an English stem left inside it. I went to count it and the count did not come back anywhere near 116. Every token in the 728 Arabic files where a lowercase Latin run is fused directly to Arabic letters, with the number of files each appears in:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nb"&gt;cd&lt;/span&gt; /tmp/corpus/content/post/insights
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;python3 -c &lt;span class="s1"&gt;&amp;#39;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s1"&gt;import re, glob, collections
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s1"&gt;AR = r&amp;#34;ء-يٱ-ۓ&amp;#34;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s1"&gt;tok = re.compile(f&amp;#34;[A-Za-z{AR}]*(?:[{AR}][a-z]|[a-z][{AR}])[A-Za-z{AR}]*&amp;#34;)
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s1"&gt;n = collections.Counter()
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s1"&gt;for f in sorted(glob.glob(&amp;#34;*.ar.md&amp;#34;)):
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s1"&gt; for t in set(tok.findall(open(f, encoding=&amp;#34;utf-8&amp;#34;).read())):
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s1"&gt; n[t] += 1
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s1"&gt;for t, c in n.most_common():
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s1"&gt; print(c, t)
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="s1"&gt;&amp;#39;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;pre tabindex="0"&gt;&lt;code&gt;23 وtur
2 لتautomate
1 تتblur
1 وb
1 السيمantics
1 يautomates
1 treatingها
&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;That is the whole population. &lt;code&gt;لتautomate&lt;/code&gt;, &lt;code&gt;يautomates&lt;/code&gt;, &lt;code&gt;تتblur&lt;/code&gt; and &lt;code&gt;treatingها&lt;/code&gt; are the corruption I described: an Arabic verbal prefix or object suffix welded to a bare English stem, because a substitution pass treated a templatic language as if it were space-delimited Latin words. Arabic builds a verb by interleaving a root with affixes. The pass split one open and left English sitting in the gap. &lt;code&gt;السيمantics&lt;/code&gt; is the same fault on a noun. Six distinct files, not 116.&lt;/p&gt;
&lt;p&gt;The top row is a different bug and it is the common one. Twenty-three files carry a Markdown link whose label was truncated mid-word and published as a fragment: &lt;code&gt;…وtur](https://blog.hubspot.com/topic/data-analytics)&lt;/code&gt;. It rendered as a fragment, on twenty-three pages, for months. Nobody read it. &lt;code&gt;وb&lt;/code&gt; I cannot classify and am not going to guess at.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Chinese.&lt;/strong&gt; I have no basis to characterise the Chinese quality in either direction. I do not read Chinese and I never had it reviewed, which is itself the finding.&lt;/p&gt;
&lt;h2 id="why-the-multiplier-goes-both-ways"&gt;Why the multiplier goes both ways&lt;/h2&gt;
&lt;p&gt;Take the general shape. hreflang multiplies whatever your per-locale quality is by the number of locales you serve. Positive quality, more reach. Negative quality, more reach.&lt;/p&gt;
&lt;p&gt;Which means that everything I did right (the self-referencing sets, the valid codes, the &lt;code&gt;x-default&lt;/code&gt; fallback, the bidirectional return links that a Search Console audit would have waved through without a note) worked precisely as designed, and what it was designed to do was take a headline no Italian would write, put it in front of Italian readers, in Italy, at the moment they were looking for exactly that topic, four locales wide, from 31 October 2025 to 27 April 2026.&lt;/p&gt;
&lt;p&gt;Google&amp;rsquo;s spam policies name the pattern under scaled content abuse: &amp;ldquo;Scraping feeds, search results, or other content to generate many pages (including through automated transformations like synonymizing, translating, or other obfuscation techniques), where little value is provided to users.&amp;rdquo; Translating is listed as an automated transformation not because translation is suspect, but because translating at volume without review turns one source into many thin pages. As a description of what I built, that is accurate, and the hreflang was the delivery mechanism.&lt;/p&gt;
&lt;h2 id="what-i-would-tell-a-solo-operator"&gt;What I would tell a solo operator&lt;/h2&gt;
&lt;p&gt;Pick one non-English locale, or zero.&lt;/p&gt;
&lt;p&gt;Zero is respectable, and it is what this site&amp;rsquo;s articles use now. If you pick one, pick the language you can read. I live in Messina. I can check Italian output against how people around me actually write, which is precisely the check I skipped at 2,183-file scale and could have done in an afternoon at three-file scale. Then:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Translate a page only when you have read the translation end to end.&lt;/li&gt;
&lt;li&gt;Never run English text transforms on non-English output: title casing, truncation at a character count, possessive handling. Gate every one of them on locale. The 228 titles and the truncated Arabic link labels are the same bug wearing two hats.&lt;/li&gt;
&lt;li&gt;Add the hreflang set only after the translation passes review. hreflang last, not first.&lt;/li&gt;
&lt;li&gt;Keep the set self-referencing and bidirectional, add &lt;code&gt;x-default&lt;/code&gt;, and let the template generate it so it cannot drift.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Three pages in two languages that a human has read will outperform two thousand in four. I am not offering that as a consolation prize. It is the only configuration in which the mechanics above are worth implementing at all.&lt;/p&gt;</content:encoded></item><item><title>A gate that reads only the artifact scores the surface, and mine passed 842 articles</title><link>https://therezaali.com/writing/the-bot-that-published-900-articles/</link><guid isPermaLink="true">https://therezaali.com/writing/the-bot-that-published-900-articles/</guid><pubDate>Wed, 05 Aug 2026 00:00:00 +0200</pubDate><author>contact@therezaali.com (Mo RezaAli)</author><category>postmortem</category><description>LLM-as-judge rubrics and keyword briefs carry the same defect. The source of a gate that had it, and seven checks that find it in yours.</description><content:encoded>&lt;p&gt;An automated evaluator that scores a generated artifact using only features it can
read inside that artifact does not measure the property it names. It measures the
surface of the property, and anything generating against it will produce the
surface. Mine required each article to carry at least two named entities and at
least one number. It passed 842 of them. 470 of the 739 English articles carried a
before-and-after metrics table, and not one figure in any of them describes
something that happened.&lt;/p&gt;
&lt;p&gt;The URL of this page rounds. The precise figure is &lt;strong&gt;842 commits titled &lt;code&gt;Add post:&lt;/code&gt;&lt;/strong&gt;, and it belongs in the first paragraph rather than a footnote, because a
slug that sands 842 up to 900 is exactly the sort of small convenience this whole
article is about. 739 of those articles were live in English when I deleted the
archive. Most were then machine-translated into Italian, Arabic and Chinese: 728,
728 and 727 files. That is how 739 English articles became 2,922 files on disk.&lt;/p&gt;
&lt;p&gt;Nothing about that defect is specific to content bots, and you do not need to have
run one to have it in a gate you own. It is the same defect as an LLM-as-judge
rubric that rewards well-formed citations, a keyword-count SEO brief, an
engineering scorecard that counts test files rather than running them, and a
hiring filter that counts years of a word in a document. Every one of those scores
a text, and a language model is very cheap at producing text. What this page has
that an argument about it would not is the source of a gate that had the defect,
quoted from the file, and seven checks that find it in yours.&lt;/p&gt;
&lt;p&gt;I wrote that gate. It lives in one file, 2,177 lines of JavaScript, and after the
deletion I read all of them again, which is why the references at the foot of this
page are line numbers rather than recollection, and why the validator further down
is quoted instead of described. That is the qualification behind everything below.
Not that I avoided this failure. That I have the source of a system which
manufactured citations 842 times, and have been back through it since.&lt;/p&gt;
&lt;p&gt;Nobody read any of them before they went out. Me included.&lt;/p&gt;
&lt;p&gt;Three dates, kept apart on purpose. The bot published from 31 October 2025 to
27 April 2026. It then went quiet, and the archive sat there, indexed, for
another three months. I deleted it on 5 August 2026.&lt;/p&gt;
&lt;p&gt;I wrote the bot. I set the cron to hourly and &lt;code&gt;AUTO_PR&lt;/code&gt; to &lt;code&gt;&amp;quot;false&amp;quot;&lt;/code&gt;, which is
the setting that made every commit go straight to the published branch. This is
not a piece about a model that misbehaved, and &amp;ldquo;AI hallucinates&amp;rdquo; explains nothing
about why this particular failure took the particular shape it took. The model was
handed three instructions: write 900 words, include at least one number, invent no
numbers. The only material it was given was 450 characters of RSS &lt;code&gt;description&lt;/code&gt;.
Only two of the three could be satisfied at once, and the gate I wrote decided
which two.&lt;/p&gt;
&lt;h2 id="the-system-and-what-i-can-show-you-of-it"&gt;The system, and what I can show you of it&lt;/h2&gt;
&lt;p&gt;The 2,177 lines are a Cloudflare Worker. The commits and the code sit in a private
repository, so on those you have my word; every external source on this page is
linked and checkable, which is the split the &lt;code&gt;limits&lt;/code&gt; line at the top states
outright.&lt;/p&gt;
&lt;p&gt;The configuration is four lines long and tells you most of the story:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-toml" data-lang="toml"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;triggers&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nx"&gt;crons&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;0 */1 * * *&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;vars&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nx"&gt;RSS_URLS&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;https://neilpatel.com/blog/feed/, https://blog.hubspot.com/marketing/rss.xml, https://feeds.feedburner.com/mitsmr, https://www.forrester.com/blogs/feed/, https://www.beautyofsaas.com/feed&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nx"&gt;OPENAI_MODEL&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;gpt-4o-mini&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nx"&gt;GITHUB_BRANCH&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;main&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nx"&gt;GITHUB_PR_BASE&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;main&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nx"&gt;AUTO_PR&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;false&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nx"&gt;MAX_ITEMS&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;5&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nx"&gt;STRICT_MODE&lt;/span&gt; &lt;span class="p"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;true&amp;#34;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Once an hour it pulled five marketing feeds (Neil Patel, the HubSpot Marketing
Blog, MIT Sloan Management Review, Forrester&amp;rsquo;s blogs, and Beauty of SaaS), took
up to five new items, sent each to a model, and wrote the result into &lt;code&gt;main&lt;/code&gt; with a
GitHub contents &lt;code&gt;PUT&lt;/code&gt;. There was no pull request. There could not have been one:
the PR function opens with &lt;code&gt;if (config.branch === config.prBase) return;&lt;/code&gt;, and I
had set both to &lt;code&gt;main&lt;/code&gt;. Every article was committed with &lt;code&gt;draft: false&lt;/code&gt; and
&lt;code&gt;author: &amp;quot;Mo Reza Ali&amp;quot;&lt;/code&gt;. That is not even how I spell my name.&lt;/p&gt;
&lt;p&gt;The reasoning at the time was ordinary enough. I wanted a publishing habit I
could not skip, I wanted the domain to have something on it, and I had read
enough about content velocity to talk myself into believing that shipping often
mattered more than shipping well, with the quality gate holding the floor
underneath. That was the deal I made with myself. Nearly all of my attention
after that went into the gate.&lt;/p&gt;
&lt;p&gt;Here is what the deal produced. The repository held 2,967 commits when I deleted
it. 842 of them said &lt;code&gt;Add post:&lt;/code&gt;. 63 said &lt;code&gt;Update post:&lt;/code&gt;, 845 auto-assigned a
cover image, 220 auto-translated, 160 auto-healed metadata. That is 2,130 commits
no human wrote and no human read. I was tuning the instrument and never once
reading the dial.&lt;/p&gt;
&lt;h2 id="the-contradiction-inside-the-gate"&gt;The contradiction inside the gate&lt;/h2&gt;
&lt;p&gt;This is the part that generalises, so it is worth being precise about.&lt;/p&gt;
&lt;p&gt;The only input the model ever received about a story was this, from the prompt
builder:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-js" data-lang="js"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nx"&gt;rss_summary&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;stripHtml&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;item&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;description&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="s2"&gt;&amp;#34;&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nx"&gt;slice&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;450&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;450 characters. That is roughly the length of this paragraph. From that, the
prompt asked for a 600–900 word &amp;ldquo;insider briefing&amp;rdquo;, and then the validator
decided whether the result was allowed to ship. Two of its checks did the damage:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-js" data-lang="js"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;// 2. Named entities — require at least 2 proper nouns (companies, products, people)
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;&lt;/span&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;entities&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;size&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nx"&gt;errors&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;push&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sb"&gt;`only &lt;/span&gt;&lt;span class="si"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;entities&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;size&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sb"&gt; named entities found (need &amp;gt;= 2): ...`&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;// 3. Specific numbers — require at least 1 concrete data point
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;&lt;/span&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;numbers&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nx"&gt;errors&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;push&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;no specific numbers or data points found (need &amp;gt;= 1)&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Read those two rules next to the 450-character input and the conflict is
structural. A 450-character excerpt very often contains no statistic at all. The
validator refused to pass any article without one. The prompt, in the same
breath, said: &amp;ldquo;NEVER fabricate statistics, percentages, or study results.&amp;rdquo; So the
model was given a requirement it could satisfy only by breaking a prohibition. It
resolved the contradiction the only way the contradiction could be resolved.&lt;/p&gt;
&lt;p&gt;And when it could not, the article shipped anyway:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-js" data-lang="js"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;config&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;strictMode&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;processed&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;validation&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;ok&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;attempts&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="nx"&gt;AI_REPAIR_ATTEMPTS&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;warn&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;strict_mode_fallback&amp;#34;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;title&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;item&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;title&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;errors&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;processed&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;validation&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;errors&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nx"&gt;processed&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;validation&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;warnings&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="p"&gt;...(&lt;/span&gt;&lt;span class="nx"&gt;processed&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;validation&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;warnings&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="p"&gt;[]),&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="p"&gt;...&lt;/span&gt;&lt;span class="nx"&gt;processed&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;validation&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;errors&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;map&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;e&lt;/span&gt; &lt;span class="p"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="sb"&gt;`[relaxed] &lt;/span&gt;&lt;span class="si"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;e&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sb"&gt;`&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="p"&gt;];&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="nx"&gt;processed&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;validation&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;ok&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;processed&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="p"&gt;}&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt; &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nb"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sb"&gt;`Validation failed: &lt;/span&gt;&lt;span class="si"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;processed&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;validation&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;errors&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;&amp;#34;; &amp;#34;&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sb"&gt;`&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;&lt;code&gt;AI_REPAIR_ATTEMPTS&lt;/code&gt; was 2. After two failed retries, every error was rewritten
as a warning with a &lt;code&gt;[relaxed]&lt;/code&gt; prefix, &lt;code&gt;validation.ok&lt;/code&gt; was set to &lt;code&gt;true&lt;/code&gt; by
hand, and the article was published. &lt;code&gt;STRICT_MODE = &amp;quot;true&amp;quot;&lt;/code&gt; in the config. It did
nothing except delay publication by two API calls.&lt;/p&gt;
&lt;p&gt;You can measure how little the gate held. The prompt shipped a kill list of 114
banned phrases and told the model that articles containing them &amp;ldquo;will be
REJECTED by our validator&amp;rdquo;. 686 of the 739 published articles (92.8 percent)
contain at least one phrase from that list. 604 of them, 81.7 percent, contain a
word from the hard-reject subset that was supposed to fail an article outright.
The word &amp;ldquo;crucial&amp;rdquo; is in 486 published articles. &amp;ldquo;Essential&amp;rdquo; is in 424. &amp;ldquo;Foster&amp;rdquo;
is in 368. The list was real, the enforcement was real, and it made no
difference to what went out the door, because the escape hatch ran after it.&lt;/p&gt;
&lt;h2 id="what-a-check-for-the-presence-of-digits-produced"&gt;What a check for the presence of digits produced&lt;/h2&gt;
&lt;p&gt;Here is a block from one published article, unedited:&lt;/p&gt;
&lt;pre tabindex="0"&gt;&lt;code&gt;## What Good Looks Like in Numbers

| Metric | Before | After | Change |
|-------------------|--------|-------|----------|
| Conversion Rate | 2% | 5% | +150% |
| Retention Rate | 60% | 75% | +25% |
| Time-to-Value | 30 days| 15 days| -50% |

Source: HubSpot Blog
&lt;/code&gt;&lt;/pre&gt;&lt;p&gt;That exact &lt;code&gt;Retention Rate | 60% | 75% | +25%&lt;/code&gt; row appears in 29 of the 739
articles, under sixteen different attributions: four crediting &lt;code&gt;Source: HubSpot Blog&lt;/code&gt;, two &lt;code&gt;Source: Forrester Research&lt;/code&gt;, nine naming HubSpot in some other
phrasing, two naming &lt;code&gt;Neil Patel&lt;/code&gt;, and ten pointing at an internal source that
does not exist: &lt;code&gt;Internal Company Data&lt;/code&gt;, &lt;code&gt;Internal Marketing Analysis&lt;/code&gt;,
&lt;code&gt;Internal Marketing Metrics 2026&lt;/code&gt;. &lt;code&gt;Conversion Rate | 2% | 5% | +150%&lt;/code&gt; appears
126 times. &lt;code&gt;Time-to-Value | 6 months | 3 months | -50%&lt;/code&gt;, 134 times. 470 of the
739 English articles carried a before-and-after table of this kind, and not one
figure in any of them describes something that happened.&lt;/p&gt;
&lt;p&gt;The attribution lines are worse than the tables. 341 articles carry a line
beginning &lt;code&gt;Source:&lt;/code&gt;. Six of those 341 contain a URL.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;180 of the 739 English posts cited a source called &amp;ldquo;Internal Analysis&amp;rdquo; that
does not exist.&lt;/strong&gt; That is the figure, with its denominator. And because the
argument of this article is that a number ought to be re-derivable by a reader
who does not trust the person who wrote it, here is the command rather than a
description of the command. &lt;code&gt;$B&lt;/code&gt; is the commit immediately before the deletion:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="nv"&gt;B&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;22b7c22d^
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# denominator: the English article files&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;git ls-tree -r --name-only &lt;span class="nv"&gt;$B&lt;/span&gt; content/post/insights/ &lt;span class="se"&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="se"&gt;&lt;/span&gt; &lt;span class="p"&gt;|&lt;/span&gt; grep &lt;span class="s1"&gt;&amp;#39;\.md$&amp;#39;&lt;/span&gt; &lt;span class="p"&gt;|&lt;/span&gt; grep -v &lt;span class="s1"&gt;&amp;#39;\.\(it\|ar\|zh-cn\)\.md$&amp;#39;&lt;/span&gt; &lt;span class="p"&gt;|&lt;/span&gt; wc -l
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# 739&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# numerator: pattern Source:.*Internal, case-insensitive, each file counted once&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;git grep -lEi &lt;span class="s1"&gt;&amp;#39;Source:.*Internal&amp;#39;&lt;/span&gt; &lt;span class="nv"&gt;$B&lt;/span&gt; -- &lt;span class="s1"&gt;&amp;#39;content/post/insights/*.md&amp;#39;&lt;/span&gt; &lt;span class="se"&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="se"&gt;&lt;/span&gt; &lt;span class="p"&gt;|&lt;/span&gt; grep -v &lt;span class="s1"&gt;&amp;#39;\.\(it\|ar\|zh-cn\)\.md$&amp;#39;&lt;/span&gt; &lt;span class="p"&gt;|&lt;/span&gt; wc -l
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# 180&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Run it against this repository and you get 180. You do not have this repository.
It is private, which the &lt;code&gt;limits&lt;/code&gt; line at the top of the page says in as many
words. What you actually have is the pattern, the denominator, and the commit,
which is more than the old site ever gave anyone.&lt;/p&gt;
&lt;p&gt;There is no internal analysis, and no internal company data. I have never run the
study. Some of the other attribution lines are worse still, because they put
invented figures in the mouths of institutions that never said them: &lt;code&gt;Bank of England Internal Metrics&lt;/code&gt;, &lt;code&gt;ABN AMRO Internal Metrics&lt;/code&gt;, &lt;code&gt;Anthropic Case Studies&lt;/code&gt;,
&lt;code&gt;Cloudflare's internal metrics post-rewrite&lt;/code&gt;.&lt;/p&gt;
&lt;p&gt;Even the fabrications disagree with each other. The same before-and-after pair,
70 percent to 85 percent, is labelled &lt;code&gt;+15%&lt;/code&gt; in one article and &lt;code&gt;+21%&lt;/code&gt; in
another. &lt;code&gt;2.5%&lt;/code&gt; to &lt;code&gt;4.0%&lt;/code&gt; is labelled &lt;code&gt;+60%&lt;/code&gt; in one and &lt;code&gt;+1.5%&lt;/code&gt; in another. A
gate that checked arithmetic would have caught this. Mine checked for the
presence of digits.&lt;/p&gt;
&lt;h2 id="what-it-cost-other-people"&gt;What it cost other people&lt;/h2&gt;
&lt;p&gt;The site&amp;rsquo;s credibility, first, which is why it was deleted rather than corrected.
The larger cost is what it did with other people&amp;rsquo;s work.&lt;/p&gt;
&lt;p&gt;31 of the published article titles contain the word &amp;ldquo;Forrester&amp;rdquo;. They are not
articles about Forrester. They are Forrester&amp;rsquo;s own headlines, republished under
my byline: &lt;code&gt;Announcing The Forrester Wave™: Digital Experience Platforms, Q4 2025&lt;/code&gt;. &lt;code&gt;Call For Entries: Forrester B2B Summit North America 2026 Awards&lt;/code&gt;.
&lt;code&gt;Meet Jess Lloyd: Forrester's New Principal Analyst Covering Consumer&lt;/code&gt;. A staff
announcement about somebody&amp;rsquo;s new job at somebody else&amp;rsquo;s company, with my name
in the author field.&lt;/p&gt;
&lt;p&gt;52 articles have a &lt;code&gt;description&lt;/code&gt; written in the first person by a person who is
not me, lifted straight out of the feed excerpt. One of them reads: &lt;em&gt;&amp;ldquo;In the
first episode of my Nudge podcast, I interviewed the fantastic psychologist Dr.&amp;rdquo;&lt;/em&gt;
That sentence belongs to Phill Agnew, who writes for the HubSpot Marketing Blog
and hosts Nudge. It is the opening line of his post about the psychology of
music, and it is truncated mid-name because my code sliced the description at 100
characters. My site published his memory of his own podcast episode as my
description, under my byline, above 900 words of invented metrics.&lt;/p&gt;
&lt;p&gt;Not one of the 739 articles carried a link back to the item it was derived from.
The frontmatter schema the bot wrote has thirteen fields. &lt;code&gt;source_url&lt;/code&gt; is not one
of them. The pipeline knew the origin URL (it is right there in the prompt
payload as &lt;code&gt;source_url&lt;/code&gt;) and dropped it before publishing. That was not an
oversight in the model. That was a field I did not add.&lt;/p&gt;
&lt;p&gt;Google&amp;rsquo;s spam policy has a name for the shape of this. &amp;ldquo;Scaled content abuse is
when many pages are generated for the primary purpose of manipulating search
rankings and not helping users,&amp;rdquo; and the policy is explicit that it applies no
matter how the pages were created. I would have argued at the time that mine was
different, because mine had a quality gate. Its two load-bearing checks are
quoted above. They read a table of invented metrics and marked it a pass.&lt;/p&gt;
&lt;h2 id="what-i-deleted"&gt;What I deleted&lt;/h2&gt;
&lt;p&gt;On 5 August 2026 I removed all of it in two commits. &lt;code&gt;content/post/insights/&lt;/code&gt;:
2,922 files. All Arabic and Chinese translations, including six Arabic files that
had shipped a verb split open with a raw English stem left in the gap, and a
further twenty-three carrying a link label truncated mid-word. They stayed live for months
in a language I do not read. An earlier version of this paragraph put the first
figure at 116; I recounted it, and the token-by-token derivation is in
&lt;a href="https://therezaali.com/writing/hreflang-small-multilingual-sites/"&gt;the hreflang piece&lt;/a&gt;.
A comparison page publishing
&lt;code&gt;AggregateRating&lt;/code&gt; structured data that scored a real product at 1.36 out of 5
while every public review platform put it above 4.4, presented as independent,
and ranking a former employer without disclosing the relationship. A 2,074,882
byte &amp;ldquo;cover image&amp;rdquo; whose visible content is the word &lt;code&gt;LOADING..&lt;/code&gt;, served on 229
posts. Of the 739 English articles, 736 carried a disqualifying defect I could
name individually. The other three were entirely derived from someone else&amp;rsquo;s
post. The honest survivor count is zero.&lt;/p&gt;
&lt;p&gt;Deleting the source file does not undeploy a Worker. The cron keeps firing and the
token keeps working until you remove the Worker in the Cloudflare dashboard and
revoke the GitHub personal access token by hand, which is why the commit message
for the pipeline removal ends with a line to myself in capital letters:
&lt;code&gt;MANUAL STEP REQUIRED&lt;/code&gt;. If you are reading this because you built something
similar, that is the step people forget.&lt;/p&gt;
&lt;h2 id="the-same-defect-in-four-other-systems"&gt;The same defect in four other systems&lt;/h2&gt;
&lt;p&gt;The claim at the top of this page is not about content bots. Stated in full:
&lt;strong&gt;a quality gate which cannot distinguish &amp;ldquo;sourced&amp;rdquo; from &amp;ldquo;present&amp;rdquo; will
manufacture whatever shape it is asked to measure.&lt;/strong&gt; My validator could see that
a paragraph contained a percent sign. It could not see whether the percentage
referred to anything. So it did not enforce evidence. It enforced the
&lt;em&gt;appearance&lt;/em&gt; of evidence. Then I pointed a generative system at it and told it
to score well. The output was not a failure of the system. It was the system
working, on the objective I actually gave it.&lt;/p&gt;
&lt;p&gt;This is Campbell&amp;rsquo;s law with a shorter feedback loop. The moment a proxy becomes
the target, the cheapest way to move the proxy is to fake the thing it stands
for, and a language model is very cheap at producing the surface of things. But
the point is not confined to content bots. Any automated evaluator that scores a
generated artifact on detectable surface features has this defect: keyword-count
SEO briefs, LLM-as-judge rubrics that reward citation &lt;em&gt;formatting&lt;/em&gt;, engineering
scorecards that count tests rather than run them, hiring filters that count
years. If your gate measures a shadow, you will get very good shadows.&lt;/p&gt;
&lt;p&gt;Volume is the symptom and never the cause. What scales the damage is how long the
loop runs unread, and mine ran hourly for six months. Tightening the rules inside
that window changed nothing that mattered, because the defect was never how strict
the rules were. It was that every one of them could be satisfied by the text under
test.&lt;/p&gt;
&lt;h2 id="how-to-detect-it-in-your-own-pipeline"&gt;How to detect it in your own pipeline&lt;/h2&gt;
&lt;p&gt;Seven checks, in the order I would run them, on any pipeline where a generator is
scored by an evaluator. Each of them is an afternoon at most, and the first is ten
minutes.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;1. Unplug the network and run the gate.&lt;/strong&gt; If every check returns the same
verdict offline, nothing in the gate resolves against the world, and every rule
in it is a property of the text. Mine passed offline. That one test in November
would have told me what I found out in August.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;2. Find every write to the pass flag.&lt;/strong&gt; Search the pipeline for each assignment
to whatever variable means &amp;ldquo;this passed&amp;rdquo;. Any write outside the function that
computes it is an escape hatch, and a gate with an escape hatch is advisory. Mine
is quoted above: &lt;code&gt;processed.validation.ok = true&lt;/code&gt;, five lines after the errors
were renamed warnings.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;3. Test the ban list against shipped output, not against the prompt.&lt;/strong&gt; I had 114
banned phrases and a sentence telling the model that articles containing them
&amp;ldquo;will be REJECTED by our validator&amp;rdquo;. 686 of the 739 that shipped contained at
least one. That measurement is a single grep over the output directory. Any rule
you have never run against what actually went out is a rule you are assuming
works.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;4. Sort the corpus by repeated line and read the top of the list.&lt;/strong&gt; Fabricated
evidence clusters, because a model asked for a plausible number reaches for the
same plausible numbers. One row, &lt;code&gt;Retention Rate | 60% | 75% | +25%&lt;/code&gt;, appears in
29 of my 739 articles under sixteen different attributions. Independent
measurements of different companies do not agree to the digit.&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-bash" data-lang="bash"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;# the most repeated table rows across a generated corpus&lt;/span&gt;
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;grep -rhE &lt;span class="s1"&gt;&amp;#39;^\|&amp;#39;&lt;/span&gt; content/ &lt;span class="p"&gt;|&lt;/span&gt; sed &lt;span class="s1"&gt;&amp;#39;s/[[:space:]]\{1,\}/ /g&amp;#39;&lt;/span&gt; &lt;span class="se"&gt;\
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="se"&gt;&lt;/span&gt; &lt;span class="p"&gt;|&lt;/span&gt; sort &lt;span class="p"&gt;|&lt;/span&gt; uniq -c &lt;span class="p"&gt;|&lt;/span&gt; sort -rn &lt;span class="p"&gt;|&lt;/span&gt; head -20
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;&lt;strong&gt;5. Check the corpus against itself before you check it against the world.&lt;/strong&gt;
Internal contradiction costs nothing to detect and needs no external source: the
disagreeing percentage labels above were sitting in my own output the whole time.
Arithmetic that does not close is fabrication that did not coordinate.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;6. Count the outputs carrying a pointer a machine can follow.&lt;/strong&gt; Not a citation,
a resolvable pointer. 341 of my articles carried a line beginning &lt;code&gt;Source:&lt;/code&gt;. Six
of those 341 contained a URL. The ratio between those two numbers is the entire
finding of this article, and getting it takes two greps.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;7. Write one deliberately false artifact by hand and run it through the gate.&lt;/strong&gt;
If it passes, you are done and you have your answer in an afternoon rather than
six months. Mine would have passed a table of pure invention, because it did, 470
times.&lt;/p&gt;
&lt;h2 id="how-to-build-a-gate-a-generator-cannot-satisfy-by-writing-more"&gt;How to build a gate a generator cannot satisfy by writing more&lt;/h2&gt;
&lt;p&gt;One rule covers all of it. &lt;strong&gt;Every check must be able to fail for a reason that
does not exist inside the artifact being checked.&lt;/strong&gt; Something outside has to
answer: a server, a file on disk after a build, the exit code of a process that
actually ran, a row in a database, a person who picks up the phone. If you cannot
name the external thing a check resolves against, you have a text property with a
serious-sounding name.&lt;/p&gt;
&lt;p&gt;Applied, in four places the same gate shape keeps appearing:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Content.&lt;/strong&gt; Do not check whether an article cites a source. That is a string test
and a generator passes it for free, which is how 341 articles carried a &lt;code&gt;Source:&lt;/code&gt;
line and six carried a URL. Check that the URL resolves, then that the fetched
document contains the figure being claimed. The first needs a regex. The second
needs a server.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;LLM-as-judge.&lt;/strong&gt; A rubric that rewards well-formed citations scores citation
formatting, and the model under evaluation will supply well-formed citations. Put
the retrieved document in front of the judge alongside the candidate answer and
score whether the claim is supported by that document. A judge shown only the
answer can only score the answer&amp;rsquo;s surface, which is my defect one layer up and
with better vocabulary.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Engineering scorecards.&lt;/strong&gt; A dashboard that counts test files counts test files.
Run them and read the exit code. Line coverage is a milder version of the same
mistake, because it records which lines executed rather than which assertions
were made, so a suite that asserts nothing can still score well on it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Hiring screens.&lt;/strong&gt; A filter counting years of a keyword in a document scores the
document, and documents are cheap to write and getting cheaper. Resolve outside
it: a work sample scored blind, a reference who answers.&lt;/p&gt;
&lt;p&gt;The strongest form of the rule is to delete the thing the gate was guarding.
Nothing that publishes. That is the whole design change I made, and it costs me
nothing I valued: the useful half of the bot read five feeds and told me what was
new, and the half that destroyed the site was one &lt;code&gt;PUT&lt;/code&gt;. Drop the five links into
a queue I open over coffee and you keep everything worth keeping, minus the
GitHub token, the write scope and the byline.&lt;/p&gt;
&lt;p&gt;The correction is never a stricter version of the same gate. I tried that. One of
my own commits in that window is literally titled &lt;code&gt;fix(insightbot): promote 30+ AI words to hard-reject quality gate&lt;/code&gt;, and 686 articles containing banned words
shipped anyway. Tightening a surface check buys a better forgery. Check the URL,
not the word &amp;ldquo;Source:&amp;rdquo;.&lt;/p&gt;
&lt;h2 id="what-each-check-on-this-site-resolves-against"&gt;What each check on this site resolves against&lt;/h2&gt;
&lt;p&gt;The checker in this repository resolves every citation on the site against the
live internet on every build, and a dead link stops the deployment. Precisely
what that means:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Every article under &lt;code&gt;content/writing/&lt;/code&gt; must carry a non-empty &lt;code&gt;sources&lt;/code&gt; array
and a &lt;code&gt;limits&lt;/code&gt; sentence in its frontmatter, and every source entry must have a
title, a publisher, a URL and an accessed date. Missing any one of those fails
the build.&lt;/li&gt;
&lt;li&gt;Each source URL is fetched on every build: &lt;code&gt;HEAD&lt;/code&gt; first, &lt;code&gt;GET&lt;/code&gt; when a server
refuses &lt;code&gt;HEAD&lt;/code&gt;, two retries with backoff on transport failures and 5xx, six
requests in flight, twenty second timeout. The response must be 2xx or 3xx.
The count grows with every article, so it is not written down here: an earlier
version of this sentence said 142 URLs across ten articles and was wrong within
a fortnight, while the homepage carried the current figure. One site, one fact,
two numbers, which is the thing this whole page is about. The checker prints
the real count on every run and the homepage reads it from the content rather
than from my memory.&lt;/li&gt;
&lt;li&gt;It runs before Hugo. If a cited document died since the last deploy, nothing
renders at all, and the failure happens before a single page exists.&lt;/li&gt;
&lt;li&gt;It runs a second time after the build against &lt;code&gt;public/&lt;/code&gt;, checking that every
site-root link written in &lt;code&gt;content/&lt;/code&gt; resolves to a file that exists, and that
every URL in the sitemap exists and is not caught by a redirect rule. A sitemap
advertising a URL the edge redirects away is its own small lie.&lt;/li&gt;
&lt;li&gt;It is invoked from &lt;code&gt;scripts/build.sh&lt;/code&gt;, which is the command Cloudflare Pages is
configured to run, so the checker&amp;rsquo;s exit code is the deployment&amp;rsquo;s exit code.
Non-zero and nothing ships. That last link is a dashboard setting rather than a
file, which makes it the one part of the chain a repository cannot enforce, and
it is written down as such at the top of that script.&lt;/li&gt;
&lt;li&gt;401, 403 and 429 are downgraded to warnings in the deploy path, because a
publisher refusing an unfamiliar user agent from a datacenter address is not a
dead link. It cannot hide a fabricated citation: an invented URL fails DNS or
returns 404, never 403.&lt;/li&gt;
&lt;li&gt;Zero dependencies. A check that guards against unreviewed automation should not
itself install several hundred unreviewed packages.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Two further checks exist because of a failure found in the rebuild rather than in
the bot: one canonical spelling of my name, and one fixed value per fact, so a
figure an earlier draft got wrong becomes unwritable everywhere on the site.&lt;/p&gt;
&lt;p&gt;The property that matters is not thoroughness. It is that every one of those
checks is satisfied only by something outside the file being checked. There is
nothing in it I can argue past by writing more confidently, which is precisely
what my validator could be argued past by doing. I have already had to obey it
twice while writing these pages.&lt;/p&gt;
&lt;p&gt;If I ever did want machine help with a draft again, the gate would run backwards
from the one I built: extract every factual claim, demand a URL for each, fetch
the URL, and fail the build when it does not answer 200 or does not contain the
figure being claimed. Harder to write. Much cheaper to trust, because no amount
of fluent prose moves it. Only a server answering does.&lt;/p&gt;
&lt;p&gt;I publish no statistics about my own work. I have not collected any, and the
previous version of this domain is a 2,922-file demonstration of what happens
when I let a machine fill that particular silence. Where I do not know something,
the article says so. The &lt;code&gt;limits&lt;/code&gt; line at the top of this one admits that the
repository I keep quoting is private, so my external citations are checkable and
my commits are not.&lt;/p&gt;
&lt;p&gt;The full rules are on the &lt;a href="https://therezaali.com/editorial-policy/"&gt;editorial policy&lt;/a&gt; page, along with
what I do when I get something wrong.&lt;/p&gt;
&lt;p&gt;One rule, in the same words both times it appears here. &lt;strong&gt;Every check must be able
to fail for a reason that does not exist inside the artifact being checked.&lt;/strong&gt; You
can settle that about your own gate this afternoon: unplug the network and run it,
and see whether a single check changes its verdict. Mine did not change one. The
four systems at the top of this page are the ones I would put through it first,
because every one of them scores a text, and a text is the cheapest artifact a
machine now makes.&lt;/p&gt;
&lt;h2 id="primary-evidence"&gt;Primary evidence&lt;/h2&gt;
&lt;p&gt;Everything above comes from artifacts in the repository behind this site rather
than from memory. For my own record, and so the numbers can be re-derived:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Worker source &lt;code&gt;extracted-projects/rezaali-insightbot/mo-content-worker.js&lt;/code&gt;,
2,177 lines by &lt;code&gt;wc -l&lt;/code&gt;: &lt;code&gt;AI_REPAIR_ATTEMPTS&lt;/code&gt; at line 11, the banned-phrase list
at lines 26–103, the &lt;code&gt;[relaxed]&lt;/code&gt; fallback at lines 686–697, the frontmatter
builder with &lt;code&gt;draft: false&lt;/code&gt; and the byline at lines 714–735, the 450-character
slice at line 827, the entity and number checks at lines 1440–1454, the
unreachable pull-request function at line 1945.&lt;/li&gt;
&lt;li&gt;Configuration: &lt;code&gt;extracted-projects/rezaali-insightbot/wrangler.toml&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Commit &lt;code&gt;fa6312b3&lt;/code&gt; removed the pipeline; commit &lt;code&gt;22b7c22d&lt;/code&gt; removed the content,
and its message contains the file-by-file accounting I have summarised here.&lt;/li&gt;
&lt;li&gt;Corpus counts were taken at &lt;code&gt;22b7c22d^&lt;/code&gt;, the commit immediately before the
deletion, over the 739 English files under &lt;code&gt;content/post/insights/&lt;/code&gt;. The
internal-source figure of 180 is the file count for
&lt;code&gt;git grep -lEi 'Source:.*Internal'&lt;/code&gt; at that commit, restricted to the English
files. The command is printed in full earlier on this page.&lt;/li&gt;
&lt;li&gt;The checker described above is &lt;code&gt;scripts/check-sources.mjs&lt;/code&gt; in the repository
behind this site, and the deploy gate that invokes it is &lt;code&gt;scripts/build.sh&lt;/code&gt;.
The citation and URL counts are that script&amp;rsquo;s own output on the build that
produced this page.&lt;/li&gt;
&lt;/ul&gt;</content:encoded></item></channel></rss>