04 October 2026
๐ Another Sunday night in Gourock, Popular Monster in the headphones, depression spikes, music helps. The last note here is from the end of May. Four months of silence - not because nothing happened, but because too much did. The day job took everything it could take, and what was left went into moving an Italian collective blog off shared hosting and onto machines I now look after. Both are starting to settle. So: back to writing. Or something like it.
La Bottega del Barbieri is an Italian collective blog, online since 2001, with something like thirty thousand articles in its archive. On the night of 9 September it moved from Aruba shared hosting to a small Hetzner box. The whole story is on the blog itself, in Italian, in two parts (one, two). This note is the part that didn't fit there: what the bots did when the door opened, and what stands at the door now.
It's not a guide. It's an account of what happened, in the order it happened, mistakes included. The rules below were written against one swarm, on one site, over a few weeks of reading logs at night; they say what worked here, not what will work anywhere else. Sensitive values stay out. The swarm reads blogs too.
the polite one
On Aruba, every client that couldn't run JavaScript met a challenge page before the real one. It broke link previews for years and nobody on our side could switch it off. It also, as it turned out, kept the crawlers out.
The DNS switch happened a little after one in the morning, local time. By two, the request graph was at ten, fifteen, twenty per second, on a blog whose readers are asleep at that hour. One query to the nginx JSON log gave the name: ClaudeBot. One IP, 6,667 requests between 02:03 and 02:18. GPTBot second with 974. Somewhere in the noise, one Italian Firefox. A real reader, awake.
Twenty-five years of archive had become readable an hour earlier. The crawlers noticed before I'd finished checking the certificate.
The answer took less time than the query: a physical robots.txt in the document root, disallowing every declared AI crawler I could list, search engines left alone. Then I watched. Reconstructed tonight from the rotated log, times in UTC:
# zcat access.json-20260929.gz \
| jq -r 'select(.time|startswith("2026-09-09T01:")) | select(.user_agent|test("ClaudeBot"))
| if .path=="/robots.txt" then "ROBOTS \(.time) \(.status)" else .time[0:16] end' \
| uniq -c
...
716 2026-09-09T01:19
554 2026-09-09T01:20
1 ROBOTS 2026-09-09T01:20:46+00:00 200
164 2026-09-09T01:20
474 2026-09-09T01:21
1 ROBOTS 2026-09-09T01:25:02+00:00 200
1 ROBOTS 2026-09-09T01:39:53+00:00 200
uniq -c only merges adjacent lines, so the robots fetch splits minute 01:20 into a before and an after. That's the whole story in one accident of formatting. It read the file at 01:20:46 - twenty-seven seconds after my own curl had confirmed it was there. It drained 638 requests already in flight. Then nothing. It came back twice, at 01:25 and at 01:39, to read the file again. No pages.
robots.txt is a sign on the door, not a lock. It works on whoever chooses to read signs. That night, one crawler chose to.
the impolite ones
That afternoon the visitor counter claimed numbers I didn't believe, so I read the raw log instead. 5,265 distinct IPs presenting as ordinary browsers. 4,593 of them asked for exactly one page and never came back. The top user agents were a handful of current devices - macOS 15, iOS 18.4, Pixel 9, Galaxy S25 - each spread across hundreds of addresses.
No neighbourhood on earth has four hundred people with the same phone on the same build, each reading one page of an Italian blog and vanishing. This was a rotating residential proxy pool with user-agent rotation. By the evening: 25,009 IPs and 261,751 requests in 24 hours.
It never declared itself, so nothing in robots.txt was addressed to it - and nothing suggests it would have listened. Blocking by IP was pointless, since every IP was used once. Blocking by user agent was pointless, since its user agents were our readers'. The server didn't even notice - the page cache absorbed everything. Cost wasn't the problem. The problem was the authors' texts walking out one page at a time, and a counter that had stopped meaning anything.
a door, not a wall
Anubis puts a small proof-of-work in front of the site. The browser solves it once, gets a cookie valid for a week, and goes through. For a person that's one second a week. For a scraper that rotates IPs and never keeps the cookie, it's one second per page. Verified search engines, link preview fetchers and feed readers go through on an exceptions list.
The uncomfortable part: that is exactly what Aruba's anti-bot did, the thing I'd spent a year cursing. The difference isn't the tool. It's who writes the list, from which logs, and who gets to see what breaks. A wall you hold the key to, and can look through, is a door. The same wall built by someone who won't tell you what it blocks is something else.
Wiring it was one field in the proxy: the WordPress host now forwards to Anubis instead of nginx, and the rollback is the same field. The first test battery failed on everything - Anubis wants the client address from the proxy in front of it, and my curls weren't sending it. Second round, headers set: browsers challenged, Mastodon and Telegram and Facebook through, feeds through, a fake Googlebot challenged, AI crawlers denied.
By 17:30 the next day: 35,119 challenges issued, 772 passed. 2.2%. Daily visitors in the counter went 700, 800, 400 - before, the day of the switch, the first full day after. Even the 400 turned out inflated when I reread the logs two weeks later, but the direction was right.
weights add up
Anubis rules either decide (ALLOW, DENY) or add weight (WEIGH), and the total weight picks a threshold:
thresholds:
- name: moderate-suspicion
expression:
all: [weight >= 10, weight < 20]
action: CHALLENGE
challenge: { algorithm: fast, difficulty: 4 }
- name: mild-proof-of-work
expression:
all: [weight >= 20, weight < 30]
action: CHALLENGE
challenge: { algorithm: fast, difficulty: 4 }
- name: extreme-suspicion
expression: weight >= 30
action: CHALLENGE
challenge: { algorithm: fast, difficulty: 6 }
For the first night, every browser matched two rules: one for generic browsers, +10, and a catch-all, +10. Weight 20. I didn't know weights accumulated. Look at the table: 10 and 20 land in two different bands with the same difficulty. Raising the difficulty, my first idea that night, would have changed nothing. The generic-browser rule went. The catch-all alone gives 10.
Then the hosting networks:
- name: unknown-catchall
user_agent_regex: .*
action: WEIGH
weight:
adjust: 10
- import: /data/datacenter.yaml
The imported file is a long list of prefixes announced by hosting providers, at +20. A browser coming from a datacenter sums to 30 and gets difficulty 6. Weight, not DENY, on purpose: VPN exits live in those same networks, and someone reading the Bottega through a VPN has their reasons.
fingerprints
The swarm doesn't sleep, but it has habits. The first one that gave it away: user agents claiming to be Chrome, sending an Accept-Language header the way Firefox formats it. No real browser does that.
- name: swarm-lang-mismatch-deny
action: DENY
expression:
all:
- '"Accept-Language" in headers'
- 'headers["Accept-Language"] == "<redacted>"'
- 'headers["User-Agent"].matches("<redacted>")'
The first line wasn't there in the first version. On a request without the header, the expression doesn't evaluate to false - it errors out on the missing key. The logs said so as soon as the rule went live. Guard every header lookup before you compare it. It's in every expression rule since.
More followed, one per habit. An exact language string with no browser condition at all. A browser-looking user agent with no language. Exact user agent and language pairs for a Linux Chrome sprayed across hundreds of IPs, one article each. Every rule keys on a behaviour. None on a country, none on an address range.
One of them caught me. The newsletter is built by a script on the other machine that reads the blog with headless Chromium, and headless Chromium sends an empty Accept-Language. The "browser with no language" rule denied it, and the extraction test came back with zero articles. The fix is an ALLOW for that one machine, placed above the rule that caught it:
- name: bottega-lab-allow
action: ALLOW
remote_addresses:
- <redacted>/32
Order matters: the first ALLOW or DENY that matches ends the evaluation. WEIGH rules keep adding, and the total only counts if nothing else has decided.
everything is 200
status_codes:
CHALLENGE: 200
DENY: 200
Which means the proxy log shows 200 for a real page, a challenge page and a deny page alike. Any count made from status codes is wrong. Anubis stamps every request it lets through with the rule that let it in, and with a PASS when that meant solving a challenge. nginx now writes both into its JSON log. That's the only place where "who actually got in" is written down.
The same blindness bit the monitor. The health check on the other machine hits the home page every five minutes. For over two weeks it got a challenge page, with a 200, and reported the site healthy. It was measuring Anubis. It only surfaced when, a few times a day, the challenge path answered 500 because the check didn't send Accept-Encoding: gzip - Anubis logs it as client was given a challenge but does not in fact support gzip compression. Two in a row, and an alert. The fix was one more ALLOW for the monitor. Working but not verified, again.
and then...
The swarm adapted within a day, every time. A new rule, a new face. So there's now a script that reads the log every hour, counts distinct IPs per user agent among the requests that passed, and pushes an alert when one agent spreads over too many addresses. It only looks. Writing the rule stays a human decision.
Which leaves the door in a state I'd file next to the others: challenged, not stopped. A part of the swarm runs JavaScript and pays the second. Anubis made reading the Bottega expensive, not impossible.
In May I wrote that running Invidious makes me a guest on someone else's web - reading YouTube's internal JSON, uninvited, because the alternative is being watched. This time I'm the house, and the guests don't knock. I don't think the two are the same thing. I also don't think the difference is as clean as I'd like it to be.
The album finished a while ago and I haven't put another on. Say what you like about Falling In Reverse, Ronnie Radke can sing. Clean, screamed, rapped, sometimes inside the same chorus, and it never sounds like three people taking turns. Have you heard the new one, Joseph? Out in mid-September, with Corey Taylor and Serj Tankian beside him - the first time those two have ever been on the same track - and he doesn't disappear between them. He holds the song.
It's gone past midnight. The door holds, mostly. Time to put Joseph on again.