When bots started hitting ArenaSim’s registration form, the expected move was obvious: add an off-the-shelf bot shield and move on.
I did not, for three reasons:
- It exposes the user to a third party. Everyone filling in my form has their browser connect to another company and hand it data.
- It does not work everywhere. In some countries the relevant domains are unreachable; users cannot submit the form and have no idea why.
- Image puzzles tire humans too. Picking out bicycles delays a bot three seconds and a person fifteen. The ratio keeps moving against the human.
I wrote my own. Five layers — and a normal user never knows the first four exist.
1. A signed token
The form page embeds a token: session, timestamp and a signature.
token = base64( session_id | time | random )
sig = HMAC-SHA256( token, server_secret )
On submission three conditions are checked: is the signature valid, is the token too old, has it been used before.
This layer removes every bot that posts directly without ever opening the form — and that is the large majority of the traffic.
2. The honeypot
There is a field on the form that is not visible. A human cannot see it; a bot that parses the form and fills every field does.
<div aria-hidden="true" class="hp">
<label>Company name</label>
<input name="company_name" tabindex="-1" autocomplete="off">
</div>
tabindex="-1" so a keyboard user never lands there by accident. autocomplete="off" so the browser’s autofill does not declare a real person a bot.
For a honeypot to work it must be impossible for a human to fill under any circumstance.
3. Proof of work
The third layer imposes a cost. The server issues a challenge; the client searches until it finds a hash beginning with a given number of zero bits.
On average 2¹⁸ ≈ 262,000 attempts. About 0.2 seconds in a modern browser — it finishes in the background while the user is typing.
| requests | total cost | |
|---|---|---|
| Normal user | 1 | 0.2 s |
| Bot | 10,000/hour | 33 min of pure CPU |
It does not stop a bot — it makes it expensive. In most spam operations the margin is thin enough that this cost is sufficient to move the target elsewhere.
4. Rate limiting
Classic but necessary: a sliding window per IP and per token. The subtlety is to slow an over-limit request rather than reject it. Rejection sends the bot looking for a new IP; slowing keeps it busy and teaches it nothing.
5. And only then: the image puzzle
If the first four layers become suspicious, the fifth appears. A puzzle I generate myself: my letters, my drawing, produced on the server with the answer kept there.
Not complex — it does not need to be. The traffic arriving here is already a small filtered remainder. The fifth layer’s job is not to stop 99% but to slow the handful that got past four.
The underlying idea
An off-the-shelf shield is one tall wall. Mine is five short ones.
Each layer is easy to clear alone. Clearing all five requires the bot to actually open the page, interpret the DOM correctly, run JavaScript, pay a computational cost and respect a rate limit.
By that point the bot has become a browser — and running a browser is hundreds of times more expensive than running a script.
The goal is not to make it impossible. The goal is to break the economics.
What I learned
In security there is no “solved / unsolved”; there is a cost curve. Anything that raises the attacker’s cost faster than the user’s works.
And the best protection is the one the protected never notices. In this system all a normal user sees is the form being submitted.