An online shop went live. Two complaints arrived the same week, and neither was a straightforward “this is broken”.
First: “the customer paid and it isn’t in the system”
An order had been paid but the panel still said “awaiting payment”. One order you fix by hand. But it was spreading.
I traced the payment flow:
payment provider → notification received ✓
application → notification processed ✓
application → order status written ✓
application → accounting record written ✗
The last step wrote to a table. That table did not exist.
A database migration had not been run during installation. With the table missing, that step threw, the error was caught and logged — and the flow continued. So the order became “paid” while the accounting side never came into being.
A missing table is an accident. The real bug is that a failed step did not stop the flow.
When a step like payment accounting fails, the system must not say “continue”. Either the whole transaction is rolled back, or the order is explicitly marked incomplete and put in front of a human.
Continuing silently turns a fault into a record-keeping error — and record-keeping errors are noticed months later.
I created the table, backfilled the missing records, and added this to the flow: if the accounting step fails, the order lands in a “needs review” queue.
Second: “the site never opens, but I can get in”
This is a bug that gives itself away in its own sentence.
The site would not open from outside. But the panel owner could get in. Nothing in the access log, nothing in the error log, server load normal.
The cause was the PHP-FPM pool.
The control panel generates a pool file per account and writes the same default into all of them: at most four concurrent requests. For a busy account that is far too low.
When the pool is full, requests are not refused — they wait in the socket queue and time out. From outside it looks like “the site is down”, but because the request never reaches the web server it appears in no log at all.
And the asymmetry: whoever is filling the pool holds the workers, so they can get in while anyone queueing behind them cannot.
“I can get in but nobody else can” is this bug’s signature.
How it was confirmed
The definitive evidence is in the pool’s own log:
# first find which PHP version it runs under
grep SetHandler /…/vhosts/<domain>.conf
# then read that version's pool log
WARNING: [pool <account>] server reached pm.max_children setting (4),
consider raising it
The pool limit was raised and the site returned to normal.
What the two share
Completely different layers. What they share: neither produced an error message.
| what happened | what it looked like | |
|---|---|---|
| Missing table | a step failed, the flow continued | a successful order |
| Pool exhausted | the request timed out in the queue | no log entry at all |
Silent failure has been the most frequent class of bug I have met this year. In command-line flags, in a mail watchman, in a CSS class name — all the same pattern.
Cleanup the same week
Besides those two root causes, a few small but grating things were fixed in the same pass:
- Generating a barcode in the panel printed more than sixty warning lines into the output — deprecation notices from an old library on a newer PHP. Suppressed without altering the output.
- The admin dashboard polled every 15 seconds. Moved to 60; nobody noticed, the server relaxed.
- Some select boxes on the account forms were invisible — a style collision left over from a theme migration.
- A debug bar had been left enabled in production. Turned off, leftovers removed.
What I learned
“The site won’t open” is a symptom. The first reflex is to look at the code; in this case there was nothing wrong with the code.
And “the payment isn’t showing” is a symptom. There was something wrong in the code, but the code ran correctly — what was wrong was how failure was handled.
The lesson from both is the same: a system being quiet does not mean it is working.