A period closes. Invoices are calculated for 2,000 flats. Now every one of them needs a PDF — to be printed, put in envelopes, delivered.
The old flow: press a button, the browser waits, the server starts rendering. Four minutes later the connection drops. Nobody knows how many were produced. You try again from the start.
First I measured where it was stuck
“Slow” is not a diagnosis. I timed the steps separately:
| step | per invoice | share of total |
|---|---|---|
| Database queries | ~0.002 s | 2.5% |
| Template compilation | negligible | — |
| PDF render | 0.11 s | 97% |
So optimising queries would have won at most 2.5% of the total. The bottleneck is entirely in PDF rendering.
And there is no way to make that render faster — the library takes what it takes. Which leaves one option: don’t speed the work up, remove the waiting.
In a single request that exceeds the patience of both the web server and the browser. For a background process it is an entirely ordinary duration.
The job queue
Every heavy task is now a record. It lives in a table, with a state and a progress count.
job #4471
type : invoice_bulk_pdf
state : running
progress : 1,284 / 2,000
started : 14:02:11
heartbeat: 14:09:38 ← is the worker still alive?
control : — ← pause / resume / cancel is written here
When the user presses the button the job is created and the page returns immediately. A worker starts in the background. The panel shows a progress bar.
Why the queue lives in the main database
In this system invoices live in period databases — each billing period has its own. The worker naturally connects to that period.
But if the job record goes there, the “running jobs” list in the main panel has no idea which period to look in. So the queue table lives in the settings database and is addressed by its fully qualified name. Whatever period a worker is attached to, it can write its progress there.
It looks like a small decision; the entire “see every job on one screen” capability depends on it.
Pause and resume
This was the most requested feature. When you start a 2,000-invoice job on the wrong period, you do not want to wait for it to finish.
Every N invoices the worker does two things: checkpoints its progress and reads the control channel.
every 25 invoices:
write progress
read control channel
"pause" → close the ZIP, set state 'paused', exit cleanly
"cancel" → clean up and exit
(empty) → continue
On resume the worker starts where it left off and appends to the ZIP rather than rewriting it. If 1,284 invoices were produced, those 1,284 are not produced again.
Reaping dead workers
A worker can crash. A server can restart. The job record stays “running” — and stays that way forever.
Hence the heartbeat: the worker says “still here” at intervals. Jobs whose heartbeat has gone stale are reaped — marked as abandoned and made restartable.
The number of concurrent jobs is capped. A job that looks “running” but is actually dead occupies one of those slots. Let three or four accumulate and the queue jams completely — no new job ever starts, and on screen everything looks normal.
A heartbeat is how a system catches its own lie.
Then it generalised
I left the job type open from the start: job_type is a text field. The first use was bulk PDF, but others followed immediately:
- Excel export
- Thermal printer output (58 / 80 / 90 / 120 mm)
- Bulk mail
- Data import
All four use the same queue, the same panel and the same pause/resume machinery. The only new code was the work itself.
What I learned
“Slow” is a complaint, not a diagnosis. Measured, 97% of the bottleneck was in one place — and that place could not be optimised.
At that point the right move was not to go faster but to remove the waiting. For the user those 3 minutes 40 no longer exist; the job is started and they go and do something else.