EYThe LogEmre Yakut
← all entries
Systems · Metering · deep

Moving two thousand invoices into the background

At period close somebody has to press “render all invoices to PDF”. The server carries it for four minutes, then the browser gives up and nobody knows what happened. First I measured where it was actually stuck.

eskiden: web isteği içinde2.000 × 0,11 sn ≈ 3 dk 40 sntarayıcı vazgeçti · kaç tanesi üretildi, bilinmiyorPDF render = toplamın %97’si — optimize edilemezşimdi: iş kuyruğubuton → iş yaratılır → sayfa hemen dönerpanelde ilerleme · durdur / devam / iptalnabız gelmeyen iş toplanıriş #4471 · 1.284 / 2.000her 25 faturada: ilerlemeyi yaz + kontrol kanalını oku → temiz çıkış noktasıdevam edilince ZIP’e eklenir — üretilen 1.284 fatura tekrar üretilmeziş kaydı AYARLAR veritabanında: işçi hangi döneme bağlı olursa olsun oraya yazabilir

A period closes. Invoices are calculated for 2,000 flats. Now every one of them needs a PDF — to be printed, put in envelopes, delivered.

The old flow: press a button, the browser waits, the server starts rendering. Four minutes later the connection drops. Nobody knows how many were produced. You try again from the start.

First I measured where it was stuck

“Slow” is not a diagnosis. I timed the steps separately:

stepper invoiceshare of total
Database queries~0.002 s2.5%
Template compilationnegligible—
PDF render0.11 s97%

So optimising queries would have won at most 2.5% of the total. The bottleneck is entirely in PDF rendering.

And there is no way to make that render faster — the library takes what it takes. Which leaves one option: don’t speed the work up, remove the waiting.

simple multiplication
2,000 × 0.11 s ≈ 3 min 40 s

In a single request that exceeds the patience of both the web server and the browser. For a background process it is an entirely ordinary duration.

The job queue

Every heavy task is now a record. It lives in a table, with a state and a progress count.

job #4471
  type     : invoice_bulk_pdf
  state    : running
  progress : 1,284 / 2,000
  started  : 14:02:11
  heartbeat: 14:09:38        ← is the worker still alive?
  control  : —               ← pause / resume / cancel is written here

When the user presses the button the job is created and the page returns immediately. A worker starts in the background. The panel shows a progress bar.

Why the queue lives in the main database

In this system invoices live in period databases — each billing period has its own. The worker naturally connects to that period.

But if the job record goes there, the “running jobs” list in the main panel has no idea which period to look in. So the queue table lives in the settings database and is addressed by its fully qualified name. Whatever period a worker is attached to, it can write its progress there.

It looks like a small decision; the entire “see every job on one screen” capability depends on it.

Pause and resume

This was the most requested feature. When you start a 2,000-invoice job on the wrong period, you do not want to wait for it to finish.

Every N invoices the worker does two things: checkpoints its progress and reads the control channel.

every 25 invoices:
  write progress
  read control channel
    "pause"  → close the ZIP, set state 'paused', exit cleanly
    "cancel" → clean up and exit
    (empty)  → continue

On resume the worker starts where it left off and appends to the ZIP rather than rewriting it. If 1,284 invoices were produced, those 1,284 are not produced again.

Reaping dead workers

A worker can crash. A server can restart. The job record stays “running” — and stays that way forever.

Hence the heartbeat: the worker says “still here” at intervals. Jobs whose heartbeat has gone stale are reaped — marked as abandoned and made restartable.

what happens without this

The number of concurrent jobs is capped. A job that looks “running” but is actually dead occupies one of those slots. Let three or four accumulate and the queue jams completely — no new job ever starts, and on screen everything looks normal.

A heartbeat is how a system catches its own lie.

Then it generalised

I left the job type open from the start: job_type is a text field. The first use was bulk PDF, but others followed immediately:

  • Excel export
  • Thermal printer output (58 / 80 / 90 / 120 mm)
  • Bulk mail
  • Data import

All four use the same queue, the same panel and the same pause/resume machinery. The only new code was the work itself.

What I learned

“Slow” is a complaint, not a diagnosis. Measured, 97% of the bottleneck was in one place — and that place could not be optimised.

At that point the right move was not to go faster but to remove the waiting. For the user those 3 minutes 40 no longer exist; the job is started and they go and do something else.

SystemsMetering