EYThe LogEmre Yakut
← all entries
Metering · Field · deep

From a 77-page PDF to 68 products

A home-textile manufacturer’s only decent digital asset was a 49 MB catalogue PDF. I extracted the product tree out of it — and walked into two traps.

KATALOG167 MB · 68 ürün · tarama kalitesinde çıkar ürün adı koleksiyon ölçü kompozisyon renk kodu görselalan alan çıkarılmış, elle doğrulanmış site 68 ürün PDF bir veri kaynağı değildir — ama tek kaynak oysa veri kaynağı odur

A home-textile manufacturer in Bursa. What they make is obvious: duvet sets, quilts, trousseau sets, bridal sets, table linen. Digitally they barely exist.

They had one decent asset: a 77-page, 49 MB catalogue PDF. Their site showed 12 products. The catalogue held 68 models.

So 82% of what they sell was not on the internet.

A PDF is not a data source

Worth saying up front: a PDF is a presentation format. There are no tables, no relations, no schema. There are boxes placed on a page.

But if it is the only source, it is the data source.

Text was relatively easy: model name, collection, dimensions, composition, colour code. Images were the hard part.

Trap one: the ghost band

Extracting the images produced far more files than expected. Because the decorative band at the top of every page was embedded as an image too. Seventy-seven times. The same image.

The fix came from aspect ratio:

the band filter
width / height = 2200 / 400 = 5.5 → cannot be a product photo

I set the threshold at 3.0. Anything above it was treated as decorative. All 77 copies went in one pass, and not a single product photo was lost — the widest product image was 1.9.

Trap two: the fused panel

This one was sneakier. On the duvet pages, a dimensions-and-care panel was printed to the right of the product photo, inside the same image. One PNG with fabric on the left and text on white on the right.

You cannot put that on a site: crop it and the text is cut in half; leave it and there is an unrelated table next to the product.

To find the cut I scanned the image column by column and computed each column’s near-white pixel ratio:

# for each x column
white_ratio(x) = (pixels with luminance > 240)
                 ────────────────────────────────
                        image height

# photo side: textured, low and noisy   (~0.1–0.4)
# panel side: white ground, high and stable (>0.85)
cut = the first x where the ratio rises above 0.85 and never falls again

“Never falls again” matters. There are white regions inside the photo too — a bright pillow surface, for instance. Look at one column and you are misled. But once the panel begins it stays white to the edge, so I built the condition on persistence.

then I looked at them by hand

I opened all 68 images one by one. Four had cuts a few pixels inside; I fixed those manually.

Automated extraction does not remove verification — it shrinks the amount to verify. Instead of cropping 68 images by hand I checked 68 by hand. Three hours of difference.

The result: a 49 MB catalogue became a web-ready 10.4 MB image set across five collections — bridal 4, trousseau 18, quilts 12, duvet sets 28, table linen 6.

I measured the live site too

A habit when preparing a proposal: rather than praising or criticising a client’s site, measure it. Every finding with a timestamp and evidence.

measurementfinding
SSL certificatedoes not belong to the domain — browsers warn
Last content update4 June 2024 — frozen for two years
robots.txt / sitemap / schemanone present
Product coverage56 of 68 models are missing from the site
Page metadatatemplate text for a beauty salon, in another language
Document author fieldthe name of the student who made the template

The last two say the most: the site was built from a bought template and the demo text was never cleaned out. Search engines therefore think this is a beauty salon.

I kept the tone careful in the proposal. You do not say “your site is terrible”; you say “on this date we measured this, and here is the result”. One is an opinion, the other is a record. Only the first is arguable.

And I sent it… to nobody

Everything was ready and the portal link went out three times. All three to info@. Total views: 0.

The lesson is not technical: info@ is not a mailbox, it is a landfill. The decision-maker’s name — confirmed from the site’s own page metadata — was in my hand; I needed the channel.

On the second attempt we moved to a direct message and turned the open-tracking pixel off. Making someone feel watched while you are trying to sell to them costs more than the sale is worth.

What I learned

The hardest part of a job is rarely the most visible part. Here the site design took two days; getting the truth out of the catalogue took longer.

And however well you prepare it, anything sent to the wrong mailbox counts as unopened.

MeteringField