EYThe LogEmre Yakut
← all entries
Protocol · Systems · deep

Don’t walk the disk, read its ledger — and two tools that tidy my desk

I wrote a disk analyser. Walking folder by folder took nine minutes to scan 1.4 million files. NTFS already keeps a record of every one of them in a single table.

klasör klasör gezmekC:\ Users\ emrey\ Documents\ … …1.4 M dosya · ~9 dk · disk sürekli arıyor$MFT’i doğrudan okumakFILE0 #1024 $STANDARD_INFOFILE0 #1031 $FILE_NAMEFILE0 #1038 $DATAFILE0 #1045 $FILE_NAMEFILE0 #1052 $DATAaynı 1.4 M dosya · 11 sn · tek sıralı okumaNTFS zaten her dosyanın kaydını tek bir tabloda tutuyor — sormak yerine tabloyu oku

One morning in early September I noticed: I write systems that bring order to clients, and my own working environment is a mess. The disk permanently red, three hundred unread messages, three copies of the same document.

I turned both into tools.

Disk: why walking folders is slow

The classic approach is recursive traversal: list the root, descend into each subfolder, repeat.

C:\               → 34 entries
  Users\          → 8 entries
    emrey\        → 91 entries
      Documents\  → 1,204 entries
        …

Each folder is a separate system call. Each call goes somewhere different on disk. For 1.4 million files that is hundreds of thousands of small scattered reads.

Measured: 9 minutes 12 seconds.

But NTFS already keeps a ledger

Then I remembered the obvious: NTFS already holds all this in one place. It is called $MFT — the master file table. Every file and folder on the volume has a fixed-size record in it:

FILE0  #128934
  $STANDARD_INFORMATION   created, modified, accessed, attributes
  $FILE_NAME              name + the parent folder's record number
  $DATA                   either the content itself or where it sits on disk
a nice detail

A very small file’s content lives inside the $DATA attribute — no separate disk allocation. So thousands of tiny files can occupy less than the sum of their sizes. The converse is also true: every file consumes at least one record, so even a one-byte file is not free.

Building the tree

A record does not know its own full path. All it knows is its parent’s record number. That is actually convenient:

1. read $MFT sequentially      → one big read, the disk is happy
2. index records by number
3. for each file, walk the parent chain:
     128934 → 4102 → 91 → 5 (root)
     "Documents\project\report.pdf"

The result: 11 seconds. Fortyfold. And the difference is not algorithmic cleverness — same data, same disk. The only difference: reading the whole thing once instead of asking piece by piece.

It has a price: administrator rights are required, it only works on NTFS (the old traversal stays as a fallback), and files can change while the read is in progress. Over an eleven-second window that is acceptable; for a backup tool it would not be.

Getting faster changed the product

The unexpected part: once the tool got faster, what it was for changed.

A nine-minute scan is something you run once a month. An eleven-second scan is something you open when you are curious.

And being frequently runnable made a new question possible: what changed between yesterday and today? “This folder is 40 GB” is useless — it always was. “This folder was 2 GB yesterday” turns directly into action.

Because the real question is not “what is big?” but “what is growing?”

The second tool: the inbox

The essence of the inbox problem is not “too much mail”. It is uncertainty about which threads are still open.

Most of those three hundred messages are finished business. But seeing that they are finished requires opening each one.

The tool tracks conversations and asks: who sent the last message in this thread, and what kind was it?

last messagestate
From me, contains a questionwith them — I am waiting
From them, contains a questionwith me — needs a reply
From them, thanks/confirmationclosed
Machine-generatedclosed

The fourth row was larger than I expected: auto-replies, system notifications, newsletters — more than a third of the inbox.

Detecting them is harder than it looks. Pattern-matching subject lines is not enough; certain headers are far more reliable. Using both together brings false positives to almost zero.

Result: 300 messages became 19 conversations genuinely awaiting a reply.

Why single-user tools are a good school

What both share: the user count is one, and that user is me.

  • No requirements debate. I know what I want, because I am the one suffering without it.
  • Feedback is immediate. A bad decision hurts the same day, not in a support ticket three months later.
  • Scope stays narrow by itself. There is no “but what if someone wants…”.
  • The pressure to finish is real. An unfinished tool does not do my own job.
an automation rule

If you write a tool that manages your own files, every action it takes must be reversible. Move, do not delete, and always keep a record. Assume from the start that it will one day apply a wrong rule; the question is not “will it” but “what do I lose when it does”.

And an unexpected addition

I later added a “work catalogue” to the disk tool: it finds project folders on the drive and lists them with their last modification dates.

My reason for writing it was faintly absurd — I had lost track of how many projects there were. When the list appeared it looked back at me: how many things I had started this year.

The idea for this log came from that list.

What I learned

A system may already have answered the question you are asking. The job is not to reproduce the answer but to find where it is kept.

And “I don’t have time for that” is usually said without ever costing the alternative. Two tools took a week; they save twenty minutes a day. Break-even in a month.

The real gain is not time either: it is attention. A full disk or three hundred unread messages keeps a load running in the background all day.

ProtocolSystems