I am building an always-on voice assistant at home. Its brain runs on an existing tool — a persistent process I talk to and that answers back.
It worked. And it was unbearably slow: I say a sentence, and 8.9 seconds median later an answer arrives. In a voice assistant that is dead; a person asks “did you hear me?” at the third second.
The first diagnosis, and why it was wrong
The sensible explanation: the model is large and takes time to think. And the fix is obvious — use a smaller model, turn off unnecessary tools.
I added the flags. No error. The process started. And the duration did not change.
I added something and nothing happened. That exact sentence led me to the right place.
The evidence is in the start-up event
The process emits a start-up event reporting its state. I had never looked at it.
{
"type": "system", "subtype": "init",
"model": "…large model…", ← not what I asked for
"tools": [ … 30 tools … ] ← all enabled
}
Neither flag had been applied. Not rejected, not warned about — silently ignored.
Silent failure. The system behaves as if it accepted the request, does nothing, and says nothing.
The most dangerous kind, because you think you fixed it. You write the wrong configuration, decide “right, that is handled, next”, and the real problem sits untouched.
What worked: a file, not a flag
Since flags did nothing, another channel was needed. I tried the configuration files in the tool’s working directory: a settings file and a rules file.
I restarted and looked at the start-up event first — that is now the first thing I do after any change:
{
"model": "…small model…",
"tools": []
}
It held. The same for personality and behaviour: the instruction passed as a flag had no effect, while the rules file in the working directory did.
The result
| state | response time (median) |
|---|---|
| Default — large model, 30 tools enabled | 8.9 s |
| Via configuration file — small model, no tools | 1.8 – 3.3 s |
And here is the real surprise: the slowness was not caused by the model’s size.
With tools enabled, the model looked at the file system before answering a simple conversational question. Asked about the weather, it was searching the project for files. Most of those eight seconds were not thinking but wandering.
Turning the tools off collapsed the latency. Shrinking the model helped, but would not have been enough alone.
The third trap: an inherited environment
One day the assistant behaved differently when launched from the desktop than from a terminal. Same code, same configuration.
The cause: when launched from inside a development session, the child process inherited its parent’s environment variables. Launched from the desktop there is no such inheritance.
The fix: clear the relevant variables explicitly before starting the process. The same pattern showed up on the server.
A setting being accepted and being applied are different things. The only way to see the difference is to look at the state the system reports.
I spent weeks adding flags and assuming the outcome; one start-up event told me more than all of them.
The same pattern elsewhere
- A configuration file never read — wrong path, file exists, no error, defaults in use.
- An unknown key — a typo in configuration; the parser silently ignores it.
- An overriding layer — the setting is read correctly and then clobbered by something later.
In all of them the fix is the same: after writing a setting, verify the effective state.
Closing
This project gave me the most concrete lesson I have about using AI, and it is not about the model: if you want to make a tool faster, first measure what it is actually doing.
I had diagnosed “the model is slow”. It was not; I was making it do unnecessary work, and every attempt I made to stop that had never been applied.