An assistant that stays on at home, wakes on its name and answers when spoken to. A long-term project; the first phase was the audio loop.
The first version worked like this: listen to the microphone, record when there is sound, transcribe, respond. And it was constantly wrong.
The mistake: treating three questions as one
“Was there sound?” actually hid three separate questions:
- Is this speech? — not the air conditioner, not the television, a human voice.
- Who is speaking? — me, someone else, or the assistant’s own voice.
- Was it said to me? — two people in the room may be talking to each other.
Trying to solve all three with one threshold meant none of them worked. I split them.
1. Is this speech: unpinning the threshold
The first version had a fixed energy threshold. It broke in two opposite ways:
- With the air conditioner on, the noise floor rose above the threshold, the system thought there was continuous speech and kept opening recordings.
- In night-time silence, a sentence at normal volume sometimes fell below the threshold — because it had been tuned for a noisy room.
The noise floor updates continuously from a slow moving average of quiet windows. Turn the air conditioner on and the floor rises; the threshold rises with it.
There is also a persistence requirement: one window above the threshold is not enough, several consecutive ones are. That removes transients like a door slam or a keystroke.
But what mattered most was stopping guessing the margin and measuring it. I recorded in several real environments — quiet room, air conditioner on, television on, water running in the kitchen — and set the margin from those recordings. A measured constant instead of an invented one.
2. Who is speaking
Identification was not needed here — separation was. “A said this, B said that” is enough.
[00:03] speaker_1 : what time are we leaving tomorrow
[00:06] speaker_2 : around nine
[00:09] speaker_1 : neo, set an alarm for nine
The third line is for the assistant; the first two are not. Without separation all three arrive in the same stream and the system tries to answer all of them.
When the assistant answers through the speaker, that audio comes back through the microphone and is processed as a new command. The system starts talking to itself.
The fix is two-layered: suppress input while speaking, and recognise and discard the text of its own output when it arrives. Both are needed — suppression alone does not catch a delayed echo.
3. Was it said to me
The hardest, and not solvable with one signal. Several weak signals are combined:
| signal | what it says | reliability |
|---|---|---|
| Wake word | direct address | high |
| Direction | is the voice facing the device | medium |
| Face | is someone on camera looking at it | medium |
| Sentence form | imperative or conversational | low |
The wake word is strongest, but people do not repeat it in every sentence — they say it once and then speak three. So after a wake a short “listening window” opens during which the threshold drops.
Remembering and verifying
Conversations and the assistant’s own utterances are written to disk. An assistant that does not remember what it said repeats itself, contradicts itself and loses trust.
There is also a filter on the output side. In a system wired to home devices this is not a theoretical concern — saying “I turned the light off” without turning it off is worse than not turning it off.
The rule: an action is verified before it is announced. If the device did not respond, the assistant does not say “done”, it says “I could not reach it”. That is the audio-side equivalent of “a field without a source never reaches the screen”.
And the permission layer is not optional
What this system can reach: microphone, camera, email, home devices and the ability to write code.
That combination makes a permission layer and an audit log mandatory. Which action was taken when and on whose request is recorded. It is not something to add later — an audit added later does not cover the past.
What I learned
When a problem will not yield, it is usually divided wrongly. “Was there sound?” looked like one question; split into three, each piece revealed its own correct solution.
And each piece became individually measurable. Combined, I could only say “it sometimes gets it wrong”; separated, I can see which layer was wrong.