EYThe LogEmre Yakut
← all entries
Sound & light · AI & Claude · Systems · deep

Who spoke, where from, what now — three separate questions

NEO is an always-listening assistant at home. My first mistake was treating those three as one question. Splitting them raised accuracy and made it visible where the system was going wrong.

mikrofon sürekli akış VAD uyarlanan eşik kim? konuşmacı ayrımı nereden? yön + yüz ne yapmalı? niyet üçünü tek soruda sormak, üçünü de yanlış cevaplamaktıreşik sabit olursa: klima açıkken sürekli konuşuyor sanır, fısıltıyı hiç duymaz.çözüm: gürültü tabanını sürekli ölç, eşiği tabanın üstünde gezdir — ölçümle kalibre et.

An assistant that stays on at home, wakes on its name and answers when spoken to. A long-term project; the first phase was the audio loop.

The first version worked like this: listen to the microphone, record when there is sound, transcribe, respond. And it was constantly wrong.

The mistake: treating three questions as one

“Was there sound?” actually hid three separate questions:

  1. Is this speech? — not the air conditioner, not the television, a human voice.
  2. Who is speaking? — me, someone else, or the assistant’s own voice.
  3. Was it said to me? — two people in the room may be talking to each other.

Trying to solve all three with one threshold meant none of them worked. I split them.

1. Is this speech: unpinning the threshold

The first version had a fixed energy threshold. It broke in two opposite ways:

  • With the air conditioner on, the noise floor rose above the threshold, the system thought there was continuous speech and kept opening recordings.
  • In night-time silence, a sentence at normal volume sometimes fell below the threshold — because it had been tuned for a noisy room.
adaptive threshold
threshold(t) = noise_floor(t) + margin

The noise floor updates continuously from a slow moving average of quiet windows. Turn the air conditioner on and the floor rises; the threshold rises with it.

There is also a persistence requirement: one window above the threshold is not enough, several consecutive ones are. That removes transients like a door slam or a keystroke.

But what mattered most was stopping guessing the margin and measuring it. I recorded in several real environments — quiet room, air conditioner on, television on, water running in the kitchen — and set the margin from those recordings. A measured constant instead of an invented one.

2. Who is speaking

Identification was not needed here — separation was. “A said this, B said that” is enough.

[00:03] speaker_1 : what time are we leaving tomorrow
[00:06] speaker_2 : around nine
[00:09] speaker_1 : neo, set an alarm for nine

The third line is for the assistant; the first two are not. Without separation all three arrive in the same stream and the system tries to answer all of them.

the sneakiest speaker: the system itself

When the assistant answers through the speaker, that audio comes back through the microphone and is processed as a new command. The system starts talking to itself.

The fix is two-layered: suppress input while speaking, and recognise and discard the text of its own output when it arrives. Both are needed — suppression alone does not catch a delayed echo.

3. Was it said to me

The hardest, and not solvable with one signal. Several weak signals are combined:

signalwhat it saysreliability
Wake worddirect addresshigh
Directionis the voice facing the devicemedium
Faceis someone on camera looking at itmedium
Sentence formimperative or conversationallow

The wake word is strongest, but people do not repeat it in every sentence — they say it once and then speak three. So after a wake a short “listening window” opens during which the threshold drops.

Remembering and verifying

Conversations and the assistant’s own utterances are written to disk. An assistant that does not remember what it said repeats itself, contradicts itself and loses trust.

There is also a filter on the output side. In a system wired to home devices this is not a theoretical concern — saying “I turned the light off” without turning it off is worse than not turning it off.

The rule: an action is verified before it is announced. If the device did not respond, the assistant does not say “done”, it says “I could not reach it”. That is the audio-side equivalent of “a field without a source never reaches the screen”.

And the permission layer is not optional

What this system can reach: microphone, camera, email, home devices and the ability to write code.

That combination makes a permission layer and an audit log mandatory. Which action was taken when and on whose request is recorded. It is not something to add later — an audit added later does not cover the past.

What I learned

When a problem will not yield, it is usually divided wrongly. “Was there sound?” looked like one question; split into three, each piece revealed its own correct solution.

And each piece became individually measurable. Combined, I could only say “it sometimes gets it wrong”; separated, I can see which layer was wrong.

Sound & lightAI & ClaudeSystems