Someone in sales is on the phone with a customer. The customer says: “it goes on a twenty line, make it ultrasonic, and it should be read remotely.”
What the salesperson has to find is one item in the catalogue. The catalogue holds more than four thousand products and the search box looks only at names and codes. Typing “ultrasonic” returns 180 results.
To close that gap I wrote a semantic search layer into the CRM.
Let me name the wrong approach first
The easy answer: hand the sentence to a language model and say “list the matching products”.
That produces something that looks like it works and cannot be sold. Because the model does not know the catalogue, it invents plausible product codes. The salesperson puts one into a quote. No such product exists.
Once is enough to destroy trust in the whole system.
The rule: the model produces filters, not data
That one sentence is the architecture.
free text
↓ (intent parser)
{ hard filters, soft filters }
↓ (database query)
real products
The model — or the regex-based parser — only performs the middle step. The product list always comes from the database. If it is not in the catalogue, it is not in the result.
The ontology: defining what things are
For this to work the system has to understand product attributes. I defined those by hand: every attribute key, its kind and its weight.
| field | kind | why |
|---|---|---|
diameter | hard | A DN25 does not fit a DN20 line. Physical. |
product_type | hard | Never show a heat meter to someone asking for a water meter. |
pipe_connection | hard | Threaded or flanged — it either fits or it does not. |
pressure_class | hard | If the rating is insufficient the product is invalid. |
technology | soft | Ultrasonic preferred; mechanical may still do the job. |
communication | soft | wM-Bus preferred; M-Bus is also read remotely. |
certification | soft | A plus if present, a minus if not — but never disqualifying. |
An unmet hard filter removes the product. A soft filter only affects order. That split makes the result both correct and useful: elimination by physical fact, ranking by preference.
The intent parser
The piece that turns free text into filters. Surprisingly, most of it is not a model but pattern matching:
"twenty" → diameter = DN20
"DN 20" → diameter = DN20
"20mm" → diameter = DN20
"ultrasonic" → technology = ultrasonic
"remote" → communication ∈ {wmbus, mbus, lora}
The subtlety is resolving conflicts. When “wmbus” matches, “mbus” is removed — otherwise it looks as if two different protocols are being asked for at once and the result blurs.
And the key rule: the parser cannot produce a key that is not in the ontology. An undefined field means the request is rejected. No door is left open for hallucination.
Then MCP: letting Claude connect directly
With that layer in place I went one step further and wrote an MCP server into the CRM, so an AI client can connect and search the catalogue directly.
The server lists its tools, the client invokes them. The data is never handed to the model in bulk — the model only ever sees the answer to the question it asked.
The legacy database layer terminated the process with die() on error. Because MCP is a JSON-RPC protocol, that sent the client a truncated, invalid response and dropped the connection.
The fix was to buffer output and register a shutdown hook: if the process ends unexpectedly the buffer is discarded and a valid JSON-RPC error response is written in its place. The client says “something went wrong”, not “the protocol broke”.
When you attach an old code base to a new protocol, the time goes not into features but into these exit paths.
What I learned
When adding AI to a system the right question is not “how smart is the model?” It is “what happens when the model is wrong?”
Here the answer is: it produces a wrong filter. The result is either irrelevant products or none at all. Both are visible and correctable.
Had I let the model invent products, a wrong answer would have looked exactly like a right one. That is the entire difference.