The shiny toy
The AI project that dazzles in the demo and dies in a drawer has a name and an antidote. A three-signal test to tell a tool from a whim.
What we discover while building AI in production with our clients.
The AI project that dazzles in the demo and dies in a drawer has a name and an antidote. A three-signal test to tell a tool from a whim.
One assistant answers several companies and none of them can see another. The four layers that guarantee it, a lesson about permissions and the redesign that erased a whole class of failures.
We sell AI agents, and we don't like how almost everyone builds them. Business rules cannot live in the prompt, they have to live in the code.
An entire market sells prompt engineering. Our production experience says you gain more by organizing the data than by polishing the prompt.
A zero is not a neutral answer, it is an unclassified alarm. The three causes of a "no data" response and why the third one is the most treacherous.
Brilliant AI projects launch well and die young. Profitable ones are easy to start and easy to maintain. The difference is decided before any code is written.
Documenting how a failure is recognized from the outside pays more than documenting how it was fixed. Three real signatures from an AI assistant in production.
A second model that distrusts by design, a deliberately short memory, and the metric that catches an assistant answering from memory.
A conversational assistant depends on systems that fail. The circuit breaker that protects the user, and the three lessons production taught us about it.
Productivity is AI's comfortable metric, it always looks good and commits to nothing. We prefer the uncomfortable one, which concrete gain did the system move.
When an AI assistant cannot answer, there are two causes of a different nature. Confusing them is the classic mistake that inflates documentation without fixing anything.
A hallucination is a well-written, false answer. The four mechanisms our systems use to corner them in production, with their numbers and their scars.
The text-to-SQL pattern promises a model that writes queries. We built it for Savian with the opposite decision, and that decision is what makes it safe.
What changes when OCR meets a language model, how it looks in a real case with utility invoices and why validation is the actual product.
Dozens of WhatsApp messages a day, five to ten minutes per inquiry and an overloaded team. What the agent that filters requests for a Barcelona agency actually does.