YUNT
Article4 min read

Glassbox: How to minimize the AI ​​hallucination

Three years ago, a piece of software we developed cited a nonexistent source. That lesson led to the creation of the glassbox: a system that detects fabricated data before it's published.

Rodrigo Medina CEO @YuntRodrigo MedinaFounder & CEO
Imagen de Yunt Colaboradores Digitales

Three years ago, we helped a startup develop its product: AI-powered software that would allow anyone, at an early stage, to see if their invention was potentially patentable and what legal protections were available. That was the goal, and it still is. A task that takes a law firm weeks was within reach of someone just starting out. The product worked and was actually used.

But this column isn't about what went well. It's about what we discovered behind the scenes.

A Patent That Didn't Exist

While developing the solution, we noticed something that initially seemed like a bug: occasionally, the system would cite a patent that didn't exist. Not a misread patent or a number with a changed digit. An invented patent, with a plausible title, a credible number, and a perfectly written description of something 100% made up. And it did so with the same certainty as it cited real patents.

We tried everything—and remember, this was the heyday of the prompt: from "please don't hallucinate" to sophisticated instructions, double-checks, and layers of review. Every layer helped, but none was enough. The real test came when we wanted to try it for a more demanding use: as an expert witness tool. The verdict was clear: it wasn't foolproof. For guiding an inventor in the early stages, with a human reviewer, it was a great tool; for a report where every citation carries legal weight, we couldn't guarantee it. Today, all those "problems" are solved, and those problems are a thing of the past.

Over time, we understood why: we weren't fighting a flaw in our code, but a property of the material itself. A language model generates what is probable, and an invented patent is, for it, a perfectly probable text. The industry calls it hallucination. We suffered from it in a domain where misquoting isn't a minor detail: it's a serious error.

Changing the Question

That lesson remained stored for several years, as relevant lessons tend to be. And it matured into a change of question that now seems important when a company puts artificial intelligence to work seriously.

We stopped asking ourselves how to ensure the system didn't make mistakes. We started asking ourselves how to review the work done and validate, through a senior professional, that the error didn't come to light. Just as an intern presents their work to their boss for review.

These are different questions. The first doesn't have a complete answer: today there's no reliable way to look inside a neural network and guarantee what it will do, and anyone who promises otherwise is selling snake oil. The second does have an answer, and it's an old one: the same one any serious organization uses with people, whom they also can't inspect from the inside. Record what is done. Review what matters. A signature before anything with consequences.

Glassbox - The Glass Box

That's where what we now call the "glassbox" came from. The glass box, the opposite of the black box.

The mechanics are less glamorous than a state-of-the-art model, and that's why we trust them. Everything the digital collaborator does is recorded in an unalterable log: what they did, at what time, and what rule justifies it. And before a piece of work is released, the system cross-references each claim against the record of actions that actually occurred. If it's supported, it passes. If the support is partial, a person reviews it. If there's no supporting action, it's blocked: it's not sent, published, or signed until a human resolves it.

The first version of that audit verified only one thing: that every patent cited in a report existed in the patent registry. Exactly the same issue as three years prior. I don't believe in coincidences in engineering: you end up building protection against the blow you've already received.

The Glass Can't Be Added Later

And there's something we learned trying to add controls to that first system: transparency can't be added later. If the software didn't leave a trace of every action from day one, no audit will fix it, because there's nothing to audit against. You can't put glass on a black box. You have to build the glass box from the beginning—slower at first, easier to defend later.

Someone asked me if this wasn't an elegant way of admitting that technology fails. It's exactly that, and I don't see the problem with acknowledging it: technology fails all the time—bank websites crash, even WhatsApp crashes, and it causes a global uproar. The illusion isn't eliminated with the next version of the model. It's managed, with architecture, and it's managed better when you stop pretending it doesn't exist.

Who's Responsible

Three years ago, we would have liked to have this. We didn't, and we learned through human review. But no

Keep reading

Un colaborador digital de Yunt sentado en el suelo junto a un laptop abierto que muestra su rostro en pantalla.
Article

Harness: What surrounds the model

The model's intelligence is not what makes an AI work. What makes it work is the harness: six decisions a company makes for a digital collaborator, the same ones it makes for a person.

Rodrigo Medina CEO @YuntRodrigo Medina · Founder & CEO
Ilustración de una figura mecánica con el logo de Yunt que intercambia dos módulos, LLM A y LLM B.
Article

Astra, and why we are AI agnostic

OpenAI shipped GPT-6 Astra; the next is weeks away. AI agnostic means swapping the model is one line, and your handbook, data and records stay yours.

Rodrigo Medina CEO @YuntRodrigo Medina · Founder & CEO
All notes