Lumy Labs
← All insights

AI

Grounding is the part users actually judge

A model that is right most of the time and cannot show its work is unusable in a regulated setting. Retrieval quality and citation are the product, not the model.

Lumy Labs9 min read
The drawers of a wooden library card catalogue

The demo always goes well. Someone asks a question, the system answers fluently, the room is impressed. Then a subject-matter expert asks a question they already know the answer to, and the interesting part begins.

In enterprise settings the question is never whether the model is clever. It is whether a specialist can verify an answer quickly enough to trust the next one without checking.

Retrieval is where the accuracy is

Most quality problems in production assistants are retrieval problems, not generation problems. The model composed a reasonable answer from the wrong three documents, or from the right document at the wrong version.

This is good news, because retrieval is an engineering problem with familiar tools. Chunking that respects document structure, filters that respect permissions and effective dates, hybrid search when exact identifiers matter, and reranking when recall is good and precision is not.

  • Chunk on structure, not on a fixed character count
  • Filter by version and effective date before relevance
  • Combine keyword and vector search when identifiers matter
  • Rerank, then show the top sources to the person

Citations are a workflow feature

A citation is not a footnote for credibility. It is the mechanism by which a specialist checks an answer in seconds instead of minutes, and it determines whether the tool saves time at all.

That means the citation has to land the reader in the right paragraph of the right version, not on the front page of a hundred-page policy. A link that requires the reader to search again has moved the work rather than removed it.

Refusal is a feature you have to build

Systems that always answer are systems that sometimes invent. In domains where a wrong answer carries a cost, the ability to say that the corpus does not contain an answer is worth more than a few points of helpfulness.

This is mostly a retrieval-confidence decision rather than a prompting trick. If the best retrieved passages are weak, saying so is the correct behaviour, and it is far easier to earn trust back from a system that admits gaps than from one that fills them.

Evaluate on questions your experts already argue about

Generic benchmarks tell you very little about a specific corpus. The evaluation set that predicts production behaviour is the one built from real questions, including the ambiguous ones where two experts disagree.

Those cases are the valuable ones. They reveal whether the system surfaces the tension in the source material or papers over it, and papering over it is precisely what erodes trust the fastest.

Let's talk

Living with this problem?

If this one landed close to home, tell us where you are stuck and we will tell you honestly whether we can help.

Or email start@lumylabs.co

What happens next

  1. A real replyWithin 48 hours
  2. Before specificsNDA first
  3. Not a sales pitchA real proposal