A language model without a verified data foundation invents pack sizes, prices and authorisation status, and does so fluently. This page sets out what a data foundation has to provide for an AI application to answer dependably in a professional setting.






Asked for a pack size, a price or a marketing status without a verified foundation, a language model still produces an answer. It is fluent, plausible and, when the fact is missing, wrong. In a professional healthcare setting that is the difference between a useful tool and a liability.
Most of the errors customers show us as hallucinations turn out to be something else: a product name that could not be mapped to anything. Trade names are not unique, not stable and do not travel across borders, so an answer anchored to one has nothing behind it.
What may be done with drug data inside your own application decides whether you can cache it, index it or pass it on. Teams that reach this question in week three of the build have usually already chosen an architecture that the licence does not support.
An application can only be as unambiguous as its input data. At pack level that is the Pharmazentralnummer; at therapeutic level the active substance and the ATC code. A trade name is a string, not a mapping.
A price without a reference date is a guess. Every field carries the date it refers to and the source it came from, which is what makes a generated answer checkable afterwards.
A licence written for on-screen use does not cover a model pipeline. Caching, enrichment and onward use inside your product are exactly what decides the architecture, so the scope is agreed before you build.
Data arrives as a web service or as a configurable download, with scope and refresh rhythm set per data type. What you build on top stays your architectural decision.
A language model asked for a pack size, a price or a marketing status without a verified data foundation will still answer. It will answer fluently, plausibly and, when it does not know, wrongly. In a professional healthcare setting that is the difference between a useful tool and a liability. We do not supply a model. We supply the layer underneath it: structured, dated and licensed drug data your application can answer against.
| Requirement | What it means | How it fails |
|---|---|---|
| A stable identifier | every statement hangs on a product identifier rather than a trade name | trade names are not unique, not stable and do not travel across borders |
| A dated state | every value carries the date it refers to and the source it came from | a price without a reference date is not information, it is a guess |
| A matching licence scope | machine processing, caching and onward use inside your own product are explicitly covered | a licence written for on-screen use does not cover a model pipeline |
The third row is the one projects reach last and pay for most. What may be done with drug data inside your own application decides the architecture, so it belongs at the start.
An AI application can only be as unambiguous as its input data. In Germany the Pharmazentralnummer carries that at pack level; the active substance and the ATC code carry it at therapeutic level. An answer anchored to a trade name has no mapping, only a string. That is where most of the hallucinations customers show us actually originate: not in the model, but in a name that could not be resolved.
The data arrives the same way as for any other integration, as a web service or as a configurable download, with scope and refresh rhythm set per data type. What you build on top, whether a runtime lookup, an index for retrieval-grounded answers or a verification step after generation, remains your architectural decision.
Three limits we would rather state up front. First, we provide no model and no training, only the data foundation. Second, structured data does not replace professional judgement: it shortens the path to a decision, it does not make it. Third, any onward processing depends on our data partners' licence, and that question gets a binding answer in conversation rather than an approximate one on a web page.
Anyone building an AI application in a healthcare context also operates within the scope of the EU AI Act. How a given system is classified depends on its intended purpose and is not ours to advise on. What we can contribute is traceability of the factual base: a value with a source and a date can be checked, a generated value with neither cannot.
The quickest way to judge fit is a 30-minute demo. Bring the questions your application is meant to answer, and we will show which fields that needs and which do not exist.
More clarity, faster research, and faster decision-making.






Because they are built to produce plausible text, not to evidence a value. Without a verified data foundation the model fills the gap with something that sounds right. Pack sizes, prices and marketing status are especially exposed, because they change often and are rarely present in training material in their current form.
Three things: every statement hangs on a stable product identifier rather than a trade name, every value carries a reference date and a source, and the licence scope explicitly covers machine processing, caching and onward use inside your own product. The third point decides the architecture and is most often raised too late.
No. We provide the data foundation, not the model and not training. What you build on it, whether a runtime lookup, an index for retrieval-grounded answers or a verification step after generation, is your architectural decision. Our part is making sure the facts underneath are correct and dated.
That depends on our data partners' licence and cannot be waved through on a web page. This is precisely the question that decides your architecture, so it belongs in the first conversation rather than in week three of the project. You will get a binding answer there rather than a comfortable one.
At pack level the Pharmazentralnummer; at therapeutic level the active substance and the ATC code. A trade name is not an identifier: it is not unique, not stable and does not travel across borders. A large share of the errors shown to us as hallucinations are in fact names that could not be resolved.
No. It shortens the path to a decision, it does not make it. The pharmaceutical assessment stays with qualified staff. What structured, dated data delivers is traceability: a value with a source and a date can be checked, a generated value with neither cannot.