SummaryThe Summary of Product Characteristics exists in two forms, and the difference decides the workload: as a document, and as structured data. Anyone who wants to read a single SmPC will find it free of charge on a regulator's portal. Anyone who wants to feed SmPC content into a system, whether that is a search function, a medication review or an AI assistant, needs it section by section, version stamped, and on an agreed update cadence. That is a licensing question rather than a download question, and the gap between the two is where most projects lose a quarter.
This guide covers the three sourcing routes, the fields a system actually needs, the versioning problem that stalls implementations, and the clauses that belong in the licence before the architecture is fixed. What an SmPC contains and how it is structured is covered separately in what is an SmPC; how authorised texts are changed and kept traceable across markets is in SmPC compliance and lifecycle management.
There are three ways to get SmPC content. They do not differ in what the text says. They differ in the form it arrives in, and in how much work is left afterwards.
| Route | Form | Carries | Where it stops |
|---|---|---|---|
| Regulator portals | PDF, one document at a time | Single lookups, authoritative source, no cost | No field structure, no bulk retrieval, currency is your problem |
| Research interface | Searchable full text | Human research, comparison across products | No transfer into your own systems |
| Data licence | Structured delivery or API | Loading into your systems, processing, analysis | Licensed, with cadence and retention set by contract |
The most common planning error is the assumption that the first route scales into the third. Pulling individual PDFs answers a question; it does not create a dataset. As soon as a system itself searches, checks or evaluates the content, you need the sections addressable.
Note also where the source documents actually live. For centrally authorised medicines, the European Medicines Agency publishes the SmPC. For nationally authorised medicines, each national competent authority publishes its own. There is no regulator-side consolidated store that holds every national version in its current state, which is precisely why cross-border projects are harder than single-market ones.
An SmPC follows a fixed section structure. In the structured form each of those sections is individually addressable rather than buried in running text. That sounds like a detail and decides whether a query is possible at all.
Take a concrete question: which products containing this active substance carry a contraindication in renal impairment? Against structured data, that is a query against the contraindications section. Against a folder of PDFs, it is a full-text search with an unknown recall rate, because the same clinical statement is worded differently by different authorisation holders. One returns a list you can act on. The other returns a list you have to check by hand, and you will not know what it missed.
The same applies to posology, undesirable effects and the pregnancy and lactation statements. Section structure is the precondition for machine evaluation, and it is the reason the regulatory direction of travel, electronic product information built on structured standards, points exactly this way.
Field requirements follow function. Three tiers cover almost every implementation.
Tier 1, identification and matching. The SmPC has to attach to the right product. That needs the identifier as a key, the brand name, the authorisation holder, the dosage form and the strength. Without this tier you hold a document that cannot be matched to anything in your own inventory, which is the single most common reason a pilot does not become a system.
Tier 2, the clinical sections. Therapeutic indications, contraindications, interactions, undesirable effects, posology, pregnancy and lactation. This is the core the SmPC is wanted for.
Tier 3, the handling layer. Storage and shelf life, including shelf life after first opening or reconstitution, plus the full composition with active substances and excipients. These matter the moment medicines are not merely prescribed but dispensed, split, prepared or stored.
| Field | What it is for | Tier |
|---|---|---|
| Product identifier (PZN in Germany) | Matching to the article in your own inventory | 1 |
| Brand name and authorisation holder | Telling identical substances apart, contact for queries | 1 |
| Dosage form and strength | Selecting the right pack | 1 |
| Active substances with quantity, excipients | Substance search, allergy checks via excipients | 1 |
| ATC code | Grouping across brand names and markets | 1 |
| Therapeutic indications | Indication-based research | 2 |
| Contraindications | Patient-specific checking | 2 |
| Interactions | Medication review | 2 |
| Undesirable effects | Assessing reported events | 2 |
| Posology and method of administration | Therapy planning, dose adjustment | 2 |
| Pregnancy and lactation | Advice for specific patient groups | 2 |
| Storage and shelf life | Storage, shelf life after opening | 3 |
Two fields are left out of specifications with striking regularity and retrofitted expensively later. The first is excipients: allergy checks fail without them, because the intolerance often concerns an excipient rather than the active substance. The second is shelf life after opening, which in ward supply and blister packing is the difference between usable stock and waste.
Rather than starting with which vendor, you get there faster by settling four things. They determine the route almost completely.
One: does a person read it, or does a machine process it? If a person reads it, a research hand-off from your own system is enough, and the data stays with the provider. If a machine processes it, you need the fields in your own store, and different licence terms apply.
Two: is one market enough? A single national source covers a single national market. The moment cross-border sourcing, parallel trade or an international group inventory enters, it does not, because the same substance carries its own SmPC in each country, each with its own revision state and its own language.
Three: which sections are needed? An application that displays indications has a different field set from one that checks contraindications against patient data. Scope is a price factor, and a field set cut too wide costs continuously rather than once.
Four: how current does the dataset have to be? For a lookup function a monthly state is often fine. For a check that gets documented, the update cadence is part of the evidence trail.
With those four answers you can describe the sourcing route in one sentence, and the licensing conversation takes one round rather than three.
SmPC and package leaflet turn up in requirement documents as though they were two editions of one text. They are not.
The SmPC addresses healthcare professionals; the package leaflet addresses patients. The direction is fixed: the leaflet is derived from the SmPC, never the other way round. A change to the SmPC pulls a change to the leaflet behind it, and a leaflet cannot carry information the SmPC does not support.
For a system that means: if you need both texts, for instance because an application serves professionals and patients, you license both and you keep both current. Budgeting for one and discovering you need two is a common and avoidable surprise.
One use case now comes up in nearly every conversation and deserves its own section, because it places specific demands on the form the data arrives in.
Using SmPC content as the knowledge base for a language-model application raises two problems that have nothing to do with model quality. The first is attributability. An answer drawn from an SmPC is only usable if you can trace which product, which section and which revision it came from. That requires the source to deliver those markers rather than emitting running text.
The second is currency. A model working from a frozen snapshot gives an answer that was correct six months ago. For contraindications and warnings that is not an academic problem. Robust applications therefore read at runtime from a continuously maintained dataset rather than relying on a model that has memorised the text.
Both lead to the same requirement as above: structured, versioned, on an agreed cadence. And to an extra clause in the licence, namely whether use in such an application is covered at all. It is not automatic, and it should be settled before the architecture decision rather than after it.
An SmPC is not a finished document. It changes across the entire lifecycle of the medicine as safety data accumulate and indications are added or restricted. That is why it carries a revision date.
In practice this means three things.
First, "the SmPC for product X" always means the current version, and current is a moving target. A dataset without an agreed update cadence stops being a reference after a few months and becomes a snapshot that people still treat as a reference.
Second, the agreed cadence also determines how long you may hold a given state in your own system. That is a licence question, not a technical one, and it is the clause most often skipped in a data contract review.
Third, anyone who needs SmPCs across several countries is dealing with separate national versions, each on its own revision cycle and in its own language. As noted above, no regulator-side source consolidates them in their current state, so either you build that reconciliation or you buy it.
A point regularly underestimated in implementation: sourcing is only half the job. The other half is recognising what changed.
If every delivery replaces the entire dataset, the system loses the information about which products were affected. For a lookup function that is harmless. For a quality system that has to demonstrate it reacted to an amended contraindication, it is a break in the evidence chain, and it is the sort of break that is only discovered during an audit.
Two things are therefore worth putting in the contract. First, whether the delivery marks changes, meaning which records are new, amended or withdrawn since the last delivery. Second, whether the revision date is delivered per record, so that your own system can show which version a decision rested on.
Designed in from the start, both cost almost nothing. Retrofitted, both are expensive, and the second is sometimes impossible for historical records.
"SmPCs are public, so they are free to use." The texts are publicly viewable, which is true. It does not follow that their structured preparation, their linkage to a product identifier and their continuous maintenance are free to use. That preparation is the licensed service, and it is the reason a build-your-own is rarely cheaper than it looks on a slide.
"We will pull the PDFs and parse them ourselves." Technically possible, and in a pilot it works. The effort is not in the first pass but in the maintenance: layout changes, inconsistent wording between authorisation holders, new products, amendments without notice. Choosing this route is not building a feature, it is accepting a permanent obligation that usually has nothing to do with your core business.
"One source covers everything." SmPCs answer questions about the medicine. They say nothing about the price, dispensing status or availability of a specific pack. Those live in the article master and are a separate sourcing question, covered in the pharmaceutical article master explained. Anyone who needs both, which in trade is the normal case, should scope across both layers from the start.
pharmazie.com is the consolidated pharmaceutical data platform operated by DACON Datenbank Consulting GmbH, on the market since 1989. It brings 25+ pharmaceutical databases into a single search, the Eisbergsuche®, and is addressed exclusively to healthcare professionals rather than patients. Four components are relevant for SmPC sourcing.
Portal access starts at 135 EUR per month, net plus VAT. How a raw licence differs from an application-level entitlement is set out in raw licence versus API for German drug data.
Not every source carries complete SmPCs. Some reference works hold condensed summaries rather than full product information. For day-to-day lookups that is often enough; for systematic evaluation it is not. Check this explicitly before a requirements document gets written, because the word "SmPC" in a vendor datasheet does not always mean the full text.
Sourcing does not govern use. What you may do with the texts, meaning display, store, evaluate, pass to third parties or use for model training, sits in the licence and differs sharply by purpose.
The SmPC does not replace professional judgement. It describes the agreed state of knowledge for a medicine. Application in the individual case stays with the professional, and the content addresses healthcare professionals rather than patients.
For single lookups, the regulator portals are enough. As soon as a system processes the content, you need it structured, keyed to a product identifier, with the clinical sections individually addressable and an update cadence set by contract. The fields most often missing are excipients and shelf life after opening. The clauses most often missing are change marking and retention.
If you want to hold a field set against your own specification, book a demo and bring the spec.
This content is addressed to healthcare professionals and does not constitute medical advice. Last reviewed: September 2026.
Through a data licence, either as a recurring delivery or via an API. Regulator portals publish SmPCs as individually retrievable PDFs, which is fine for lookups but offers no field structure and no bulk retrieval. As soon as a system searches, checks or evaluates the content itself, you need the sections individually addressable and a version stamp per record.
Continuously, across the whole lifecycle of a medicine, as new safety data accumulate and indications are added or restricted. That is why every SmPC carries a revision date. A dataset without an agreed update cadence stops being a reference after a few months and becomes a snapshot that people still treat as current.
In three tiers. For identification: the product identifier, brand name, authorisation holder, dosage form, strength, active substances and excipients, plus the ATC code. Clinically: indications, contraindications, interactions, undesirable effects, posology, pregnancy and lactation. For handling: storage and shelf life, including after first opening. The two most often forgotten are excipients and shelf life after opening.
That is governed by the licence and is not automatic. Displaying, storing, evaluating, passing to third parties and using content for model training are separate purposes with separate terms. Technically, two things decide usefulness: whether an answer can be traced to a product, a section and a revision, and whether the underlying dataset is current at runtime.
The SmPC addresses healthcare professionals, the package leaflet addresses patients. The direction is fixed: the leaflet is derived from the SmPC, never the other way round, and it cannot carry information the SmPC does not support. A system that serves both audiences has to license both texts and keep both current.
In a pilot it works; as a permanent arrangement it rarely does. The effort is not in the first pass but in the maintenance: layout changes, inconsistent wording between authorisation holders, new products and amendments without notice. Choosing that route means accepting a permanent obligation that usually has nothing to do with your core business.