Regulatory News

How we screen a week of pharmacovigilance literature

Someone on a PV team reads through a stack of abstracts every week, checking each one against the same criteria: right product, safety signal, valid case. That checklist barely changes from one week to the next. What keeps changing is how much of it there is to get through.

Main problem is volume

MEDLINE and PubMed, the indexes behind most literature surveillance, held 28.9 million citations in fiscal year 2018. By fiscal year 2023 that had grown to 36.6 million, a 26% increase in five years. The pace isn't perfectly even year to year, but the corpus keeps getting bigger, and it applies across every therapeutic area a PV team might be covering.

PubMed citations, cumulative

+26% in 5 years
28.9MFY201830.2MFY201931.6MFY202033.1MFY202134.7MFY202236.6MFY2023

Source: U.S. National Library of Medicine, MEDLINE/PubMed production statistics.

For a literature surveillance process built around a fixed team size, that's the whole problem. The regulatory logic behind screening hasn't gotten any lighter. There's just more of it to apply, every single week, indefinitely.

Extraction blocks

We turned that checking into something an article runs through, rather than something a person repeats from memory. We call each individual check an extraction block. Is it in scope for this product. Does it name a substance or class the client actually monitors. Does it carry a safety signal such as an adverse event or a quality defect. Each of these is a block.

None of these blocks come from a hunch. Each one traces back to a specific line in the regulation behind it, and each one can be switched on or off per client. Every change to that configuration gets logged, so the logic behind a given week's results can be reconstructed later.

What comes out the other end

Every article ends up with a labeled outcome and a written reason, not a flat yes or no. Take a case report on an elderly patient with a documented reporter and product: it clears every block and moves to inclusion. A pharmacokinetic study in rodents doesn't, and gets dropped with the reason logged as preclinical, no clinical application. Something in between, a case missing one of the four validity elements, goes to a person, with the missing element named.

The processing itself, reading the article, running it through the blocks, drafting the classification and the reviewer's notes, runs in seconds. What happens after that is up to the person reviewing it. A well-prepared case can be approved quickly. A harder one takes the time it needs. The part we're automating is the setup, so the reviewer starts from a clearer judgment call instead of a blank abstract.

Why the trail matters more than the speed

The part worth dwelling on isn't the seconds. Anyone can classify an abstract quickly if they're willing to guess. What makes the output usable in a regulated setting is that every classification carries a reason that points back to a specific block, and every block points back to a specific piece of regulation. If an inspector asks why an article was dismissed, the answer isn't "the model thought so." It's the exact block that fired, and who configured it, and when.

That's also why the blocks themselves stay visible. A client's portfolio changes, a new substance gets added, a class gets dropped, and the set of active blocks has to reflect that without anyone rewriting logic from scratch. Turning a block on or off is a recorded change, not a silent one.

The same blocks across four domains

The architecture behind extraction blocks covers four regulated areas: human medicines, veterinary medicines, medical devices, and cosmetics. The blocks differ by domain, a device vigilance block looks nothing like a human PV block, but the underlying structure is the same everywhere: configurable blocks, each one mapped to a specific piece of regulatory text.

If you run literature surveillance every week and want to see how this works for your portfolio, get in touch.