Creating medical registry datasets from unstructured text

The project aims to automate the data population process for clinical registries through the use of large language models (LLMs).

Factsheet

  • Schools involved School of Engineering and Computer Science
  • Institute(s) Institute for Patient-centered Digital Health (PCDH)
  • Research unit(s) PCDH / AI for Health
  • Funding organisation Innosuisse
  • Duration 17.06.2024 - 31.12.2025
  • Head of project Prof. Dr. Kerstin Denecke
  • Partner ID Suisse AG
  • Keywords Artificial intelligence, large language model, information extraction

Situation

This project aims to automate the process of data collection for clinical registries through the use of large language models (LLM). There are currently 116 registries represented in the Swiss Forum of Clinical Registries (Forum medizinische Register Schweiz ), which is managed by the Swiss Medical Association (FMH). Registry data is essential for quality assurance (e.g. the implant registry), including the tracking of adverse events and outcomes and the identification of treatment gaps. These and similar use cases require complete, high-quality data that is available in registries. Traditional methods of extracting clinical data from routine data and hospital information systems involve the manual copying and pasting of data, a time-consuming and error-prone process that leads to inconsistent and incomplete data. Our approach aims to automate this process by developing advanced natural language processing (NLP) algorithms that are able to accurately analyse and extract relevant clinical information from unstructured text in medical records.

Course of action

The project evaluated and refined various LLM-based approaches to automatically extract structured information from unstructured clinical records and make the data ready for clinical registries. For this, we developed suitable extraction strategies, evaluated them for quality and optimised them for different registry fields. Taking it a step further, we analysed how easily the methods developed here can scale to other registries and clinical use cases. Finally, we examined technical and organisational aspects of scalability to derive requirements for future productive deployment and long-term ongoing development of the solution.

Result

The project successfully demonstrated the technical feasibility of using LLM-based support to populate clinical registries. The methods developed enabled the automated extraction of relevant clinical information from unstructured texts and its mapping to structured registry fields. The project evaluation confirmed that modern LLMs have great potential to support medical record documentation and are able to identify relevant information with sufficient quality and verifiability.

Looking ahead

The insights gained from this project pave the way for the further development of a scalable, automated registry population solution. Future work will focus on integrating the system into existing documentation and registry processes, extending it to include additional registry types and continuously improving extraction quality. In the long term, there is potential to reduce manual documentation workloads, ensure the real-time accuracy and quality of registry data, and expand the availability of structured clinical information for research, quality assurance and healthcare provision.

This project contributes to the following SDGs

  • 3: Good health and well-being
  • 9: Industry, innovation and infrastructure