Biomedical Text Simplification Models Trained on Aligned Abstracts and Lay Summaries
| Authors |
|
|---|---|
| Publication date | 2025 |
| Host editors |
|
| Book title | The Thirty-Third Text REtrieval Conference (TREC 2024) |
| Book subtitle | Proceedings |
| Series | NIST Special Publication, SP1329 |
| Event | 33rd Text REtrieval Conference (TREC 2024) |
| Number of pages | 8 |
| Publisher | Gaithersburg, MD: National Institute of Standards and Technology |
| Organisations |
|
| Abstract |
This paper documents the University of Amsterdam’s participation in the TREC 2024 Plain Language Adaptation of Biomedical Abstracts (PLABA) Track. We investigated the effectiveness of text simplification models trained on aligned pairs of sentences in biomedical abstracts and plain language summaries. We participated in Task 2 on Complete Abstract Adaptation and conducted post-submission experiments in Task 1 on Term Replacement. Our main findings are the following. First, we used text simplification models trained on aligned real-world scientific abstracts and plain language summaries. We observed better performance for the context-aware model relative to the sentence-level model. Second, our experiments show the value of training on external corpora and demonstrate very reasonable out-of-domain performance on the PLABA data. Third, more generally, our models are conservative and cautious in gratuitous edits or information insertions. This approach ensures the fidelity of the generated output and limits the risk of overgeneration or hallucination.
|
| Document type | Conference contribution |
| Language | English |
| Published at |
https://trec.nist.gov/pubs/trec33/papers/UAmsterdam.plaba.pdf
(Final published version)
|
| Other links | |
| Downloads |
UAmsterdam.plaba
(Final published version)
|
| Permalink to this page | |
