Simulation and annotation of global acronyms.

Maxim Filimonov Daphné Chopard Irena Spasić

Bioinformatics

School of Computer Science and Informatics, Cardiff University, Cardiff CF24 4AG, UK.

Published: May 2022

Motivation: Global acronyms are used in written text without their formal definitions. This makes it difficult to automatically interpret their sense as acronyms tend to be ambiguous. Supervised machine learning approaches to sense disambiguation require large training datasets. In clinical applications, large datasets are difficult to obtain due to patient privacy. Manual data annotation creates an additional bottleneck.

Results: We proposed an approach to automatically modifying scientific abstracts to (i) simulate global acronym usage and (ii) annotate their senses without the need for external sources or manual intervention. We implemented it as a web-based application, which can create large datasets that in turn can be used to train supervised approaches to word sense disambiguation of biomedical acronyms.

Availability And Implementation: The datasets will be generated on demand based on a user query and will be downloadable from https://datainnovation.cardiff.ac.uk/acronyms/.

Download full-text PDF	Source
http://www.ncbi.nlm.nih.gov/pmc/articles/PMC9154234	PMC
http://dx.doi.org/10.1093/bioinformatics/btac298	DOI Listing

Publication Analysis

Top Keywords

global acronyms

sense disambiguation

large datasets

simulation annotation

annotation global

acronyms motivation

motivation global

acronyms written

written text

text formal

Similar Publications

Want AI Summaries of new PubMed Abstracts delivered to your In-box?

Enter search terms and have AI summaries delivered each week - change queries or unsubscribe any time!