MapReduce in the Cloud: A Use Case Study for Efficient Co-Occurrence Processing of MEDLINE Annotations with MeSH.

Markus Kreuzthaler Jose Antonio Miñarro-Giménez Stefan Schulz

Stud Health Technol Inform

Institute for Medical Informatics, Statistics and Documentation, Medical University of Graz, Austria.

Published: April 2017

Big data resources are difficult to process without a scaled hardware environment that is specifically adapted to the problem. The emergence of flexible cloud-based virtualization techniques promises solutions to this problem. This paper demonstrates how a billion of lines can be processed in a reasonable amount of time in a cloud-based environment. Our use case addresses the accumulation of concept co-occurrence data in MEDLINE annotation as a series of MapReduce jobs, which can be scaled and executed in the cloud. Besides showing an efficient way solving this problem, we generated an additional resource for the scientific community to be used for advanced text mining approaches.

Download full-text PDF	Source

Publication Analysis

Top Keywords

mapreduce cloud

cloud case

case study

study efficient

efficient co-occurrence

co-occurrence processing

processing medline

medline annotations

annotations mesh

mesh big

Similar Publications

Want AI Summaries of new PubMed Abstracts delivered to your In-box?

Enter search terms and have AI summaries delivered each week - change queries or unsubscribe any time!