Semantic Textual Similarity (STS) is the task of identifying the semantic correlation between two sentences of the same or different languages. STS is an important task in natural language processing because it has many applications in different domains such as information retrieval, machine translation, plagiarism detection, document categorization, semantic search, and conversational systems. The availability of STS training and evaluation data resources for some languages such as English has led to good performance systems that achieve above 80% correlation with human judgment. Unfortunately, such required STS data resources are not available for many languages like Arabic. To overcome this challenge, this paper proposes three different approaches to generate effective STS Arabic models. The first one is based on evaluating the use of automatic machine translation for English STS data to Arabic to be used in fine-tuning. The second approach is based on the interleaving of Arabic models with English data resources. The third approach is based on fine-tuning the knowledge distillation-based models to boost their performance in Arabic using a proposed translated dataset. With very limited resources consisting of just a few hundred Arabic STS sentence pairs, we managed to achieve a score of 81% correlation, evaluated using the standard STS 2017 Arabic evaluation set. Also, we managed to extend the Arabic models to process two local dialects, Egyptian (EG) and Saudi Arabian (SA), with a correlation score of 77.5% for EG dialect and 76% for the SA dialect evaluated using dialectal conversion from the same standard STS 2017 Arabic set.

Download full-text PDF

Source
http://www.ncbi.nlm.nih.gov/pmc/articles/PMC9371328PMC
http://journals.plos.org/plosone/article?id=10.1371/journal.pone.0272991PLOS

Publication Analysis

Top Keywords

data resources
12
arabic models
12
arabic
10
sts
9
semantic textual
8
textual similarity
8
sts task
8
machine translation
8
resources languages
8
sts data
8

Similar Publications

Lifestyle intervention has proven effective in managing older adults' frailty and mild cognitive impairment issues. What remains unclear is how best to encourage lifestyle changes among older adults with frailty and Mild Cognitive Impairment (MCI). We conducted searches in electronic literature searches such as PubMed, Scopus, Cochrane Reviews, ProQuest, and grey resources to find articles published in English between January 2010 and October 2023.

View Article and Find Full Text PDF

Characterization of Tumor Antigens from Multi-omics Data: Computational Approaches and Resources.

Genomics Proteomics Bioinformatics

January 2025

Center for Epigenetics and Disease Prevention, Institute of Biosciences and Technology, Texas A&M University, Houston, TX 77030, USA.

Tumor-specific antigens, also known as neoantigens, have potential utility in anti-cancer immunotherapy, including immune checkpoint blockade (ICB), neoantigen-specific T cell receptor-engineered T (TCR-T), chimeric antigen receptor T (CAR-T), and therapeutic cancer vaccines (TCVs). After recognizing presented neoantigens, the immune system becomes activated and triggers the death of tumor cells. Neoantigens may be derived from multiple origins, including somatic mutations (single nucleotide variants, insertion/deletions, and gene fusions), circular RNAs, alternative splicing, RNA editing, and polymorphic microbiome.

View Article and Find Full Text PDF

Background: Podcasts are an unconventional method of disseminating information through audio to the masses. They are an emerging portable technology and a valuable resource that provides unlimited access for promoting health among participants. Podcasts related to health care have been used as a source of medical education, but there is a dearth of studies on the use of podcasts as a source of health information.

View Article and Find Full Text PDF

Background: Cardiovascular diseases (CVDs) are the leading cause of death globally. Demographic, behavioral, socioeconomic, health care, and psychosocial variables considered risk factors for CVD are routinely measured in population health surveys, providing opportunities to examine health transitions. Studying the drivers of health transitions in countries where multiple burdens of disease persist (eg, South Africa), compared with countries regarded as models of "epidemiologic transition" (eg, England), can provide knowledge on where best to intervene and direct resources to reduce the disease burden.

View Article and Find Full Text PDF

Exploring drought dynamics has become urgent due to unprecedented climate change. Projections indicate that drought events will become increasingly widespread globally, posing a significant threat to the sustainability of the agricultural sector. This growing challenge has resulted in heightened interest in understanding drought dynamics and their impacts on agriculture.

View Article and Find Full Text PDF

Want AI Summaries of new PubMed Abstracts delivered to your In-box?

Enter search terms and have AI summaries delivered each week - change queries or unsubscribe any time!