Pesquisa | Portal Regional da BVS (teste)

DravidianCodeMix: sentiment analysis and offensive language identification dataset for Dravidian languages in code-mixed text.

Chakravarthi, Bharathi Raja; Priyadharshini, Ruba; Muralidaran, Vigneshwaran; Jose, Navya; Suryawanshi, Shardul; Sherly, Elizabeth; McCrae, John P.

Lang Resour Eval ; 56(3): 765-806, 2022.

Artigo em Inglês | MEDLINE | ID: mdl-35996566

RESUMO

This paper describes the development of a multilingual, manually annotated dataset for three under-resourced Dravidian languages generated from social media comments. The dataset was annotated for sentiment analysis and offensive language identification for a total of more than 60,000 YouTube comments. The dataset consists of around 44,000 comments in Tamil-English, around 7000 comments in Kannada-English, and around 20,000 comments in Malayalam-English. The data was manually annotated by volunteer annotators and has a high inter-annotator agreement in Krippendorff's alpha. The dataset contains all types of code-mixing phenomena since it comprises user-generated content from a multilingual country. We also present baseline experiments to establish benchmarks on the dataset using machine learning and deep learning methods. The dataset is available on Github and Zenodo.

Hope speech detection in YouTube comments.

Chakravarthi, Bharathi Raja.

Soc Netw Anal Min ; 12(1): 75, 2022.

Artigo em Inglês | MEDLINE | ID: mdl-35821874

RESUMO

Recent work on language technology has tried to recognize abusive language such as those containing hate speech and cyberbullying and enhance offensive language identification to moderate social media platforms. Most of these systems depend on machine learning models using a tagged dataset. Such models have been successful in detecting and eradicating negativity. However, an additional study has lately been conducted on the enhancement of free expression through social media. Instead of eliminating ostensibly unpleasant words, we created a multilingual dataset to recognize and encourage positivity in the comments, and we propose a novel custom deep network architecture, which uses a concatenation of embedding from T5-Sentence. We have experimented with multiple machine learning models, including SVM, logistic regression, K-nearest neighbour, decision tree, logistic neighbours, and we propose new CNN based model. Our proposed model outperformed all others with a macro F1-score of 0.75 for English, 0.62 for Tamil, and 0.67 for Malayalam.

Multilingual hope speech detection in English and Dravidian languages.

Chakravarthi, Bharathi Raja.

Int J Data Sci Anal ; 14(4): 389-406, 2022.

Artigo em Inglês | MEDLINE | ID: mdl-35844297

RESUMO

Recent work on language technology has aimed to identify negative language such as hate speech and cyberbullying as well as improve offensive language detection to mediate social media platforms. Most of these systems rely on using machine learning models along with the labelled dataset. Such models have succeeded in identifying negativity and removing it from the platform deleting it. However, recently, more research has been conducted on the improvement of freedom of speech on social media. Instead of deleting supposedly offensive speech, we developed a multilingual dataset to identify hope speech in the comments and promote positivity. This paper presents a multilingual hope speech dataset that promotes equality, diversity and inclusion (EDI) in English, Tamil, Malayalam and Kannada. It was collected to promote positivity and ensure EDI in language technology. Our dataset is unique, as it contains data collected from the LGBTQIA+ community, persons with disabilities and women in science, engineering, technology and management (STEM). We also report our benchmark system results in various machine learning models. We experimented on the Hope Speech dataset for Equality, Diversity and Inclusion (HopeEDI) using different state-of-the-art machine learning models and deep learning models to create benchmark systems.

A Survey of Orthographic Information in Machine Translation.

Chakravarthi, Bharathi Raja; Rani, Priya; Arcan, Mihael; McCrae, John P.

SN Comput Sci ; 2(4): 330, 2021.

Artigo em Inglês | MEDLINE | ID: mdl-34723204

RESUMO

Machine translation is one of the applications of natural language processing which has been explored in different languages. Recently researchers started paying attention towards machine translation for resource-poor languages and closely related languages. A widespread and underlying problem for these machine translation systems is the linguistic difference and variation in orthographic conventions which causes many issues to traditional approaches. Two languages written in two different orthographies are not easily comparable but orthographic information can also be used to improve the machine translation system. This article offers a survey of research regarding orthography's influence on machine translation of under-resourced languages. It introduces under-resourced languages in terms of machine translation and how orthographic information can be utilised to improve machine translation. We describe previous work in this area, discussing what underlying assumptions were made, and showing how orthographic knowledge improves the performance of machine translation of under-resourced languages. We discuss different types of machine translation and demonstrate a recent trend that seeks to link orthographic information with well-established machine translation methods. Considerable attention is given to current efforts using cognate information at different levels of machine translation and the lessons that can be drawn from this. Additionally, multilingual neural machine translation of closely related languages is given a particular focus in this survey. This article ends with a discussion of the way forward in machine translation with orthographic information, focusing on multilingual settings and bilingual lexicon induction.

RESUMO

RESUMO

RESUMO

RESUMO

ENVIAR RESULTADO:

SELEÇÃO DE REFERÊNCIAS

DETALHE DA PESQUISA