Pesquisa | Portal Regional da BVS (teste)

An Instance Segmentation Model for Strawberry Diseases Based on Mask R-CNN.

Afzaal, Usman; Bhattarai, Bhuwan; Pandeya, Yagya Raj; Lee, Joonwhoan.

Sensors (Basel) ; 21(19)2021 Sep 30.

Artigo em Inglês | MEDLINE | ID: mdl-34640893

RESUMO

Plant diseases must be identified at the earliest stage for pursuing appropriate treatment procedures and reducing economic and quality losses. There is an indispensable need for low-cost and highly accurate approaches for diagnosing plant diseases. Deep neural networks have achieved state-of-the-art performance in numerous aspects of human life including the agriculture sector. The current state of the literature indicates that there are a limited number of datasets available for autonomous strawberry disease and pest detection that allow fine-grained instance segmentation. To this end, we introduce a novel dataset comprised of 2500 images of seven kinds of strawberry diseases, which allows developing deep learning-based autonomous detection systems to segment strawberry diseases under complex background conditions. As a baseline for future works, we propose a model based on the Mask R-CNN architecture that effectively performs instance segmentation for these seven diseases. We use a ResNet backbone along with following a systematic approach to data augmentation that allows for segmentation of the target diseases under complex environmental conditions, achieving a final mean average precision of 82.43%.

Assuntos

Fragaria , Processamento de Imagem Assistida por Computador , Humanos , Redes Neurais de Computação , Doenças das Plantas

Music video emotion classification using slow-fast audio-video network and unsupervised feature representation.

Pandeya, Yagya Raj; Bhattarai, Bhuwan; Lee, Joonwhoan.

Sci Rep ; 11(1): 19834, 2021 10 06.

Artigo em Inglês | MEDLINE | ID: mdl-34615904

RESUMO

Affective computing has suffered by the precise annotation because the emotions are highly subjective and vague. The music video emotion is complex due to the diverse textual, acoustic, and visual information which can take the form of lyrics, singer voice, sounds from the different instruments, and visual representations. This can be one reason why there has been a limited study in this domain and no standard dataset has been produced before now. In this study, we proposed an unsupervised method for music video emotion analysis using music video contents on the Internet. We also produced a labelled dataset and compared the supervised and unsupervised methods for emotion classification. The music and video information are processed through a multimodal architecture with audio-video information exchange and boosting method. The general 2D and 3D convolution networks compared with the slow-fast network with filter and channel separable convolution in multimodal architecture. Several supervised and unsupervised networks were trained in an end-to-end manner and results were evaluated using various evaluation metrics. The proposed method used a large dataset for unsupervised emotion classification and interpreted the results quantitatively and qualitatively in the music video that had never been applied in the past. The result shows a large increment in classification score using unsupervised features and information sharing techniques on audio and video network. Our best classifier attained 77% accuracy, an f1-score of 0.77, and an area under the curve score of 0.94 with minimum computational cost.

Assuntos

Emoções , Aprendizado de Máquina , Modelos Teóricos , Música , Gravação em Vídeo/classificação , Bases de Dados Factuais , Curva ROC

Deep-Learning-Based Multimodal Emotion Classification for Music Videos.

Pandeya, Yagya Raj; Bhattarai, Bhuwan; Lee, Joonwhoan.

Sensors (Basel) ; 21(14)2021 Jul 20.

Artigo em Inglês | MEDLINE | ID: mdl-34300666

RESUMO

Music videos contain a great deal of visual and acoustic information. Each information source within a music video influences the emotions conveyed through the audio and video, suggesting that only a multimodal approach is capable of achieving efficient affective computing. This paper presents an affective computing system that relies on music, video, and facial expression cues, making it useful for emotional analysis. We applied the audio-video information exchange and boosting methods to regularize the training process and reduced the computational costs by using a separable convolution strategy. In sum, our empirical findings are as follows: (1) Multimodal representations efficiently capture all acoustic and visual emotional clues included in each music video, (2) the computational cost of each neural network is significantly reduced by factorizing the standard 2D/3D convolution into separate channels and spatiotemporal interactions, and (3) information-sharing methods incorporated into multimodal representations are helpful in guiding individual information flow and boosting overall performance. We tested our findings across several unimodal and multimodal networks against various evaluation metrics and visual analyzers. Our best classifier attained 74% accuracy, an f1-score of 0.73, and an area under the curve score of 0.926.

Assuntos

Aprendizado Profundo , Música , Emoções , Expressão Facial , Redes Neurais de Computação

RESUMO

Assuntos

RESUMO

Assuntos

RESUMO

Assuntos

ENVIAR RESULTADO:

SELEÇÃO DE REFERÊNCIAS

DETALHE DA PESQUISA