Pesquisa | Portal Regional da BVS

Multi-Moments in Time: Learning and Interpreting Models for Multi-Action Video Understanding.

Monfort, Mathew; Pan, Bowen; Ramakrishnan, Kandan; Andonian, Alex; McNamara, Barry A; Lascelles, Alex; Fan, Quanfu; Gutfreund, Dan; Feris, Rogerio Schmidt; Oliva, Aude.

IEEE Trans Pattern Anal Mach Intell ; 44(12): 9434-9445, 2022 12.

Artigo em Inglês | MEDLINE | ID: mdl-34752386

RESUMO

Videos capture events that typically contain multiple sequential, and simultaneous, actions even in the span of only a few seconds. However, most large-scale datasets built to train models for action recognition in video only provide a single label per video. Consequently, models can be incorrectly penalized for classifying actions that exist in the videos but are not explicitly labeled and do not learn the full spectrum of information present in each video in training. Towards this goal, we present the Multi-Moments in Time dataset (M-MiT) which includes over two million action labels for over one million three second videos. This multi-label dataset introduces novel challenges on how to train and analyze models for multi-action detection. Here, we present baseline results for multi-action recognition using loss functions adapted for long tail multi-label learning, provide improved methods for visualizing and interpreting models trained for multi-label action detection and show the strength of transferring models trained on M-MiT to smaller datasets.

Assuntos

Algoritmos , Aprendizagem

Edge-Guided Single Depth Image Super Resolution.

Feris, Rogerio Schmidt.

IEEE Trans Image Process ; 25(1): 428-38, 2016 Jan.

Artigo em Inglês | MEDLINE | ID: mdl-26599968

RESUMO

Recently, consumer depth cameras have gained significant popularity due to their affordable cost. However, the limited resolution and the quality of the depth map generated by these cameras are still problematic for several applications. In this paper, a novel framework for the single depth image superresolution is proposed. In our framework, the upscaling of a single depth image is guided by a high-resolution edge map, which is constructed from the edges of the low-resolution depth image through a Markov random field optimization in a patch synthesis based manner. We also explore the self-similarity of patches during the edge construction stage, when limited training data are available. With the guidance of the high-resolution edge map, we propose upsampling the high-resolution depth image through a modified joint bilateral filter. The edge-based guidance not only helps avoiding artifacts introduced by direct texture prediction, but also reduces jagged artifacts and preserves the sharp edges. Experimental results demonstrate the effectiveness of our method both qualitatively and quantitatively compared with the state-of-the-art methods.

RESUMO

Assuntos

RESUMO

ENVIAR RESULTADO:

SELEÇÃO DE REFERÊNCIAS

DETALHE DA PESQUISA