NFDI4DS | UHH-SEMS - Publication Details

Latent space improved masked reconstruction model for human skeleton-based action recognition

self-supervised learning masked reconstruction model vector quantized variational autoencoder variational autoencoder Neurosciences. Biological psychiatry. Neuropsychiatry human skeleton-based action recognition RC321-571

DOI: 10.3389/fnbot.2024.1482281 Publication Date: 2025-02-12T07:31:10Z

Abstract Supplemental Material References Cited by

AUTHORS (5)

Enqing Chen

Xueting Wang

Xin Guo

Ying Zhu

Dong Li

ABSTRACT

Human skeleton-based action recognition is an important task in the field of computer vision. In recent years, masked autoencoder (MAE) has been used in various fields due to its powerful self-supervised learning ability and has achieved good results in masked data reconstruction tasks. However, in visual classification tasks such as action recognition, the limited ability of the encoder to learn features in the autoencoder structure results in poor classification performance. We propose to enhance the encoder's feature extraction ability in classification tasks by leveraging the latent space of variational autoencoder (VAE) and further replace it with the latent space of vector quantized variational autoencoder (VQVAE). The constructed models are called SkeletonMVAE and SkeletonMVQVAE, respectively. In SkeletonMVAE, we constrain the latent variables to represent features in the form of distributions. In SkeletonMVQVAE, we discretize the latent variables. These help the encoder learn deeper data structures and more discriminative and generalized feature representations. The experiment results on the NTU-60 and NTU-120 datasets demonstrate that our proposed method can effectively improve the classification accuracy of the encoder in classification tasks and its generalization ability in the case of few labeled data. SkeletonMVAE exhibits stronger classification ability, while SkeletonMVQVAE exhibits stronger generalization in situations with fewer labeled data.

SUPPLEMENTAL MATERIAL

Coming soon ....

REFERENCES (43)

CITATIONS (0)

EXTERNAL LINKS

CROSSREF - Publications OPENAIRE - Products

PlumX Metrics

Latent space improved masked reconstruction model for human skeleton-based action recognition

RECOMMENDATIONS

FAIR ASSESSMENT

Coming soon ....

JUPYTER LAB

Coming soon ....