文献详情 >Masked cosine similarity predi... 收藏

Masked cosine similarity prediction for self-supervised skeleton-based action representation learning

作者：Ren, Ziliang Liu, Ronggui Qin, Yong Gao, Xiangyang Zhang, Qieshi

作者机构：Dongguan Univ Technol Sch Comp Sci & Technol Dongguan 523808 Guangdong Peoples R China Chinese Acad Sci Shenzhen Inst Adv Technol CAS Key Lab Human Machine Intelligence Synergy Sys Shenzhen 518055 Guangdong Peoples R China

出版物：《PATTERN ANALYSIS AND APPLICATIONS》 (Pattern Anal. Appl.)

年卷期：2025年第28卷第2期

页面：1-17页

核心收录：

学科分类：08[工学] 0812[工学-计算机科学与技术（可授工学、理学学位）]

基　　金：Natural Science Foundation of Guangdong Province Intergovernmental International Scientific and Technological Innovation Cooperation Project of the National Key Research and Development Program [2025YFE0199900] National Natural Science Foundation of China [62376261, U21A20487] 2024A1515011754 2023A1515011307 2022A1515140119

主　　题：Skeleton-based action recognition Self-supervised learning Masked autoencoders

摘要：Skeleton-based human action recognition faces challenges owing to the limited availability of annotated data, which constrains the performance of supervised methods in learning representations of skeleton sequences. To address this issue, researchers have introduced self-supervised learning as a method of reducing the reliance on annotated data. This approach exploits the intrinsic supervisory signals embedded within the data itself. In this study, we demonstrate that considering relative positional relationships between joints, rather than relying on joint coordinates as absolute positional information, yields more effective representations of skeleton sequences. Based on this, we introduce the Masked Cosine Similarity Prediction (MCSP) framework, which takes randomly masked skeleton sequences as input and predicts the corresponding cosine similarity between masked joints. Comprehensive experiments show that the proposed MCSP self-supervised pre-training method effectively learns representations in skeleton sequences, improving model performance while decreasing dependence on extensive labeled datasets. After pre-training with MCSP, a vanilla transformer architecture is employed for fine-tuning in action recognition. The results obtained from six subsets of the NTU-RGB+D 60, NTU-RGB+D 120 and PKU-MMD datasets show that our method achieves significant performance improvements on five subsets. Compared to training from scratch, performance improvements are 9.8%, 4.9%, 13%, 11.5%, and 3.6%, respectively, with top-1 accuracies of 92.9%, 97.3%, 89.8%, 91.2%, and 96.1% being achieved. Furthermore, our method achieves comparable results on the PKU-MMD Phase II dataset, achieving a top-1 accuracy of 51.5%. These results are competitive without the need for intricate designs, such as multi-stream model ensembles or extreme data augmentation. The source code of our MOSP is available at https://***/skyisyourlimit/MCSP.

本地馆藏 | 借阅须知 | 我要预约

已订购，未入库

sda

目录详情 | 试阅读 |

读者评论与其他读者分享你的观点

学校读者

用户名:未登录

我的评分

建议与咨询留下您的常用邮箱和电话号码，以便我们向您反馈解决方案和替代方法

时间限定

文献类型

馆藏选择

核心期刊

语言

文献类型

帮助

文字说明：

检索规则说明：

检索范例：

分类表

所选分类

看过本文的还看了

相关文献

该作者的其他文献

CADAL相关文献

Masked cosine similarity prediction for self-supervised skeleton-based action representation learning

读者评论与其他读者分享你的观点

请选择收藏分类：

建议与咨询 留下您的常用邮箱和电话号码，以便我们向您反馈解决方案和替代方法

时间限定

文献类型

馆藏选择

核心期刊

语言

文献类型

帮助

文字说明：

检索规则说明：

检索范例：

分类表

所选分类

看过本文的还看了

相关文献

该作者的其他文献

CADAL相关文献

Masked cosine similarity prediction for self-supervised skeleton-based action representation learning

读者评论 与其他读者分享你的观点

请选择收藏分类： 新增自定义分类 确定 取消

建议与咨询留下您的常用邮箱和电话号码，以便我们向您反馈解决方案和替代方法

读者评论与其他读者分享你的观点

请选择收藏分类：