Go to main content

Human Action Recognition (HAR) plays a critical role in a wide range of applications, including intelligent surveillance, human-computer interaction, healthcare monitoring, and sports analytics. Among various modalities, skeleton-based HAR has emerged as a robust paradigm due to its invariance to appearance, lighting, and background changes, as well as its compact and structured representation of human motion. Unlike RGB or depth data that are sensitive to viewpoint variations and environmental noise, skeleton sequences provide a high-level abstraction of human dynamics that facilitates both interpretability and generalization. Recent advances in deep learning have greatly improved recognition accuracy; however, challenges remain in effectively modeling long-range spatio-temporal dependencies, learning from limited annotations, and ensuring generalization to unseen actions while maintaining computational efficiency. This dissertation addresses key challenges through an integrated research framework that advances self-supervised representation learning, frequency-domain modeling, and spectral–temporal fusion strategies. We first tackle the scarcity of labeled data with SkeletonMAE, a spatial–temporal masked autoencoder that incorporates structure-aware masking to learn discriminative motion features from large-scale unlabeled skeleton sequences. To capture richer motion dynamics, FreqMixFormer introduces a frequency-aware mixed-attention transformer that leverages the discrete cosine transform (DCT) to extract complementary low- and high-frequency cues, enabling more nuanced temporal modeling. Building upon this, FreqMixFormerV2 refines the integration of spectral and spatio–temporal information via a streamlined attention design, reducing redundancy while preserving representational expressiveness and improving computational efficiency. Finally, to extend HAR capability to unseen action categories, FS-VAE embeds frequency-guided semantic features into a variational autoencoder framework, aligning latent representations between seen and unseen classes to improve generalization in zero-shot recognition scenarios. Beyond methodological innovations, this dissertation situates skeleton-based HAR in the broader landscape of efficient and scalable action recognition. The proposed models are designed not only to push state-of-the-art accuracy on benchmarks but also to address critical deployment constraints such as data scarcity, domain transferability, and computational overhead. Extensive experiments across large-scale datasets demonstrate that the presented methods consistently achieve superior performance in supervised recognition, self-supervised representation learning, and zero-shot classification, while remaining resource-conscious. Overall, this body of work demonstrates that combining principled representation learning, frequency-domain analysis, and architectural efficiency can yield models that are both accurate and generalizable. The results advance skeleton-based HAR towards scalable, adaptable, and resource-efficient human motion understanding systems with wide-ranging practical applications.

Metric
From
To
Interval
Export
Download Full History