Repository logo

Beyond Benchmark Accuracy: A Generalization-Centered Analysis of Deep Learning for Skeleton-Based Human Activity Recognition

aut.relation.articlenumber101030
aut.relation.endpage101030
aut.relation.journalComputer Science Review
aut.relation.startpage101030
aut.relation.volume62
dc.contributor.authorXia, Y
dc.contributor.authorYongchareon, S
dc.contributor.authorLutui, R
dc.contributor.authorSheng, QZ
dc.date.accessioned2026-07-07T22:22:34Z
dc.date.issued2026-07-03
dc.description.abstractDeep learning models for skeleton-based Human Activity Recognition (HAR) routinely exceed 93% accuracy on controlled benchmarks such as NTU RGB+D, yet degrade by 30–60% when confronted with the noise, demographic diversity, and environmental variability of real-world deployment. Existing surveys catalog this progress by network architecture (RNNs, CNNs, GCNs, Transformers) but provide limited guidance on when or why one approach outperforms another under specific deployment constraints. This survey reframes the field through a generalization-centered lens. We identify and quantify six deployment challenges that drive performance degradation: sensor noise and pose estimation errors (46–60 percentage point gap between clean Kinect and RGB-estimated skeletons), open vocabulary and unseen actions (47–63 pp between supervised and zero-shot results), demographic shift (20–40 pp across elderly-inclusive benchmarks), cross-dataset transfer (10–33 pp), annotation scarcity for purely supervised training (10–15 pp), and cross-subject or cross-view variation on a single dataset (3–7 pp). For each challenge, we systematically evaluate how representation paradigms (sequential, spatial, graph, attention, and cross-modal), learning strategies (self-supervised, semi-supervised, few-shot, and domain adaptation), and integration techniques (fusion, distillation, and multi-task) perform, and we provide cross-method comparisons grounded in published results. Our analysis shows that no method effectively addresses more than three challenges, that research effort focuses on the smallest gaps while the largest remain understudied, and that self-supervised learning and CNN-based representations emerge as the most versatile paradigms across multiple deployment scenarios. We synthesize these findings into a method–challenge suitability matrix and a decision framework that links deployment constraints to method recommendations. These recommendations are derived from single-challenge evidence and require validation under simultaneous-challenge deployment conditions. A three-tier research roadmap identifies concrete directions for bridging the generalization gap across near-, mid-, and long-term horizons.
dc.identifier.citationComputer Science Review, ISSN: 1574-0137 (Print), Elsevier BV, 62, 101030-101030. doi: 10.1016/j.cosrev.2026.101030
dc.identifier.doi10.1016/j.cosrev.2026.101030
dc.identifier.issn1574-0137
dc.identifier.urihttp://hdl.handle.net/10292/21560
dc.languageen
dc.publisherElsevier BV
dc.relation.urihttps://www.sciencedirect.com/science/article/pii/S1574013726001383
dc.rights.accessrightsOpenAccess
dc.rights.urihttp://creativecommons.org/licenses/by/4.0/
dc.subject46 Information and Computing Sciences
dc.subject4611 Machine Learning
dc.subjectNetworking and Information Technology R&D (NITRD)
dc.subjectMachine Learning and Artificial Intelligence
dc.subjectBioengineering
dc.subjectComputation Theory & Mathematics
dc.subject46 Information and computing sciences
dc.titleBeyond Benchmark Accuracy: A Generalization-Centered Analysis of Deep Learning for Skeleton-Based Human Activity Recognition
dc.typeJournal Article
pubs.elements-id766809

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Xia et al_2026_Beyond benchmark accuracy.pdf
Size:
2.67 MB
Format:
Adobe Portable Document Format
Description:
Journal article

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.37 KB
Format:
Plain Text
Description: