Ensemble Encoder-Enabled Proactive Human Assembly Intention Recognition With Multimodal and Flexible Scale Data

Dongxu Ma et al.

IEEE Transactions on Cybernetics2026https://doi.org/10.1109/tcyb.2026.3650942article
AJG 3
Weight
0.50

What the paper says

Human-robot collaboration (HRC) assembly necessitates precise mutual cognition to guarantee safe and efficient execution. In this context, human assembly intention recognition (HAIR) serves as a critical approach to achieving this mutual understanding. However, most current HAIR approaches struggle to extract sufficient spatiotemporal information from limited industrial data, particularly under complex conditions like varying scales and visual occlusions. Thereby, this article proposes an ensemble encoder approach to extract and fuse spatial and temporal features from visual and skeleton streams of the HRC assembly process, thus significantly improving HAIR accuracy and efficiency. First, an RGB feature extraction encoder is designed to model spatiotemporal dependencies of the assembly process with different scales of features from flexible input RGB encoders (RGBEs). Distinctively, a cross-attention module is utilized to fuse information from different-scale RGBEs, ensuring comprehensive assembly action representation with different granularities. Second, to address the occlusion challenge, a mask-aware skeleton feature extraction encoder is devised. By utilizing frame and joint masking strategies, it robustly models the relationship between operator pose evolution and assembly actions, maintaining high performance even under occlusion. Third, a global feature fusion encoder integrates and aligns features from RGB and skeleton feature extraction encoders. Experimental results demonstrate the state-of-the-art performance of the proposed approach, which achieves the highest accuracy of 99.12%, 99.23%, and 84.59% on MCV-Intention, HA4M, and HA-VID datasets, respectively. Six ablation studies demonstrate the performance effects of fusion positions, the number of depth channels, cross-attention fusion module, occlusions, illuminations, and computational efficiency.

Open paper page →

Cite this paper

https://doi.org/https://doi.org/10.1109/tcyb.2026.3650942

Or copy a formatted citation

@article{dongxu2026,
  title        = {{Ensemble Encoder-Enabled Proactive Human Assembly Intention Recognition With Multimodal and Flexible Scale Data}},
  author       = {Dongxu Ma et al.},
  journal      = {IEEE Transactions on Cybernetics},
  year         = {2026},
  doi          = {https://doi.org/https://doi.org/10.1109/tcyb.2026.3650942},
}

Paste directly into BibTeX, Zotero, or your reference manager.

Flag this paper

Ensemble Encoder-Enabled Proactive Human Assembly Intention Recognition With Multimodal and Flexible Scale Data

Flags are reviewed by the Arbiter methodology team within 5 business days.


Evidence weight

0.50

Balanced mode · F 0.40 / M 0.15 / V 0.05 / R 0.40

F · citation impact0.50 × 0.4 = 0.20
M · momentum0.50 × 0.15 = 0.07
V · venue signal0.50 × 0.05 = 0.03
R · text relevance †0.50 × 0.4 = 0.20

† Text relevance is estimated at 0.50 on the detail page — for your query’s actual relevance score, open this paper from a search result.