Ensemble Encoder-Enabled Proactive Human Assembly Intention Recognition With Multimodal and Flexible Scale Data
Dongxu Ma et al.
What the paper says
Human-robot collaboration (HRC) assembly necessitates precise mutual cognition to guarantee safe and efficient execution. In this context, human assembly intention recognition (HAIR) serves as a critical approach to achieving this mutual understanding. However, most current HAIR approaches struggle to extract sufficient spatiotemporal information from limited industrial data, particularly under complex conditions like varying scales and visual occlusions. Thereby, this article proposes an ensemble encoder approach to extract and fuse spatial and temporal features from visual and skeleton streams of the HRC assembly process, thus significantly improving HAIR accuracy and efficiency. First, an RGB feature extraction encoder is designed to model spatiotemporal dependencies of the assembly process with different scales of features from flexible input RGB encoders (RGBEs). Distinctively, a cross-attention module is utilized to fuse information from different-scale RGBEs, ensuring comprehensive assembly action representation with different granularities. Second, to address the occlusion challenge, a mask-aware skeleton feature extraction encoder is devised. By utilizing frame and joint masking strategies, it robustly models the relationship between operator pose evolution and assembly actions, maintaining high performance even under occlusion. Third, a global feature fusion encoder integrates and aligns features from RGB and skeleton feature extraction encoders. Experimental results demonstrate the state-of-the-art performance of the proposed approach, which achieves the highest accuracy of 99.12%, 99.23%, and 84.59% on MCV-Intention, HA4M, and HA-VID datasets, respectively. Six ablation studies demonstrate the performance effects of fusion positions, the number of depth channels, cross-attention fusion module, occlusions, illuminations, and computational efficiency.
Evidence weight
Balanced mode · F 0.40 / M 0.15 / V 0.05 / R 0.40
| F · citation impact | 0.50 × 0.4 = 0.20 |
| M · momentum | 0.50 × 0.15 = 0.07 |
| V · venue signal | 0.50 × 0.05 = 0.03 |
| R · text relevance † | 0.50 × 0.4 = 0.20 |
† Text relevance is estimated at 0.50 on the detail page — for your query’s actual relevance score, open this paper from a search result.