The rapid adoption of collaborative robots in industrial and laboratory environments has increased the demand for intuitive and deployment-efficient human–robot interaction systems. While recent research emphasizes perception-intensive autonomy and dynamic scene reconstruction, such approaches often introduce calibration complexity and extended commissioning effort in structured workspaces. This paper presents a deployment-oriented multimodal human–robot interaction framework designed specifically for organized collaborative environments. The proposed system integrates constrained-vocabulary voice command interpretation with region-of-interest (ROI)-based HSV color verification and predefined spatial mapping, implemented on a Kinova Gen3 collaborative robot using event-driven control. Instead of relying on full pose estimation or continuous scene interpretation, the architecture intentionally limits visual processing to color verification at predefined observation poses, thereby reducing computational overhead and integration effort while maintaining real-time responsiveness. Experimental evaluation under controlled indoor conditions demonstrated 94% speech recognition accuracy with an average latency of 310 ms. The perception module achieved 95% color classification consistency with an average processing time of 68 ms per cycle. Across 50 pick-and-place trials, the system achieved a 92% successful grasp rate with an average task completion time of 8.4 seconds. Comparative testing against joystick-based manual control showed reduced execution variability and improved repeatability for repetitive manipulation tasks. The results demonstrate that aligning system complexity with environmental structure provides a practical and scalable strategy for collaborative robot deployment in laboratories, medical preparation areas, and organized assembly settings.
Deployment-Efficient Multimodal Human–Robot Interaction Using Voice Commands and Color Verification in Structured Workspaces
55 views
6 Downloads