Human-in-the-Loop Perception: Lightweight Segmentation for Low-Cost Robotic Systems
DOI:
https://doi.org/10.54097/t15b4377Keywords:
Human-in-the-Loop (HITL), Lightweight Segmentation, Low-Cost Robotics, Restaurant Robot, Interactive Perception, CAPTCHA Learning, Cost-Efficient AI, Human–Machine Collaboration.Abstract
A lightweight Human-in-the-Loop (HITL) perception framework is proposed in this paper to improve the segmentation accuracy of low-cost service robots operating in complex real-world environments, such as restaurants. Instead of using high-cost sensors and many training rounds, prompt-based segmentation with SAM-style embeddings can be adopted to reduce human intervention to an extent, and only one or two clicks are needed for real-time mask refinement. We simulate human feedback via click-based heuristics and evaluate the model on the COCO 2017 validation set under low-resolution constraints. Experiments have shown that the first 1-3 clicks provide a significant increase in accuracy (IoU +27% relative), and then it reaches a plateau; therefore, a small number of human interventions can be used instead of hardware upgrades while maintaining high efficiency and low latency. Therefore, HITL perception offers a relatively low-cost and scalable path to a robust deployment of low-cost robots. The framework has also supported the application of Chinese restaurant robots in practice and provided a feasible direction for the development of service robots.
References
[1] Li, W.; Shi, Y.; Yang, W.; Wang, H.; Gao, Y, X. “Interactive image segmentation via cascaded metric learning.” IEEE International Conference on Image Processing (ICIP), 2015, pp. 2900–2904.
[2] Xu, N., Price, B., Cohen, S., Yang, J., & Huang, T. “Deep Interactive Object Selection.” Proceedings of IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 373–381.
[3] Sofiiuk, K., Petrov, I., Barinova, O., & Konushin, A. “f-BRS: Rethinking backpropagating refinement for interactive segmentation.” Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 8623–8632.
[4] Kirillov, A., Mintun, E., Ravi, N., Mao, H., Rolland, C., Gustafson, L., Xiao, T., Whitehead, S., Berg, A., Lo, W., Dollár, P., & Girshick, R. “Segment Anything.” Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023, pp. 4015–4026.
[5] Zhang, Y., & Zhu, J. “Human-in-the-Loop Learning for Interactive AI Systems: A Survey and Outlook.” ACM Computing Surveys, 2022, 55(8), Article 164.
[6] Goodfellow, I., Bengio, Y., & Courville, A. Deep Learning. MIT Press, 2016 — Chapter 20: Semi-Supervised and Human-in-the-Loop Learning.
[7] Fang, W., Chen, C., & Li, Z. “Low-Cost Visual Perception for Service Robots in Complex Environments.” IEEE International Conference on Robotics and Automation (ICRA), 2021, pp. 8427–8434.
[8] von Ahn, L., Maurer, B., McMillen, C., Abraham, D., & Blum, M. “reCAPTCHA: Human-Based Character Recognition via Web Security Measures.” Science, 2008, 321(5895), 1465–1468.
[9] Wang, X., Zhang, Y., & Yu, C. “Adaptive Human–Robot Collaboration in Manufacturing: A Review of Perception and Learning Mechanisms.” IEEE Transactions on Industrial Informatics, 2020, 16(11), 7265–7276.
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.







