Small-Object Enhanced YOLOv8 Model for Visual Grasping Detection in Industrial Manipulators

Authors

  • Junichi Ishikawa Faculty of Information Science and Electrical Engineering, Kyoto University, Kyoto, 606-8501, Japan
  • Chia-wen Wu College of Electrical Engineering and Computer Science, National Taiwan University of Science and Technology, Taipei, 10607, China
  • Hsin-yu Lin College of Electrical Engineering and Computer Science, National Taiwan University of Science and Technology, Taipei, 10607, China

DOI:

https://doi.org/10.64972/jaat.2024v2.324p29e:398-412

Keywords:

YOLOv8, Small Object Detection, Visual Grasping, Feature Fusion, Boundary Attention, Robotic Perception

Abstract

For handling small parts mixed with clamps, trays, labels, and partially visible workpieces, industrial manipulators require dependable optical grasping detectors. Standard YOLOv8 is quick, but its detail is too shallow, and it may lose the ability to detect small things like screws, connectors, washers, and narrow brackets after several downsampling operations. This study proposes a small-object enhanced YOLOv8 model that incorporates boundary-guided attention, grasp-aware localisation refinement, receptive-field balanced feature aggregation, and a high-resolution detection branch. Bin-picking, conveyor sorting, and fixture-loading scenarios have produced a reusable industrial dataset with 18,600 annotated photos and 74,280 object instances. The suggested model outperforms the YOLOv8s baseline in terms of mAP50 (96.4% vs. 91.8%), boosts small-object AP from 72.6% to 84.9%, lowers grasp-center error by 2.4 pixels (from 5.8 pixels to 3.4 pixels), and keeps a single-image inference performance of 7.9 ms on an RTX 4060 GPU. Ablation results show that boundary-guided attention reduces localisation failure close to occluded edges, and the high-resolution branch is responsible for a 5.1 percentage-point boost in small-object AP. In mixed-part situations, robotic verification on a six-axis manipulator increased the first-attempt grasp success rate from 86.3% to 94.1%. As shown in the above results, small-object representation is necessary for practical visual grasping in industrial manipulation and does not depend on the depth or scale of the detector.

Downloads

Published

2024-08-29

How to Cite

Ishikawa, J., Wu, C.- wen, & Lin, H.- yu. (2024). Small-Object Enhanced YOLOv8 Model for Visual Grasping Detection in Industrial Manipulators. Journal of Applied Automation Technologies, 2, 29e:398–412. https://doi.org/10.64972/jaat.2024v2.324p29e:398-412

Issue

Section

Articles