Strawberry harvesting perception in greenhouse environments requires visual models that remain reliable under occlusion while staying compact enough for edge-side inference. To address this requirement, this study develops StrawPose-Lite, a lightweight pose network for strawberry picking point prediction based on YOLOv11n-pose. The network combines ADown and C3Ghost to reduce redundant computation while preserving informative structure, and it adopts a six-keypoint pose definition derived from strawberry phenotypic characteristics. In this representation, the pedicel–fruit junction is used as the final visual picking point, whereas the remaining peak, curvature, and bottom keypoints provide geometric support when the visible contour is incomplete. The keypoint branch is further enhanced by P2-guided multi-scale fusion and SimAM-based refinement to improve sensitivity to fine pedicel-related cues under strict lightweight constraints. On the public validation split, StrawPose-Lite contains 0.73 M parameters and requires 3.0 GFLOPs while achieving a pose mAP@0.5:0.95 of 79.2%. In the independent field deployment set, the TensorRT INT8 version achieved a pure network inference throughput of 277 FPS on a Jetson Orin NX 16G Super platform, with a measured total software latency of 5.01 ms under the embedded pipeline. These results indicate that StrawPose-Lite provides an effective balance between pose accuracy, model compactness, and edge-side inference speed for strawberry picking point perception on edge devices.
Liu et al. (Thu,) studied this question.