Crack segmentation is central to computer vision-based infrastructure inspection. However, deployment remains limited due to scarce pixel-level labels and domain shift across environments. This paper introduces CrackSegFlow, a controllable Flow Matching (FM) synthesis method that renders crack images from masks with pixel-level alignment. The renderer combines topology-preserving mask injection with boundary-gated modulation to maintain thin-structure continuity. Class-conditional FM samples diverse masks, and CrackSegFlow renders paired images. Compared with latent-diffusion-based conditioning, CrackSegFlow achieves higher mask–image adherence and thin-structure topology fidelity. It also samples 7 × faster than pixel-domain semantic diffusion while achieving lower FID. Cracks are further injected onto crack-free backgrounds to diversify confounders and reduce false positives. Across five datasets using the established hybrid CNN–Transformer backbone, average gains are +5.4 mIoU/+5.1 F1 in-domain and +13.0 mIoU/+14.7 F1 cross-domain under target-guided synthesis, corresponding to +45.9%/+34.8% relative improvement. The CSF-50K benchmark is also released, comprising 50,000 image–mask pairs.
Asadi et al. (Mon,) studied this question.