Automated kitting inspection must determine whether each kit contains the correct parts and quantities even when pose, lighting, background, and plastic packaging change how components appear. This paper presents an ambience-aware visual-inspection method that separates what a component is from how it is seen. A reusable synthetic digital model varies station-feasible pose, camera, lighting, background, and packaging conditions while preserving component identity. For every scenario, a synchronized Bag-to-Clean process creates a packaged image for learning and a matching packaging-free reference for mask generation. Grounding DINO and the Segment Anything Model create initial component masks, which are reviewed before training a YOLO-family instance-segmentation model. The trained model supplies component instances to recipe-aware decision logic that compares observed classes and quantities with the active kit and records evidence for operator review. Evaluation on synthetic and real packaged-component images showed consistently strong localization and segmentation performance. The results support a traceable path from controlled ambience variation and aligned supervision to robust kit-conformance decisions.
Keywords
kitting inspection; scenario learning; synthetic digital modeling; instance segmentation; domain randomization.