Few-shot counting addresses the problem of counting objects of arbitrary categories for which only a few exemplars are given. Detection-based counters predict, from the predicted bounding boxes, the number of objects of the category specified by the exemplars. A weakness of counters that predict only from color information is a drop in accuracy in scenes where objects blend into their surroundings or densely overlap. We therefore propose GeCoD2, an extension of the currently best-performing detection-based counter GeCo2 with depth information obtained from the same color image by a pretrained monocular depth-estimation model. We also extend the model for zero-shot counting, in which the given exemplars are replaced by directly learnable prototypes. We further add a density-estimation head that predicts the object count as the sum of the predicted density map, so the estimate remains accurate even in dense scenes. The predicted map additionally enables density-guided detection. For datasets in which objects are annotated only with points, we propose a method that automatically derives bounding boxes from the point annotations and thus enables the training of detection-based counters on such datasets. We evaluate GeCoD2 on three datasets: on FSCD-147 it surpasses the published GeCo2 model, on the multi-class MCAC dataset it surpasses the published ABC123 model by 8 %, and in the indiscernible underwater scenes of the IOCfish5K-D dataset it achieves a 6 % lower MAE than GeCo2.
|