0.5615743110418727DINO mAP50
0.4161519439428504RT-DETR mAP50
+0.14542236709902234DINO advantage
31 / 34classes won by DINO at AP50
Corrected DINO's advantage comes primarily from better confidence/ranking behavior, false-positive and duplicate suppression, broad class performance, and improved localization—not from universally higher low-threshold candidate recall.
| Metric | RT-DETR | DINO |
|---|---|---|
| TP | 3,654 | 3,444 |
| FP | 7,711 | 2,164 |
| Precision | 0.321513 | 0.614123 |
| Recall | 0.566599 | 0.534036 |
| F1 | 0.410239 | 0.571286 |
| Duplicates | 3,287 | 110 |
2,858both correct
586DINO-only correct
796RT-DETR-only correct
2,209both fail
Why RT-DETR still matters
RT-DETR retains meaningful complementary value: 796 RT-only correct GT objects, higher recall at score 0.01, and an advantage on the >20 objects/image subset point to crowded-scene specialization.
Why naive union is undesirable
The unique detections arrive with many model-specific false-positive and duplicate regions. Fusion must recover complementary objects without importing that burden.