Manako Vision AI Benchmark
SN44 (subnet-trained) vs SAM3 (Meta) vs Roboflow
Person Detection + Vehicle Detection • Independent ground truth • No data leakage
Detection Accuracy
| Metric | SN44 | SAM3 | Roboflow |
| mAP@50 | 84.1% | 97.3% | 33.5% |
| Precision | 95.8% | 99.0% | 94.2% |
| Recall | 72.4% | 82.0% | 27.3% |
| mAP@50 | 79.4% | 98.2% | 7.8% |
| Car | 77.7% | 98.7% | 13.3% |
| Bus | 90.1% | 98.0% | 9.2% |
| Truck | 63.4% | 98.3% | 8.6% |
| Motorcycle | 86.6% | 97.7% | 0.0% |
Compute Efficiency
| Metric | SN44 | SAM3 | Roboflow |
| Model Size | 19 MB | 3,450 MB | — |
| Parameters | 4.4M | 848M | — |
| FLOPs per image | 31–52 GFLOPs | ~2,900 GFLOPs | — |
| GPU Required | No (runs on CPU) | Yes (24GB+ VRAM) | Cloud API |
| Cost per 1,000 images | ≈0 (CPU) | $0.06 (GPU) | $0.04 (API) |
Vision Efficiency Score
E = mAP / S
E = Vision Efficiency Score • mAP = mean Average Precision • S = model size (MB)
| Model | Person E | Vehicle E | Combined E | Rank |
| SN44 | 4.47 | 4.03 | 4.25 | #1 |
| Roboflow | — | — | — | — |
| SAM3 | 0.0282 | 0.0285 | 0.0283 | #3 |
150x
more efficient
than SAM3
4.1x
more efficient
than Roboflow
19 MB
model size
runs on any CPU
84%
avg accuracy vs SAM3
at 0.5% of the size
Methodology: Person detection: 200 images, 5,883 annotations (manak0/Detect-Person-winner).
Vehicle detection: 200 images, 1,792 annotations, 3 classes (manak0/Detect-detect-vehicle-winner).
All evaluated against independently generated ground truth — no model was involved in GT creation.
SAM3: facebook/sam3, 848M params. Roboflow: cctv-naxyo/1 (person), vehicles-q0x2v/1 (vehicle).
Vision Efficiency Score E = mAP / S rewards models that maximize accuracy per megabyte of deployment cost.