Nawah-VL Grounding — أين الشيء في الصورة؟
Give it an image and an Arabic label; it returns where that thing is. A 50M-parameter finetune of Nawah-VL-50M on blip3-grounding-1m-arabic.
What to expect — read the second column first.
| model | fixed box, image ignored | |
|---|---|---|
| acc@0.5 | 0.498 | 0.262 |
| mean IoU | 0.498 | 0.326 |
| valid box emitted | 0.994 | — |
26% of reference boxes cover more than half the image, so a fixed full-frame box already scores 0.262. The real margin over not looking is about 0.20.
It finds big things and misses small ones. By how much of the frame the object fills: >50% → 1.00, 20–50% → 0.68, 5–20% → 0.27, <5% → 0.11. Below ~20% it is worse than a box that never looks. It also emits one box only, so labels with several referents fail.
Feed it original files where you can — re-saving a small image as JPEG was enough to move one prediction from a correct box to the whole frame.
جرب واحدة من دول
| الصورة | الشيء المطلوب |
|---|