Lightweight Image Classification of Raptor Species for Edge Devices: Rare-Species Dataset Expansion via Video Frame Extraction, Knowledge Distillation, and TensorRT Deployment
Quick Answer
This study presents a lightweight raptor species classification system for edge devices, utilizing DINOv2-L to distill MobileNetV4, ViT-Small, and EfficientNet-B0.
Quick Take
The dataset was expanded to 12,519 images, achieving a macro recall of 0.935 with EfficientNet-B0 deployed on NVIDIA Jetson Orin Nano at 313 images/s, significantly improving misclassification rates for White-tailed Eagles.
Key Points
- Expanded dataset from 463 to 2,050 images for Steller's Sea Eagle.
- Achieved macro recall of 0.935 with a three-student ensemble.
- EfficientNet-B0 deployed at 3.19 ms/image on NVIDIA Jetson Orin Nano.
- Misclassification of White-tailed Eagles reduced from 61% to 15%.
- Dataset expansion and teacher re-fine-tuning were key to performance gains.
Paper Resources
📖 Reader Mode
~2 min readAbstract:We investigate lightweight raptor-species classification for real-time edge deployment in wind-turbine collision mitigation. Using DINOv2-L (304M parameters) as a teacher, we distilled three lightweight students (MobileNetV4, ViT-Small, and EfficientNet-B0). To reduce confusion between closely related species, we expanded the dataset to 12,519 images, including an increase in Steller's Sea Eagle images from 463 to 2,050 via video-frame extraction. Under a group split that separates samples at the video- and source-image level to mitigate source leakage at that granularity, the three-student ensemble achieved a macro recall of 0.935 +/- 0.004 over five distillation seeds (0.955 on a conventional image-level split, retaining 97.5% of the teacher's macro recall) with roughly one-eighth as many parameters. On a subset of 1,258 images disjoint from the former training images, White-tailed Eagle recall improved by up to 38.6 percentage points, while the rate at which it was misclassified as the Steller's Sea Eagle decreased from 61% to 15% of errors. TensorRT FP16 deployment of EfficientNet-B0 on an NVIDIA Jetson Orin Nano achieved 3.19 ms/image including host-device transfer (313 images/s), with 99.95% argmax agreement with FP32. In five-seed controlled comparisons, neither distillation (versus CE-only) nor the change from a DINOv2-L to a DINOv3-L teacher yielded a clear ensemble-level improvement; the primary gains stem from the dataset expansion and teacher re-fine-tuning.
| Comments: | 21 pages, 4 figures, 14 tables. English translation of a paper submitted to IPSJ (in Japanese); the Japanese version is the source of record |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Image and Video Processing (eess.IV) |
| ACM classes: | I.4.9; I.5.4; I.2.6 |
| Cite as: | arXiv:2607.26238 [cs.CV] |
| (or arXiv:2607.26238v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2607.26238 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Takeshi Nishikawa [view email]
[v1]
Tue, 28 Jul 2026 20:19:25 UTC (72 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.