Device-First Feedback: Toward Mobile-Native LLM-Driven Neural Architecture Search
Quick Answer
The study presents an automated mobile deployment pipeline for LLM-generated CNNs, achieving a 25.6x improvement in mobile deployment score on CIFAR-10.
Quick Take
However, while GPU accuracy improved in later cycles, it did not translate to better mobile performance, particularly on CIFAR-100, highlighting the need for multi-dataset on-device testing.
Key Points
- Automated pipeline integrates QLoRA fine-tuning, GPU evaluation, and on-device benchmarking.
- CIFAR-10 shows a 25.6x improvement in mobile deployment score with 46.9% mean quantized accuracy.
- Later cycles improved GPU accuracy but failed to enhance mobile performance.
- CIFAR-100's best mobile score was retained by the pre-QLoRA baseline.
- Study emphasizes the importance of multi-dataset testing for deployment objectives.
DeepSignal Analysis
What happened
The study introduces an automated pipeline for deploying CNNs generated by large language models on mobile devices. It reports a significant improvement in mobile deployment scores on the CIFAR-10 dataset, but later cycles did not yield better mobile performance on CIFAR-100 despite improved GPU accuracy.
Key evidence
- The automated pipeline integrates QLoRA fine-tuning, GPU evaluation, INT8 export, and physical-device benchmarking for mobile deployment.
- On CIFAR-10, the first cycle achieved a 25.6x improvement in mobile deployment score, with a mean quantized accuracy of 46.9%.
- For CIFAR-100, the pre-QLoRA baseline maintained the best mobile score, while later cycles improved GPU accuracy but did not enhance on-device performance.
Why it matters
This research highlights the complexities of deploying AI models on mobile devices, emphasizing that GPU performance does not guarantee effective mobile deployment. The findings suggest a need for rigorous on-device testing across multiple datasets to ensure models perform well in real-world scenarios, particularly for challenging classification tasks.
What to watch
Paper Resources
📖 Reader Mode
~2 min readAbstract:Deploying convolutional neural networks generated by large language models (LLMs) on real mobile hardware requires more than GPU validation accuracy: INT8 TensorFlow Lite export, delegate selection, and on-device latency jointly determine whether a model is usable. We present an automated mobile deployment pipeline that closes the loop from QLoRA fine-tuning of an architecture-generating LLM through GPU evaluation, INT8 export, and physical-device benchmarking to gated augmentation of the training corpus. The pipeline is fully scripted and runs cycle-by-cycle without manual intervention, with resume support after interruptions. We evaluate the same frozen protocol on two benchmarks, CIFAR-10 and CIFAR-100, on a Samsung SM-P613 tablet (seed 42, 20 models per cycle, cycles 0-6). On CIFAR-10, cycle 1 is gate-accepted and improves the mobile deployment score approximately 25.6x over the baseline with a mean quantized accuracy of 46.9%; later cycles raise GPU accuracy but fail the non-decreasing mobile gate. On CIFAR-100, the pre-QLoRA baseline retains the best mobile score; iterative rounds improve GPU accuracy (up to 26.2%) yet cannot surpass cycle 0 on-device, and the training pool stalls at 19 examples after the first accepted round. Together, the two studies show that closed-loop GPU fine-tuning does not guarantee monotonic mobile gains, especially on harder classification tasks, and that multi-dataset, on-device measurement is needed to stress-test deployment objectives. We release per-cycle metrics with 95% confidence intervals, all figures, and complete reproduction commands.
| Comments: | 11 pages, 4 figures, 4 tables. Code: this https URL |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2608.00078 [cs.CV] |
| (or arXiv:2608.00078v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2608.00078 arXiv-issued DOI via DataCite |
Submission history
From: Saif U-Din [view email]
[v1]
Wed, 29 Jul 2026 14:36:45 UTC (859 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.