Bridging the Deployment Gap: A Decision Oriented Survey of Efficient On-Device Foundation Models for Vision

Authors

DOI:

https://doi.org/10.62411/faith.3048-3719-354

Keywords:

Edge AI, Edge computing, Efficient deep learning, Foundation models, Knowledge distillation, Model compression, On-device AI, Quantization, Sustainable AI, Vision transformers

Abstract

Foundation models have significantly advanced computer vision, yet their deployment on resource-constrained edge devices remains limited by memory, latency, and power constraints. Existing surveys typically discuss efficient architectures, compression methods, or edge AI systems separately, with limited emphasis on deployment-oriented decision making across the full optimization pipeline. This survey addresses that gap by presenting a decision-oriented synthesis of efficient on-device vision foundation models, integrating architectural design, training and adaptation strategies, compression techniques, and deployment optimizations into a unified framework. Through a structured literature review across major scientific databases, we identify several key findings. First, no single efficiency strategy universally dominates; the optimal approach depends strongly on deployment priorities such as memory, latency, power, or accuracy constraints. Second, quantization consistently provides the highest efficiency gain with minimal engineering overhead, making it the most practical first-stage optimization for edge deployment. Third, hybrid CNN-transformer architectures generally achieve superior accuracy-efficiency trade-offs compared to purely convolutional or transformer-based models across heterogeneous hardware platforms. Fourth, combined optimization pipelines, particularly pruning followed by quantization, outperform isolated compression methods and can achieve 8–16 times compression while preserving competitive accuracy. Finally, hardware-aware deployment optimizations such as kernel fusion and hardware-specific delegates are critical for translating theoretical efficiency gains into real-world performance. Beyond summarizing prior work, this survey contributes a practical decision framework that maps deployment constraints to recommended optimization strategies and representative application scenarios, including mobile augmented reality, industrial IoT cameras, and multimodal smart assistants. We further identify emerging research challenges in efficient multimodal foundation models, adaptive edge intelligence, hardware-algorithm co-design, and sustainable green AI for edge deployment. Overall, this survey serves as both a comprehensive reference for researchers and a practical deployment guide for engineers developing on-device vision foundation models.

Downloads

Download data is not yet available.

Author Biographies

Vincent Aduvuku, Muni University

Department of Computer and Information Sciences, Muni University, P.O. Box 725, Arua, Uganda

Geoffrey Andogah, Muni University

Department of Computer and Information Sciences, Muni University, P.O. Box 725, Arua, Uganda

Taban Habibu, Muni University

Department of Computer and Information Sciences, Muni University, P.O. Box 725, Arua, Uganda

References

R. Bommasani et al., “On the Opportunities and Risks of Foundation Models,” ArXiv. Jul. 12, 2022. [Online]. Available: http://arxiv.org/abs/2108.07258

A. A. R. Khan et al., “A survey of the vision transformers and their CNN-transformer based variants,” Artif. Intell. Rev., vol. 56, no. S3, pp. 2917–2970, Dec. 2023, doi: 10.1007/s10462-023-10595-0.

Y. Liu et al., “A Survey of Visual Transformers,” IEEE Trans. Neural Networks Learn. Syst., vol. 35, no. 6, pp. 7478–7498, Dec. 2022, doi: 10.1109/TNNLS.2022.3227717.

A. Dosovitskiy et al., “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,” ArXiv, Jun. 2021, [Online]. Available: http://arxiv.org/abs/2010.11929

Z. Liu et al., “Swin Transformer: Hierarchical Vision Transformer using Shifted Windows,” in 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 2021, pp. 9992–10002. doi: 10.1109/ICCV48922.2021.00986.

Z. Dai, H. Liu, Q. V. Le, and M. Tan, “CoAtNet: Marrying Convolution and Attention for All Data Sizes,” ArXiv. Sep. 15, 2021. [Online]. Available: http://arxiv.org/abs/2106.04803

H. Wu et al., “CvT: Introducing Convolutions to Vision Transformers,” in 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Mar. 2021, pp. 22–31. doi: 10.1109/ICCV48922.2021.00009.

J. Deng, W. Dong, R. Socher, L.-J. Li, Kai Li, and Li Fei-Fei, “ImageNet: A large-scale hierarchical image database,” in 2009 IEEE Conference on Computer Vision and Pattern Recognition, Jun. 2009, pp. 248–255. doi: 10.1109/CVPR.2009.5206848.

O. Russakovsky et al., “ImageNet Large Scale Visual Recognition Challenge,” Int. J. Comput. Vis., vol. 115, no. 3, pp. 211–252, Dec. 2015, doi: 10.1007/s11263-015-0816-y.

A. Vaswani et al., “Attention Is All You Need,” in 31st Conference on Neural Information Processing Systems (NIPS 2017), 2017. [Online]. Available: http://arxiv.org/abs/1706.03762

L. Yuan et al., “Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNet,” in 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 2021, pp. 538–547. doi: 10.1109/ICCV48922.2021.00060.

W. Wang et al., “Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions,” in 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 2021, pp. 548–558. doi: 10.1109/ICCV48922.2021.00061.

T. Xiao, M. Singh, E. Mintun, T. Darrell, P. Dollár, and R. Girshick, “Early Convolutions Help Transformers See Better,” ArXiv. Oct. 25, 2021. [Online]. Available: http://arxiv.org/abs/2106.14881

Y. Cheng, D. Wang, P. Zhou, and T. Zhang, “A Survey of Model Compression and Acceleration for Deep Neural Networks,” ArXiv. Jun. 14, 2020. [Online]. Available: http://arxiv.org/abs/1710.09282

V. Sze, Y.-H. Chen, J. Emer, A. Suleiman, and Z. Zhang, “Hardware for machine learning: Challenges and opportunities,” in 2017 IEEE Custom Integrated Circuits Conference (CICC), Apr. 2017, pp. 1–8. doi: 10.1109/CICC.2017.7993626.

A. Fayyazi, M. Kamal, and M. Pedram, “MARCO: Hardware-Aware Neural Architecture Search for Edge Devices with Multi-Agent Reinforcement Learning and Conformal Prediction Filtering,” in 2026 31st Asia and South Pacific Design Automation Conference (ASP-DAC), Jun. 2025, pp. 880–886. doi: 10.1109/ASP-DAC66049.2026.11420542.

S. Han, H. Mao, and W. J. Dally, “Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding,” ArXiv. Feb. 15, 2016. [Online]. Available: http://arxiv.org/abs/1510.00149

A. Gholami, S. Kim, Z. Dong, Z. Yao, M. W. Mahoney, and K. Keutzer, “A Survey of Quantization Methods for Efficient Neural Network Inference,” in Low-Power Computer Vision, Boca Raton: Chapman and Hall/CRC, 2022, pp. 291–326. doi: 10.1201/9781003162810-13.

A. G. Howard et al., “MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications,” ArXiv, Apr. 2017, [Online]. Available: http://arxiv.org/abs/1704.04861

M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen, “MobileNetV2: Inverted Residuals and Linear Bottlenecks,” in 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2018, pp. 4510–4520. doi: 10.1109/CVPR.2018.00474.

A. Howard et al., “Searching for MobileNetV3,” in 2019 IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 2019, pp. 1314–1324. doi: 10.1109/ICCV.2019.00140.

X. Zhang, X. Zhou, M. Lin, and J. Sun, “ShuffleNet: An Extremely Efficient Convolutional Neural Network for Mobile Devices,” in 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2018, pp. 6848–6856. doi: 10.1109/CVPR.2018.00716.

N. Ma, X. Zhang, H.-T. Zheng, and J. Sun, “ShuffleNet V2: Practical Guidelines for Efficient CNN Architecture Design,” in Computer Vision – ECCV 2018, vol. 11218, V. Ferrari, M. Hebert, C. Sminchisescu, and Y. Weiss, Eds. Cham: Springer International Publishing, 2018, pp. 122–138. doi: 10.1007/978-3-030-01264-9_8.

M. Tan and Q. V. Le, “EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks,” in arXiv, May 2019. [Online]. Available: http://arxiv.org/abs/1905.11946

T. Liang, J. Glossner, L. Wang, S. Shi, and X. Zhang, “Pruning and quantization for deep neural network acceleration: A survey,” Neurocomputing, vol. 461, pp. 370–403, Oct. 2021, doi: 10.1016/j.neucom.2021.07.045.

J. Redmon and A. Farhadi, “YOLOv3: An Incremental Improvement,” ArXiv. Apr. 08, 2018. [Online]. Available: http://arxiv.org/abs/1804.02767

M. Horowitz, “1.1 Computing’s energy problem (and what we can do about it),” in 2014 IEEE International Solid-State Circuits Conference Digest of Technical Papers (ISSCC), Feb. 2014, pp. 10–14. doi: 10.1109/ISSCC.2014.6757323.

G. Menghani, “Efficient Deep Learning: A Survey on Making Deep Learning Models Smaller, Faster, and Better,” ACM Comput. Surv., vol. 55, no. 12, pp. 1–37, Dec. 2023, doi: 10.1145/3578938.

S. Foroutani, N. Fahimian, R. Jalalinejad, M. Hezarkhani, S. Mahmoudi, and B. Gharleghi, “Navigating Knowledge Management Implementation Success in Government Organizations: A type-2 fuzzy approach,” ArXiv. Jun. 18, 2024. [Online]. Available: http://arxiv.org/abs/2406.12345

S. Takahashi et al., “Comparison of Vision Transformers and Convolutional Neural Networks in Medical Image Analysis: A Systematic Review,” J. Med. Syst., vol. 48, no. 1, p. 84, Sep. 2024, doi: 10.1007/s10916-024-02105-8.

S. Grigorescu, B. Trasnea, T. Cocias, and G. Macesanu, “A survey of deep learning techniques for autonomous driving,” J. F. Robot., vol. 37, no. 3, pp. 362–386, Apr. 2020, doi: 10.1002/rob.21918.

Z. Liu, H. Mao, C.-Y. Wu, C. Feichtenhofer, T. Darrell, and S. Xie, “A ConvNet for the 2020s,” in 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2022, pp. 11966–11976. doi: 10.1109/CVPR52688.2022.01167.

A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet classification with deep convolutional neural networks,” Commun. ACM, vol. 60, no. 6, pp. 84–90, May 2017, doi: 10.1145/3065386.

K. Simonyan and A. Zisserman, “Very Deep Convolutional Networks for Large-Scale Image Recognition,” in 3rd International Conference on Learning Representations, Apr. 2015, pp. 1–14. [Online]. Available: http://arxiv.org/abs/1409.1556

K. He, X. Zhang, S. Ren, and J. Sun, “Deep Residual Learning for Image Recognition,” in 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2016, vol. 2016-Decem, pp. 770–778. doi: 10.1109/CVPR.2016.90.

C. Szegedy et al., “Going deeper with convolutions,” in 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2015, pp. 1–9. doi: 10.1109/CVPR.2015.7298594.

C. Szegedy et al., “Rethinking the Inception Architecture for Computer Vision,” in 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2016, pp. 2818–2826. doi: 10.1109/CVPR.2016.308.

Y. Tay, M. Dehghani, D. Bahri, and D. Metzler, “Efficient Transformers: A Survey,” ACM Comput. Surv., vol. 55, no. 6, pp. 1–28, Jun. 2023, doi: 10.1145/3530811.

R. Child, S. Gray, A. Radford, and I. Sutskever, “Generating Long Sequences with Sparse Transformers,” ArXiv, Apr. 2019, [Online]. Available: http://arxiv.org/abs/1904.10509

M. Zaheer et al., “Big Bird: Transformers for Longer Sequences,” ArXiv. Jan. 08, 2021. [Online]. Available: http://arxiv.org/abs/2007.14062

N. Kitaev, Ł. Kaiser, and A. Levskaya, “Reformer: The Efficient Transformer,” ArXiv. Feb. 18, 2020. [Online]. Available: http://arxiv.org/abs/2001.04451

S. Wang, B. Z. Li, M. Khabsa, H. Fang, and H. Ma, “Linformer: Self-Attention with Linear Complexity,” ArXiv, Jun. 2020, [Online]. Available: http://arxiv.org/abs/2006.04768

N. Houlsby et al., “Parameter-Efficient Transfer Learning for NLP,” in Proceedings of the 36th International Conference on Machine Learning, 2019, vol. 97, pp. 2790–2799. [Online]. Available: https://proceedings.mlr.press/v97/houlsby19a.html

V. Sze, Y.-H. Chen, T.-J. Yang, and J. S. Emer, “Efficient Processing of Deep Neural Networks: A Tutorial and Survey,” Proc. IEEE, vol. 105, no. 12, pp. 2295–2329, Dec. 2017, doi: 10.1109/JPROC.2017.2761740.

A. Bochkovskiy, C.-Y. Wang, and H.-Y. M. Liao, “YOLOv4: Optimal Speed and Accuracy of Object Detection,” ArXiv, Apr. 2020, [Online]. Available: http://arxiv.org/abs/2004.10934

V. J. Reddi et al., “MLPerf Mobile Inference Benchmark,” ArXiv, Apr. 2022, [Online]. Available: http://arxiv.org/abs/2012.02328

A. Ignatov et al., “AI Benchmark: Running Deep Neural Networks on Android Smartphones,” in Lecture Notes in Computer Science, 2018, pp. 288–314. doi: 10.1007/978-3-030-11021-5_19.

M. Sabih, A. Karim, J. Wittmann, F. Hannig, and J. Teich, “Hardware/Software Co-Design of RISC-V Extensions for Accelerating Sparse DNNs on FPGAs,” in 2024 International Conference on Field Programmable Technology (ICFPT), Dec. 2024, pp. 01–09. doi: 10.1109/ICFPT64416.2024.11113397.

M. Yan, H. Wang, and S. Venkataraman, “PolyThrottle: Energy-efficient Neural Network Inference on Edge Devices,” ArXiv. Jan. 09, 2024. [Online]. Available: http://arxiv.org/abs/2310.19991

B. Varghese, N. Wang, S. Barbhuiya, P. Kilpatrick, and D. S. Nikolopoulos, “Challenges and Opportunities in Edge Computing,” in 2016 IEEE International Conference on Smart Cloud (SmartCloud), Nov. 2016, pp. 20–26. doi: 10.1109/SmartCloud.2016.18.

L. Lai, N. Suda, and V. Chandra, “CMSIS-NN: Efficient Neural Network Kernels for Arm Cortex-M CPUs,” ArXiv. Jan. 19, 2018. [Online]. Available: http://arxiv.org/abs/1801.06601

R. David et al., “TensorFlow Lite Micro: Embedded Machine Learning on TinyML Systems,” ArXiv. Mar. 13, 2021. [Online]. Available: http://arxiv.org/abs/2010.08678

Qualcomm Technologies Inc., “Snapdragon 8 Series Mobile Platforms,” Qualcomm, 2025. https://www.qualcomm.com/snapdragon/processors/8-series-mobile-platforms

Apple Inc., “A18 Pro Chip,” Apple, 2025. https://www.apple.com/iphone-16-pro/

ARM Ltd., “Arm Ethos-U85 NPU,” Arm.com, 2025. https://www.arm.com/products/silicon-ip-cpu/ethos/ethos-u85

C.-J. Wu et al., “Machine Learning at Facebook: Understanding Inference at the Edge,” in 2019 IEEE International Symposium on High Performance Computer Architecture (HPCA), Feb. 2019, pp. 331–344. doi: 10.1109/HPCA.2019.00048.

X. Dong et al., “CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped Windows,” in 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2022, pp. 12114–12124. doi: 10.1109/CVPR52688.2022.01181.

D. Bolya, C.-Y. Fu, X. Dai, P. Zhang, C. Feichtenhofer, and J. Hoffman, “Token Merging: Your ViT But Faster,” arXiv. Mar. 01, 2023. [Online]. Available: http://arxiv.org/abs/2210.09461

Y. Rao, W. Zhao, B. Liu, J. Lu, J. Zhou, and C.-J. Hsieh, “DynamicViT: Efficient Vision Transformers with Dynamic Token Sparsification,” in 35th Conference on Neural Information Processing Systems (NeurIPS 2021), Oct. 2021. [Online]. Available: http://arxiv.org/abs/2106.02034

Y. Chen, X. Dai, M. Liu, D. Chen, L. Yuan, and Z. Liu, “Dynamic Convolution: Attention Over Convolution Kernels,” in 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2020, pp. 11027–11036. doi: 10.1109/CVPR42600.2020.01104.

B. Yang, G. Bender, Q. V. Le, and J. Ngiam, “CondConv: Conditionally Parameterized Convolutions for Efficient Inference,” ArXiv, Sep. 2020, [Online]. Available: http://arxiv.org/abs/1904.04971

Y. LeCun, J. Denker, and S. Solla, “Optimal Brain Damage,” in Advances in Neural Information Processing Systems, 1989, vol. 2. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/1989/file/6c9882bbac1c7093bd25041881277658-Paper.pdf

Z. Liu, J. Li, Z. Shen, G. Huang, S. Yan, and C. Zhang, “Learning Efficient Convolutional Networks through Network Slimming,” in 2017 IEEE International Conference on Computer Vision (ICCV), Oct. 2017, pp. 2755–2763. doi: 10.1109/ICCV.2017.298.

Y. He, X. Zhang, and J. Sun, “Channel Pruning for Accelerating Very Deep Neural Networks,” in 2017 IEEE International Conference on Computer Vision (ICCV), Oct. 2017, pp. 1398–1406. doi: 10.1109/ICCV.2017.155.

H. Li, A. Kadav, I. Durdanovic, H. Samet, and H. P. Graf, “Pruning Filters for Efficient ConvNets,” ArXiv, Mar. 2017, [Online]. Available: http://arxiv.org/abs/1608.08710

A. Romero, N. Ballas, S. E. Kahou, A. Chassang, C. Gatta, and Y. Bengio, “FitNets: Hints for Thin Deep Nets,” ArXiv. Mar. 27, 2015. [Online]. Available: http://arxiv.org/abs/1412.6550

S. Mehta and M. Rastegari, “MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision Transformer,” ArXiv, Mar. 2022, [Online]. Available: http://arxiv.org/abs/2110.02178

G. Evangelidis et al., “EfficientFormer: Vision Transformers at MobileNet Speed,” in 36th Conference on Neural Information Processing Systems (NeurIPS 2022), Oct. 2022, pp. 12934–12949. doi: 10.52202/068431-0940.

M. Maaz et al., “EdgeNeXt: Efficiently Amalgamated CNN-Transformer Architecture for Mobile Vision Applications,” in Lecture Notes in Computer Science, 2022, pp. 3–20. doi: 10.1007/978-3-031-25082-8_1.

F. N. Iandola, S. Han, M. W. Moskewicz, K. Ashraf, W. J. Dally, and K. Keutzer, “SqueezeNet: AlexNet-level accuracy with 50x fewer parameters and <0.5MB model size,” arXiv. Nov. 04, 2016. [Online]. Available: http://arxiv.org/abs/1602.07360

K. Han, Y. Wang, Q. Tian, J. Guo, C. C. Xu, and C. C. Xu, “GhostNet: More Features From Cheap Operations,” in 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2020, pp. 1577–1586. doi: 10.1109/CVPR42600.2020.00165.

E. J. Hu et al., “LoRA: Low-Rank Adaptation of Large Language Models,” arXiv. Oct. 16, 2021. [Online]. Available: http://arxiv.org/abs/2106.09685

A. Graves, “Adaptive Computation Time for Recurrent Neural Networks,” ArXiv, Feb. 2017, [Online]. Available: http://arxiv.org/abs/1603.08983

T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A Simple Framework for Contrastive Learning of Visual Representations,” ArXiv, Jul. 2020, [Online]. Available: http://arxiv.org/abs/2002.05709

K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick, “Momentum Contrast for Unsupervised Visual Representation Learning,” in 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2020, pp. 9726–9735. doi: 10.1109/CVPR42600.2020.00975.

J.-B. Grill et al., “Bootstrap your own latent a new approach to self-supervised learning,” in Proceedings of the 34th International Conference on Neural Information Processing Systems, 2020.

M. Caron et al., “Emerging Properties in Self-Supervised Vision Transformers,” in 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 2021, pp. 9630–9640. doi: 10.1109/ICCV48922.2021.00951.

K. He, X. Chen, S. Xie, Y. Li, P. Dollar, and R. Girshick, “Masked Autoencoders Are Scalable Vision Learners,” in 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2022, pp. 15979–15988. doi: 10.1109/CVPR52688.2022.01553.

S. Mehta and M. Rastegari, “Separable Self-attention for Mobile Vision Transformers,” ArXiv, Jun. 2022, [Online]. Available: http://arxiv.org/abs/2206.02680

A. Hassani, S. Walton, N. Shah, A. Abuduweili, J. Li, and H. Shi, “Escaping the Big Data Paradigm with Compact Transformers,” ArXiv, Jun. 2022, [Online]. Available: http://arxiv.org/abs/2104.05704

S. N. Wadekar and A. Chaurasia, “MobileViTv3: Mobile-Friendly Vision Transformer with Simple and Effective Fusion of Local, Global and Input Features,” ArXiv, Oct. 2022, [Online]. Available: http://arxiv.org/abs/2209.15159

Y. Li et al., “Rethinking Vision Transformers for MobileNet Size and Speed,” ArXiv. Sep. 04, 2023. [Online]. Available: http://arxiv.org/abs/2212.08059

J. Pan et al., “EdgeViTs: Competing Light-Weight CNNs on Mobile Devices with Vision Transformers,” in Lecture Notes in Computer Science, 2022, pp. 294–311. doi: 10.1007/978-3-031-20083-0_18.

P. K. A. Vasu, J. Gabriel, J. Zhu, O. Tuzel, and A. Ranjan, “FastViT: A Fast Hybrid Vision Transformer using Structural Reparameterization,” ArXiv, Aug. 2023, [Online]. Available: http://arxiv.org/abs/2303.14189

K. Choromanski et al., “Rethinking Attention with Performers,” ArXiv, Nov. 2022, [Online]. Available: http://arxiv.org/abs/2009.14794

Y. Xiong et al., “Nyströmformer: A Nyström-based Algorithm for Approximating Self-Attention,” Proc. AAAI Conf. Artif. Intell., vol. 35, no. 16, pp. 14138–14148, May 2021, doi: 10.1609/aaai.v35i16.17664.

X. Chu et al., “Twins: Revisiting the Design of Spatial Attention in Vision Transformers,” ArXiv, Sep. 2021, [Online]. Available: http://arxiv.org/abs/2104.13840

Z. Kong et al., “SPViT: Enabling Faster Vision Transformers via Soft Token Pruning,” in Lecture Notes in Computer Science, 2022, pp. 620–640. doi: 10.1007/978-3-031-20083-0_37.

X. Chen, Z. Liu, H. Tang, L. Yi, H. Zhao, and S. Han, “SparseViT: Revisiting Activation Sparsity for Efficient High-Resolution Vision Transformer,” in 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2023, pp. 2061–2070. doi: 10.1109/CVPR52729.2023.00205.

K. Li et al., “UniFormer: Unifying Convolution and Self-Attention for Visual Recognition,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 45, no. 10, pp. 12581–12600, Oct. 2023, doi: 10.1109/TPAMI.2023.3282631.

Y. Chen et al., “Mobile-Former: Bridging MobileNet and Transformer,” in 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2022, pp. 5260–5269. doi: 10.1109/CVPR52688.2022.00520.

S. Teerapittayanon, B. McDanel, and H. T. Kung, “BranchyNet: Fast inference via early exiting from deep neural networks,” in 2016 23rd International Conference on Pattern Recognition (ICPR), Dec. 2016, pp. 2464–2469. doi: 10.1109/ICPR.2016.7900006.

H. Yin, A. Vahdat, J. M. Alvarez, A. Mallya, J. Kautz, and P. Molchanov, “A-ViT: Adaptive Tokens for Efficient Vision Transformer,” in 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2022, pp. 10799–10808. doi: 10.1109/CVPR52688.2022.01054.

M. Dehghani, S. Gouws, O. Vinyals, J. Uszkoreit, and Ł. Kaiser, “Universal Transformers,” ArXiv, Mar. 2019, [Online]. Available: http://arxiv.org/abs/1807.03819

M. Tan et al., “MnasNet: Platform-Aware Neural Architecture Search for Mobile,” in 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2019, pp. 2815–2823. doi: 10.1109/CVPR.2019.00293.

H. Cai, C. Gan, T. Wang, Z. Zhang, and S. Han, “Once-for-All: Train One Network and Specialize it for Efficient Deployment,” ArXiv, Apr. 2020, [Online]. Available: http://arxiv.org/abs/1908.09791

B. Lester, R. Al-Rfou, and N. Constant, “The Power of Scale for Parameter-Efficient Prompt Tuning,” in Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 2021, pp. 3045–3059. doi: 10.18653/v1/2021.emnlp-main.243.

M. Jia et al., “Visual Prompt Tuning,” in Lecture Notes in Computer Science, 2022, pp. 709–727. doi: 10.1007/978-3-031-19827-4_41.

G. Hinton, O. Vinyals, and J. Dean, “Distilling the Knowledge in a Neural Network,” ArXiv, Mar. 2015, [Online]. Available: http://arxiv.org/abs/1503.02531

S. Zagoruyko and N. Komodakis, “Paying More Attention to Attention: Improving the Performance of Convolutional Neural Networks via Attention Transfer,” ArXiv. Feb. 12, 2017. [Online]. Available: http://arxiv.org/abs/1612.03928

P. Micikevicius et al., “Mixed Precision Training,” ArXiv. Feb. 15, 2018. [Online]. Available: http://arxiv.org/abs/1710.03740

B. Jacob et al., “Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference,” in 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2018, pp. 2704–2713. doi: 10.1109/CVPR.2018.00286.

S. Xu, Y. Li, M. Wang, C. Liu, and B. Zhang, “Associative Recurrent Bilinear Optimization for Domain-Generalized Binary Neural Networks,” Int. J. Comput. Vis., vol. 134, no. 4, p. 188, Apr. 2026, doi: 10.1007/s11263-026-02762-x.

Y. Cai, Z. Yao, Z. Dong, A. Gholami, M. W. Mahoney, and K. Keutzer, “ZeroQ: A Novel Zero Shot Quantization Framework,” in 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2020, pp. 13166–13175. doi: 10.1109/CVPR42600.2020.01318.

Y. Li et al., “BRECQ: Pushing the Limit of Post-Training Quantization by Block Reconstruction,” arXiv. Jul. 25, 2021. [Online]. Available: http://arxiv.org/abs/2102.05426

Z. Yuan, C. Xue, Y. Chen, Q. Wu, and G. Sun, “PTQ4ViT: Post-training quantization for vision transformers with twin uniform quantization,” arXiv. Jun. 23, 2024. [Online]. Available: http://arxiv.org/abs/2111.12293

P. Michel, O. Levy, and G. Neubig, “Are Sixteen Heads Really Better than One?,” ArXiv. Nov. 04, 2019. [Online]. Available: http://arxiv.org/abs/1905.10650

A. Fan, E. Grave, and A. Joulin, “Reducing Transformer Depth on Demand with Structured Dropout,” ArXiv, Sep. 2019, [Online]. Available: http://arxiv.org/abs/1909.11556

R. Denton, W. Zaremba, J. Bruna, Y. LeCun, and R. Fergus, “Exploiting Linear Structure Within Convolutional Networks for Efficient Evaluation,” arXiv. Jun. 09, 2014. [Online]. Available: http://arxiv.org/abs/1404.0736

T. Dao et al., “FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness,” in Neural Information Processing Systems Foundation, Inc. (NeurIPS), Jun. 2022, pp. 16344–16359. doi: 10.52202/068431-1189.

T. Dao, “FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning,” ArXiv, Jul. 2023, [Online]. Available: http://arxiv.org/abs/2307.08691

T. Chen et al., “TVM: an automated end-to-end optimizing compiler for deep learning,” in Proceedings of the 13th USENIX Conference on Operating Systems Design and Implementation, 2018, pp. 579–594.

PyTorch, “PyTorch Mobile,” PyTorch Documentation, 2025. https://pytorch.org/mobile

Google, “TensorFlow Lite for Mobile and Edge Devices,” Google AI, 2025. https://www.tensorflow.org/lite

MediaTek Inc., “MediaTek Dimensity 9400+,” MediaTek, 2026. https://www.mediatek.com/products/smartphones/mediatek-dimensity-9400-plus

S. Sarkar, S. Kundu, K. Zheng, and P. A. Beerel, “Block Selective Reprogramming for On-device Training of Vision Transformers,” in 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Jun. 2024, pp. 8094–8103. doi: 10.1109/CVPRW63382.2024.00809.

L. Zhu, B. Liao, Q. Zhang, X. Wang, W. Liu, and X. Wang, “Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model,” in Proceedings of the 41st International Conference on Machine Learning, 2024, vol. 235, pp. 62429–62442. [Online]. Available: https://proceedings.mlr.press/v235/zhu24f.html

J. Li, D. Li, S. Savarese, and S. Hoi, “BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models,” in Proceedings of the 40th International Conference on Machine Learning, 2023, vol. 202, pp. 19730–19742. [Online]. Available: https://proceedings.mlr.press/v202/li23q.html

H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual Instruction Tuning,” ArXiv. Apr. 17, 2023. [Online]. Available: http://arxiv.org/abs/2304.08485

W. A. Nunes, A. Vinicius Corrêa Dos Santos, and F. G. Moraes, “Accelerating Machine Learning with RISC-V Vector Extension and Auto-Vectorization Techniques,” in 2025 IEEE International Symposium on Circuits and Systems (ISCAS), May 2025, pp. 1–5. doi: 10.1109/ISCAS56072.2025.11043225.

Y. Wang, R. Huang, S. Song, Z. Huang, and G. Huang, “Not All Images are Worth 16x16 Words: Dynamic Transformers for Efficient Image Recognition,” ArXiv, Oct. 2021, [Online]. Available: http://arxiv.org/abs/2105.15075

L. Liu, Z. Qu, Z. Chen, F. Tu, Y. Ding, and Y. Xie, “Dynamic Sparse Attention for Scalable Transformer Acceleration,” IEEE Trans. Comput., pp. 1–14, 2022, doi: 10.1109/TC.2022.3208206.

R. Schwartz, J. Dodge, N. A. Smith, and O. Etzioni, “Green AI,” Commun. ACM, vol. 63, no. 12, pp. 54–63, Nov. 2019, doi: 10.1145/3381831.

E. Strubell, A. Ganesh, and A. McCallum, “Energy and Policy Considerations for Deep Learning in NLP,” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019, pp. 3645–3650. doi: 10.18653/v1/P19-1355.

A. Lacoste, A. Luccioni, V. Schmidt, and T. Dandres, “Quantifying the Carbon Emissions of Machine Learning,” ArXiv. Nov. 04, 2019. [Online]. Available: http://arxiv.org/abs/1910.09700

P. Henderson, J. Hu, J. Romoff, E. Brunskill, D. Jurafsky, and J. Pineau, “Towards the systematic reporting of the energy and carbon footprints of machine learning,” J. Mach. Learn. Res., vol. 21, no. 1, Jan. 2020.

Downloads

Published

2026-05-16

How to Cite

[1]
V. Aduvuku, G. Andogah, and T. Habibu, “Bridging the Deployment Gap: A Decision Oriented Survey of Efficient On-Device Foundation Models for Vision”, J. Fut. Artif. Intell. Tech., vol. 3, no. 1, pp. 99–131, May 2026.

Issue

Section

Articles

Similar Articles

1 2 3 4 5 6 7 8 9 10 > >> 

You may also start an advanced similarity search for this article.