Tuned Attention-Guided Fusion of Deep and Handcrafted Features for Distinguishing AI-Generated and Human-Created Artworks

Authors

DOI:

https://doi.org/10.62411/faith.3048-3719-291

Keywords:

AI-generated art detection, Attention mechanism, Discrete cosine transform, Feature fusion, Frequency-domain analysis, Interpretability, MobileNetV2, ResNet50

Abstract

Generative AI models can now produce artwork that is virtually indistinguishable from human-created art, posing critical verification challenges for art galleries, educational institutions, and digital content platforms. Existing detection approaches face notable limitations: handcrafted feature–based methods offer interpretability but achieve limited accuracy, while deep learning approaches typically require substantial computational resources and provide minimal explanation for their predictions. We propose an attention-guided fusion framework that integrates Discrete Cosine Transform (DCT)–based frequency-domain features with deep learned representations through a learned attention mechanism, enabling both improved detection performance and interpretable decision-making grounded in signal processing theory. The attention module dynamically weights each feature modality based on input-specific reliability patterns, allowing adaptive fusion across diverse artistic styles and generation methods. We evaluate the proposed framework using two convolutional backbones: the lightweight MobileNetV2 for efficiency-critical deployment scenarios and the higher-capacity ResNet50 for settings where computational resources permit stronger feature extraction. Experiments are conducted on 18,288 artwork images spanning traditional paintings, digital illustrations, and outputs from multiple generative models, using stratified train–validation–test splits and five random seeds for statistical robustness. Under frozen-backbone settings, MobileNetV2-based attention fusion achieves an F1-score of 90.1%, while ResNet50-based attention fusion reaches 90.5%. When backbones are fine-tuned end-to-end, incorporating handcrafted features via attention fusion yields consistent performance gains: MobileNetV2 improves from 93.9% to 94.5% F1-score (+0.6%), and ResNet50 improves from 94.2% to 95.1% F1-score (+0.9%). Feature importance analysis further reveals that low-frequency DCT energy is the most discriminative handcrafted feature, confirming that frequency-domain characteristics effectively distinguish AI-generated from human-created artwork across diverse artistic styles. These results demonstrate that attention-guided fusion of signal processing features and deep learned representations provides consistent accuracy–efficiency benefits across both lightweight and heavyweight architectures, offering a practical and interpretable solution to the emerging challenge of AI-generated artwork detection.

Downloads

Download data is not yet available.

Author Biographies

Vedaan Kumar, Independent Researcher

Independent Researcher, Gurugram, Haryana 122003, India

Apoorva Purohit, Independent Researcher

Independent Researcher, Gurugram, Haryana 122003, India

References

K. Agrawal and R. Banerjee, “Synthetic Art Generation and Deepfake Detection: A Study on Jamini Roy Inspired Dataset,” SSRN. Mar. 29, 2025. doi: 10.2139/ssrn.5358869.

J. Frank, T. Eisenhofer, L. Schönherr, A. Fischer, D. Kolossa, and T. Holz, “Leveraging frequency analysis for deep fake image recognition,” in Proceedings of the 37th International Conference on Machine Learning, 2020. [Online]. Available: https://dl.acm.org/doi/abs/10.5555/3524938.3525242

K. Dong, C. Zhou, Y. Ruan, and Y. Li, “MobileNetV2 Model for Image Classification,” in 2020 2nd International Conference on Information Technology and Computer Application (ITCA), Dec. 2020, pp. 476–480. doi: 10.1109/ITCA52113.2020.00106.

B. Koonce, “ResNet 50,” in Convolutional Neural Networks with Swift for Tensorflow, Berkeley, CA: Apress, 2021, pp. 63–72. doi: 10.1007/978-1-4842-6168-2_6.

H. Farid, “Image forgery detection,” IEEE Signal Process. Mag., vol. 26, no. 2, pp. 16–25, Mar. 2009, doi: 10.1109/MSP.2008.931079.

S. McCloskey and M. Albright, “Detecting GAN-Generated Imagery Using Saturation Cues,” in 2019 IEEE International Conference on Image Processing (ICIP), Sep. 2019, pp. 4584–4588. doi: 10.1109/ICIP.2019.8803661.

F. Marra, D. Gragnaniello, D. Cozzolino, and L. Verdoliva, “Detection of GAN-Generated Fake Images over Social Networks,” in 2018 IEEE Conference on Multimedia Information Processing and Retrieval (MIPR), Apr. 2018, pp. 384–389. doi: 10.1109/MIPR.2018.00084.

R. Corvi, D. Cozzolino, G. Poggi, K. Nagano, and L. Verdoliva, “Intriguing properties of synthetic images: from generative adversarial networks to diffusion models,” in 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), Jun. 2023, pp. 973–982. doi: 10.1109/CVPRW59228.2023.00104.

Y. Zhang, S. Huang, L. Huangfu, and D. Dajun Zeng, “Learning Feature Exploration and Selection With Handcrafted Features for Few-Shot Learning,” IEEE Trans. Syst. Man, Cybern. Syst., vol. 55, no. 4, pp. 2599–2610, Apr. 2025, doi: 10.1109/TSMC.2024.3524390.

D. V. Fevralev, N. N. Ponomarenko, V. V. Lukin, S. K. Abramov, K. O. Egiazarian, and J. T. Astola, “Efficiency analysis of DCT-based filters for color image database,” in Proceedings Volume 7870, Image Processing: Algorithms and Systems IX, Feb. 2011, p. 78700R. doi: 10.1117/12.871944.

F. Franzen, “Image Classification in the Frequency Domain with Neural Networks and Absolute Value DCT,” in Image and Signal Processing, 2018, pp. 301–309. doi: 10.1007/978-3-319-94211-7_33.

X. Zhang, S. Karaman, and S.-F. Chang, “Detecting and Simulating Artifacts in GAN Fake Images,” in 2019 IEEE International Workshop on Information Forensics and Security (WIFS), Dec. 2019, pp. 1–6. doi: 10.1109/WIFS47025.2019.9035107.

J. Johnson, A. Alahi, and L. Fei-Fei, “Perceptual Losses for Real-Time Style Transfer and Super-Resolution,” in Computer Vision – ECCV 2016, 2016, pp. 694–711. doi: 10.1007/978-3-319-46475-6_43.

S. Kan, Y. Cen, Z. He, Z. Zhang, L. Zhang, and Y. Wang, “Supervised Deep Feature Embedding With Handcrafted Feature,” IEEE Trans. Image Process., vol. 28, no. 12, pp. 5809–5823, Dec. 2019, doi: 10.1109/TIP.2019.2901407.

F. Martin-Rodriguez, R. Garcia-Mojon, and M. Fernandez-Barciela, “Detection of AI-Created Images Using Pixel-Wise Feature Extraction and Convolutional Neural Networks,” Sensors, vol. 23, no. 22, p. 9037, Nov. 2023, doi: 10.3390/s23229037.

A. Khan et al., “A survey of the vision transformers and their CNN-transformer based variants,” Artif. Intell. Rev., vol. 56, no. S3, pp. 2917–2970, Dec. 2023, doi: 10.1007/s10462-023-10595-0.

S.-Y. Wang, O. Wang, R. Zhang, A. Owens, and A. A. Efros, “CNN-Generated Images Are Surprisingly Easy to Spot… for Now,” in 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2020, pp. 8692–8701. doi: 10.1109/CVPR42600.2020.00872.

R. Tolosana, R. Vera-Rodriguez, J. Fierrez, A. Morales, and J. Ortega-Garcia, “Deepfakes and beyond: A Survey of face manipulation and fake detection,” Inf. Fusion, vol. 64, pp. 131–148, Dec. 2020, doi: 10.1016/j.inffus.2020.06.014.

N. Yu, L. Davis, and M. Fritz, “Attributing Fake Images to GANs: Learning and Analyzing GAN Fingerprints,” in 2019 IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 2019, pp. 7555–7565. doi: 10.1109/ICCV.2019.00765.

U. Ojha, Y. Li, and Y. J. Lee, “Towards Universal Fake Image Detectors that Generalize Across Generative Models,” in 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2023, pp. 24480–24489. doi: 10.1109/CVPR52729.2023.02345.

M. Groh, Z. Epstein, C. Firestone, and R. Picard, “Deepfake detection by human crowds, machines, and machine-informed crowds,” Proc. Natl. Acad. Sci., vol. 119, no. 1, Jan. 2022, doi: 10.1073/pnas.2110013119.

W. Fang, F. Zhang, V. S. Sheng, and Y. Ding, “A Method for Improving CNN-Based Image Recognition Using DCGAN,” Comput. Mater. Contin., vol. 57, no. 1, pp. 167–178, 2018, doi: 10.32604/cmc.2018.02356.

L. R. Zuama, D. R. I. M. Setiadi, A. Susanto, S. Santosa, H.-S. Gan, and A. A. Ojugo, “High-Performance Face Spoofing Detection using Feature Fusion of FaceNet and Tuned DenseNet201,” J. Futur. Artif. Intell. Technol., vol. 1, no. 4, pp. 385–400, Feb. 2025, doi: 10.62411/faith.3048-3719-62.

H. Yu and B. Xu, “Multi-modal texture fusion network for detecting AI-generated images,” Front. Artif. Intell., vol. 8, Oct. 2025, doi: 10.3389/frai.2025.1663292.

A. Vaswani et al., “Attention Is All You Need,” in 31st Conference on Neural Information Processing Systems (NIPS 2017), Jun. 2017, vol. 30. [Online]. Available: http://arxiv.org/abs/1706.03762

D. Bahdanau, K. Cho, and Y. Bengio, “Neural Machine Translation by Jointly Learning to Align and Translate,” in 3rd International Conference on Learning Representations, ICLR 2015 - Conference Track Proceedings, Sep. 2014, pp. 1–15. [Online]. Available: http://arxiv.org/abs/1409.0473

N. Tinago, S. F. Verkijika, and K. Eva Mamabolo, “Deepfakes in Visual Art: Differentiating AI-Generated Art From Human Art Using Convolutional Neural Networks (CNN),” IEEE Access, vol. 13, pp. 141484–141495, 2025, doi: 10.1109/ACCESS.2025.3596882.

A. Mahara and N. Rishe, “Methods and Trends in Detecting AI-Generated Images: A Comprehensive Review,” arXiv. Oct. 17, 2025. [Online]. Available: http://arxiv.org/abs/2502.15176

S. Mavali, J. Ricker, D. Pape, A. Fischer, and L. Schönherr, “Adversarial Robustness of AI-Generated Image Detectors in the Real World,” arXiv. Jun. 03, 2025. [Online]. Available: http://arxiv.org/abs/2410.01574

L. Nataraj et al., “Detecting GAN generated Fake Images using Co-occurrence Matrices,” arXiv. Oct. 03, 2019. [Online]. Available: http://arxiv.org/abs/1903.06836

Y. Li, X. Yang, P. Sun, H. Qi, and S. Lyu, “Celeb-DF: A Large-Scale Challenging Dataset for DeepFake Forensics,” in 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2020, pp. 3204–3213. doi: 10.1109/CVPR42600.2020.00327.

H. H. Nguyen, J. Yamagishi, and I. Echizen, “Capsule-forensics: Using Capsule Networks to Detect Forged Images and Videos,” in ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), May 2019, pp. 2307–2311. doi: 10.1109/ICASSP.2019.8682602.

R. Bello-Cerezo, F. Bianconi, F. Di Maria, P. Napoletano, and F. Smeraldi, “Comparative Evaluation of Hand-Crafted Image Descriptors vs. Off-the-Shelf CNN-Based Features for Colour Texture Classification under Ideal and Realistic Conditions,” Appl. Sci., vol. 9, no. 4, p. 738, Feb. 2019, doi: 10.3390/app9040738.

L. K. Pavithra and T. S. Sharmila, “An efficient framework for image retrieval using color, texture and edge features,” Comput. Electr. Eng., vol. 70, pp. 580–593, Aug. 2018, doi: 10.1016/j.compeleceng.2017.08.030.

J. Ma, X. Jiang, A. Fan, J. Jiang, and J. Yan, “Image Matching from Handcrafted to Deep Features: A Survey,” Int. J. Comput. Vis., vol. 129, no. 1, pp. 23–79, Jan. 2021, doi: 10.1007/s11263-020-01359-2.

L. Fei-Fei, J. Deng, and K. Li, “ImageNet: Constructing a large-scale image database,” J. Vis., vol. 9, no. 8, pp. 1037–1037, Mar. 2010, doi: 10.1167/9.8.1037.

O. Russakovsky et al., “ImageNet Large Scale Visual Recognition Challenge,” Int. J. Comput. Vis., vol. 115, no. 3, pp. 211–252, Dec. 2015, doi: 10.1007/s11263-015-0816-y.

M. Sandler, A. Howard, M. Zhu, A. Zhmoginov, and L.-C. Chen, “MobileNetV2: Inverted Residuals and Linear Bottlenecks,” in 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2018, pp. 4510–4520. doi: 10.1109/CVPR.2018.00474.

D. Hussain, M. Ismail, I. Hussain, R. Alroobaea, S. Hussain, and S. S. Ullah, “Face Mask Detection Using Deep Convolutional Neural Network and MobileNetV2‐Based Transfer Learning,” Wirel. Commun. Mob. Comput., vol. 2022, no. 1, Jan. 2022, doi: 10.1155/2022/1536318.

K. He, X. Zhang, S. Ren, and J. Sun, “Deep Residual Learning for Image Recognition,” in 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2016, vol. 2016-Decem, pp. 770–778. doi: 10.1109/CVPR.2016.90.

A. Dosovitskiy et al., “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,” arXiv, Jun. 2021, [Online]. Available: http://arxiv.org/abs/2010.11929

Downloads

Published

2026-01-04

How to Cite

[1]
V. Kumar and A. Purohit, “Tuned Attention-Guided Fusion of Deep and Handcrafted Features for Distinguishing AI-Generated and Human-Created Artworks”, J. Fut. Artif. Intell. Tech., vol. 2, no. 4, pp. 629–647, Jan. 2026.

Issue

Section

Articles

Similar Articles

1 2 3 4 5 6 7 8 9 > >> 

You may also start an advanced similarity search for this article.