• e - ISSN No : 2832-4277
IJRTTE Logo

INTERNATIONAL JOURNAL OF RECENT TRENDS IN TECHNOLOGY AND ENGINEERING (IJRTTE)

Multi-Modal Deepfake Detection Using Cross-Frequency Patterns

S Muthuselvan
Professor, Department of Information Technology, KCG College of Technology, India.
B Yamini
Assistant Professor, Department of Computer Science and Engineering, Vels Institute of Science Technology and Advanced Studies, India.

Keywords: Deepfake Detection, Multimodal Learning, Cross-Frequency Analysis, Audio–Visual Forensics, Frequency-Domain Features, Spectral Consistency, Media Authentication

Abstract

With the swift development of deep generative models, it is now possible to produce extremely realistic synthetic audio-visual content often referred to as deepfakes that is posing significant risks to digital trust, security and media authenticity. Despite the reported significant progress in recent deepfake detection algorithms, most of the existing systems are densely based on the spatial or temporal characteristics and cannot to generalize against state-of-the-art generative models, particularly diffusion-based methods. In addition, existing multimodal models do not address much about frequency-domain inconsistency and inter-modal spectral associations that occur during media manipulation. In order to curb such constraints, the present paper suggests a new cross-frequency multi-modal deepfake detector that jointly trains based on audio and visual cues in the frequency domain. The suggested approach breaks down both modalities into multi-band spectral feature and trains the cross-frequency associations between the respective audio and visual elements. The framework manages to capture minute manipulation artifacts that are normally invisible on the space-wise domain by modelling inter-modal spectral alignment with the aid of a cross-frequency correlation and attention mechanism. The results of extensive experiments developed on several benchmark datasets (FF++, Celeb-DF, DFDC, FakeAVCeleb, and WaveFake) show that the presented approach is better than the existing unimodal and multimodal ones in terms of accuracy, robustness, and generalization. The findings confirm that cross frequency reasoning offers a robust and resilient cue when next generation deep fake detection is needed especially when compression, noise, and invisible manipulating are involved.
Download Certificate
Details

References

  1. Salvi, D., Liu, H., Mandelli, S., Bestagini, P., Zhou, W., Zhang, W., & Tubaro, S. (2023). A Robust Approach to Multimodal Deepfake Detection. Journal of Imaging, 9(6), 122. https://doi.org/10.3390/jimaging9060122
  2. Almohawes, M. (2025). Deepfake video detection methods, approaches, and challenges. Alexandria Engineering Journal, 125, 265–277.
  3. Nie, S., Qiao, T., Li, S., Zhang, X., Zhou, J., & Feng, G. (2026). Deepfake detection in the AIGC era: A survey, benchmarks, and future perspectives. Information Fusion, 127(Part A), 103740. https://doi.org/10.1016/j.inffus.2025.103740
  4. Kaur, A., Nawi, N. M., Subramaniam, V., & others. (2024). Deepfake video detection: Challenges and opportunities. Artificial Intelligence Review, 57, 159. https://doi.org/10.1007/s10462-024-10810-4
  5. Sood, N. (2025, September). Deepfake detection and multimedia forensics: Investigating synthetic media, image forgery, and video manipulation in cybercrime cases. ARC Journal of Forensic Science, 9, 36–39.
  6. Khan, A. A., Lakshmi, A. A., Islam, S. A., & others. (2025). A survey on multimodal-enabled deepfake detection: State-of-the-art tools and techniques, emerging trends, current challenges & limitations, and future directions. Discover Computing, 28, 48.
  7. Gong, L. Y., & Li, X. J. (2024). A contemporary survey on deepfake detection: Datasets, algorithms, and challenges. Electronics, 13(7), 555.
  8. Cheng, H., Pang, W., Li, K., Wei, Y., Song, Y., & Chen, J. (2025). HTMD-Net: Enhanced feature interaction and multi-domain fusion deep forgery detection network. Journal of Imaging, 11(8), 312. https://doi.org/10.3390/jimaging11080312
  9. Zhao, L., Zhang, M., Ding, H., & Cui, X. (2021). MDT-Net: Deepfake detection network based on multi-feature fusion. Sensors, 21(3), 1002. https://doi.org/10.3390/s21031002
  10. Arora, M., & Bansal, M. (2025, April). Lightweight deepfake detection on mobile devices using semi-supervised MobileNet and frequency-domain analysis. Journal of Technology Informatics and Engineering, 4, 95–114.
  11. Kim, T., Choi, J., Jeong, Y., Noh, H., Yoo, J., Baek, S., & Choi, J. (2025, July). Beyond spatial features: Pixel-wise temporal frequency-based deepfake video detection. arXiv. https://doi.org/10.48550/arXiv.2507.07208
  12. Sun, F., Zhang, N., Xu, P., & Song, Z. (2021, November). Deepfake detection method based on cross-domain fusion. Security and Communication Networks, 2021, 1–11.
  13. Wang, F., Chen, Q., Jing, B., Tang, Y., Song, Z., & Wang, B. (2024, November). Deepfake detection based on the adaptive fusion of spatial-frequency features. International Journal of Intelligent Systems, 2024, Article 7578036. https://doi.org/10.1155/2024/7578036
  14. Kou, Y., Li, P., Ma, H., & others. (2025). STADE: Deepfake detection through spatial-frequency feature integration and dynamic margin optimization. Artificial Intelligence Review, 58, 217. https://doi.org/10.1007/s10462-025-11225-7
  15. Gao, J., Xia, Z., Marcello, G. L., Deng, C., Dai, J., & Feng, X. (2024). Deepfake detection based on high-frequency enhancement network for highly compressed content. Expert Systems with Applications, 249(Part C), 123732. https://doi.org/10.1016/j.eswa.2024.123732
  16. Wang, J., Wu, Z., Ouyang, W., Han, X., Chen, J., Jiang, Y.-G., & Li, S.-N. (2022). M2TR: Multi-modal multi-scale transformers for deepfake detection. In Proceedings of the 2022 International Conference on Multimedia Retrieval (ICMR ’22) (pp. 615–623). Association for Computing Machinery. https://doi.org/10.1145/3512527.3531415
  17. Gao, J., Huang, D., Zhang, J., Pilet, H., Hu, C., & Zhu, J. (2025). DDFormer: Deepfake detection with multimodal fusion transformer. In D. S. Huang, W. Chen, Y. Pan, & H. Chen (Eds.), Advanced Intelligent Computing Technology and Applications. ICIC 2025. Lecture Notes in Computer Science (Vol. 15863). Springer.
  18. Nelson, L., Batra, H., & Babu, P. (2025, June). Deepfake detection in manipulated images/audio/video: A three-stage multi-modal deep learning framework. 31, 30–39.
  19. Wang, Y., Sun, Q., Zhang, J., Rong, D., Shen, C., & Wang, X. (2025). Improving deepfake detection with predictive inter-modal alignment and future reconstruction in audio-visual synchrony scenarios. Information Fusion, 127(Part A), 103706. https://doi.org/10.1016/j.inffus.2025.103706
  20. Javed, M., Zhang, Z., Dutta, F. H., & others. (2025). Enhancing multimodal deepfake detection with local-global feature integration and diffusion module. Signal, Image and Video Processing, 19, 400. https://doi.org/10.1007/s11760-025-03970-7
  21. Miao, H., Gao, Y., Liu, Z., & Wang, Y. (2025). Multi-modal deepfake detection via multi-task audio-visual prompt learning. In Proceedings of the Thirty-Ninth AAAI Conference on Artificial Intelligence, AI Theory.
  22. Seventh Conference on Innovative Applications of Artificial Intelligence, and the Fifteenth Symposium on Educational Advances in Artificial Intelligence (AAAI-25/IAAI-25/EAAI-25) (Article 68, 10 pp.). AAAI Press. https://doi.org/10.1609/aaai.v39i1.32042
  23. Park, C., Moon, B., Joo, M., Jung, J., & Woo, S. S. (2025). XIA: Efficient multimodal deepfake detection with cross-level fusion. In Proceedings of the ACM Symposium on Applied Computing (SAC ’25) (pp. 767–774). Association for Computing Machinery.
  24. Mostafa, A., Rosanova, D., & Patel, P. (2024, January). Multimodal deepfake detection for short videos (pp. 67–73).
  25. Wang, L., Zhao, J., Zhang, X., & others. (2025). IRF-HA-TDD: A multimodal model for audio-visual deepfake detection. Videography, 2(10).