• e - ISSN No : 2832-4277
IJRTTE Logo

INTERNATIONAL JOURNAL OF RECENT TRENDS IN TECHNOLOGY AND ENGINEERING (IJRTTE)

Deep Learning-Based Automated Bug Detection Through Code Representation Analysis

Pradeep. H
Assistant Professor, Department of Mechanical Engineering, BGS Institute of Technology, India.
Mohan R
Professor, Department of Mechanical Engineering, Sona College of Technology, India.

Keywords: Deep Learning, Automated Bug Detection, Code Representation, Abstract Syntax Trees (ASTs), Explainable AI (XAI), Source Code Analysis.

Abstract

Modern software systems are becoming increasingly complex, hence requiring smarter bug finding tools, rather than convention rule-based systems. We propose a deep learning approach for the detection of bugs in source code by means of rich code representation analysis, which includes abstract syntax trees (AST), control-flow and token embeddings. Whereas prior approaches suffer from high false positive rates and lack language generalizability, we aim for a sound and language agnostic approach, robust to large codebases, in order to be easily integrated in continuous development pipelines. By automatically extracting the structural and semantic characteristic from code, it eliminates the need for manual inspection and guarantees the early discovery of both syntactic and semantic flaws. Additionally, the use of attention mechanisms and SHAP explain ability further provides interpretable interpretation of the decisioning process to developers. This work contributes to develop the state of the art of automated bug detection by overcoming the relevant limitations of previous approaches and providing a solution which is powerful, flexible, and explainable for real-world software engineering scenarios.
Download Certificate
Details

References

  1. X. Guan and C. Treude, “Enhancing source code representations for deep learning with static analysis,” arXiv preprint arXiv:2402.09557, 2024.
  2. L. Wartschinski, Y. Noller, T. Vogel, T. Kehrer, and L. Grunske, “VUDENC: Vulnerability Detection with Deep Learning on a Natural Codebase for Python,” arXiv preprint arXiv:2201.08441, 2022.
  3. Z. Chen, S. Kommrusch, and M. Monperrus, “Neural Transfer Learning for Repairing Security Vulnerabilities in C Code,” arXiv preprint arXiv:2104.08308, 2021.
  4. S. Yang et al., “Asteria-Pro: Enhancing Deep-Learning Based Binary Code Similarity Detection by Incorporating Domain Knowledge,” arXiv preprint arXiv:2301.00511, 2023.
  5. B. Steenhoek, H. Gao, and W. Le, “Dataflow Analysis-Inspired Deep Learning for Efficient Vulnerability Detection,” arXiv preprint arXiv:2212.08108, 2022.
  6. M. Böhme, V.-T. Pham, and A. Roychoudhury, “Coverage-Based Greybox Fuzzing as Markov Chain,” IEEE Trans. Softw. Eng., vol. 45, no. 5, pp. 521–536, May 2019.
  7. V.-T. Pham, M. Böhme, A. E. Santosa, A. R. Căciulescu, and A. Roychoudhury, “Smart Greybox Fuzzing,” IEEE Trans. Softw. Eng., vol. 47, no. 9, pp. 1874–1889, Sep. 2021.
  8. S. Chakraborty, R. Krishna, Y. Ding, and B. Ray, “Deep Learning Based Vulnerability Detection: Are We There Yet,” IEEE Trans. Softw. Eng., vol. 47, no. 3, pp. 700–717, Mar. 2021.
  9. L. Zhou, M. Huang, Y. Li, Y. Nie, and J. Li, “A Survey on Deep Learning-Based Vulnerability Detection,” in Proc. IEEE 6th Int. Conf. Data Sci. Cyberspace (DSC), 2021, pp. 1–8.
  10. T. Ganz, M. Härterich, A. Warnecke, and K. Rieck, “Code Property Graphs: A Survey of Techniques and Applications,” in Proc. 14th ACM Workshop Artif. Intell. Secur., 2021, pp. 1–12.
  11. X. Du, B. Chen, Y. Li, J. Guo, and Y. Zhou, “Devign: Effective Vulnerability Identification by Learning Comprehensive Program Semantics via Graph Neural Networks,” in Proc. NeurIPS, 2019.
  12. W. Zheng, Y. Jiang, and X. Su, “VulDeePecker: A Deep Learning-Based System for Vulnerability Detection,” in Proc. IEEE 32nd Int. Symp. Softw. Rel. Eng. (ISSRE), 2021, pp. 1–10.
  13. S. Khodayari and G. Pellegrino, “JAW: Studying Client-side CSRF with Hybrid Property Graphs and Declarative Traversals,” in Proc. 2021 IEEE 14th Int. Conf. Cloud Comput. (CLOUD), 2021, pp. 1–8.
  14. T. Brito, P. Lopes, N. Santos, and J. F. Santos, “Wasmati: An Efficient Static Vulnerability Scanner for WebAssembly,” Comput. Secur., vol. 110, p. 102420, Jul. 2022.
  15. S. Wi, S. Woo, J. J. Whang, and S. Son, “Deep Learning-Based Vulnerability Detection: A Survey,” in Proc. ACM Web Conf. 2022, 2022, pp. 1–10.
  16. B. Bowman and H. H. Huang, “Graph-Based Vulnerability Detection in Smart Contracts,” in Proc. 2020 IEEE Eur. Symp. Secur. Priv. (EuroS&P), 2020, pp. 1–10.
  17. X. Du, B. Chen, Y. Li, J. Guo, and Y. Zhou, “Backporting Security Patches of Web Applications: A Prototype Design and Implementation on Injection Vulnerability Patches,” in Proc. 2022 IEEE Int. Conf. Softw. Maint. Evol. (ICSME), 2022, pp. 1–10.
  18. A. Alhuzali, R. Gjomemo, B. Eshete, and V. N. Venkatakrishnan, “NAVEX: Precise and Scalable Exploit Generation for Dynamic Web Applications,” in Proc. 2018 ACM SIGSAC Conf. Comput. Commun. Secur., 2018, pp. 1–14.
  19. F. Al Kassar, G. Clerici, L. Compagna, D. Balzarotti, and F. Yamaguchi, “Testability Tarpits: The Impact of Code Patterns on the Security Testing of Web Applications,” in Proc. NDSS Symp., 2018, pp. 1–15.
  20. A. Fioraldi, D. Maier, H. Eißfeldt, and M. Heuse, “AFL++: Combining Incremental Steps of Fuzzing Research,” in Proc. 2020 IEEE Symp. Secur. Priv. Workshops (SPW), 2020, pp. 1–10.
  21. C. Lyu, S. Ji, C. Zhang, Y. Li, and W.-H. Lee, “MOPT: Optimized Mutation Scheduling for Fuzzers,” in Proc. 2019 IEEE Symp. Secur. Priv. (SP), 2019, pp. 1–12.
  22. M. Böhme, V.-T. Pham, M.-D. Nguyen, and A. Roychoudhury, “Coverage-Based Greybox Fuzzing as Markov Chain,” in Proc. ACM CCS, 2017, pp. 1–12.
  23. S. Poeplau and A. Francillon, “Symbolic Execution with SymCC: Don’t Interpret, Compile!,” in Proc. 2020 IEEE Symp. Secur. Priv. (SP), 2020, pp. 1–12.
  24. J. Sakhinini, H. Karimipour, and A. Dehghantanha, “A Deep and Scalable Unsupervised Machine Learning System for Cyber-Attack Detection in Large-Scale Smart Grids,” IEEE Access, vol. 8, pp. 1–10, 2020.
  25. A. Yazdinejad, H. Haddadpajouh, A. Dehghantanha, R. M. Parizi, and G. Srivastava, “Cryptocurrency Malware Hunting: A Deep Recurrent Neural Network Approach,” Appl. Soft Comput., vol. 96, p. 106630, Oct. 2020.
  26. O. Osanaiye, H. Cai, K.-K. R. Choo, A. Dehghantanha, and Z. Xu, “Ensemble-Based Multi-Filter Feature Selection Method for DDoS Detection in Cloud Computing,” EURASIP J. Wirel. Commun. Netw., vol. 2019, no. 1, p. 223, Dec. 2019.
  27. P. J. Taylor, T. Dargahi, A. Dehghantanha, R. M. Parizi, and K.-K. R. Choo, “A Systematic Literature Review of Blockchain Cyber Security,” Digit. Commun. Netw., vol. 6, no. 2, pp. 147–156, May 2020.
  28. N. Milosevic, A. Dehghantanha, and K.-K. R. Choo, “Machine Learning Aided Android Malware Classification,” Comput. Electr. Eng., vol. 61, pp. 1–10, Jul. 2017.