A complete list of my publications and preprints. For a shorter, curated set, see Selected Publications on the homepage.
* equal contribution (co-first author); †corresponding author.
đź“„ Preprints
- Skills That Don’t Exist: A Large-Scale Study of Hallucinated Skill Recommendation in LLM Agents
- Weifeng Yuan, Wenbo Guo†, Feng Dong, Haoyu Wang, Yang Liu
- arXiv preprint, 2026
- This paper presents a large-scale study of hallucinated skill recommendation in LLM agents, characterizing when and why agents recommend skills that do not exist.
- MalSkillBench: A Runtime-Verified Benchmark of Malicious Agent Skills
- Wenbo Guo, Wei Zeng, Chengwei Liu, Xiaojun Jia, Yijia Xu, Lei Tang, Yong Fang, Yang Liu
- arXiv preprint, 2026
- This paper presents MalSkillBench, the first runtime-verified benchmark of malicious agent skills—3,944 malicious skills labeled along a three-dimensional taxonomy—showing that detecting malicious skills requires reasoning jointly over task intent, code, and instructions.
📝 Publications
- How Effective Are NPM Malicious Package Detectors? A Large-Scale Empirical Study
- Wenbo Guo, Zhongwen Chen, Zhengzi Xu, Chengwei Liu, Ming Kang, Shiwen Song, Chengyue Liu, Yijia Xu, Weisong Sun, Yang Liu
- 41st IEEE/ACM International Conference on Automated Software Engineering (ASE), 2026
- This paper presents the first large-scale empirical study of NPM malicious package detection, evaluating 8 tools with 13 variants on a unified benchmark of 6,420 malicious and 7,288 benign packages, and inspecting each tool’s source code to explain why it succeeds or fails.
- Latent Reuse in Agent Skills: Multi-Modal Clone Detection at Ecosystem Scale
- Jiaying Zhu, Lyuye Zhang, Wenbo Guo, Yang Liu
- 41st IEEE/ACM International Conference on Automated Software Engineering (ASE), 2026
- This paper presents SkillClone, the first multi-modal clone detection approach for agent skills, which recovers hidden reuse links across YAML metadata, natural language instructions, and embedded code at ecosystem scale.
- An Empirical Study of Observability Limits in Advanced Software Supply Chain Attacks
- Zhuoran Tan, Wenbo Guo, Jiewen Luo, Taylor Brierley, Jeremy Singer, Christos Anagnostopoulos
- ACM Conference on Computer and Communications Security (CCS), 2026
- This paper presents SynthChain, a multi-source runtime dataset with chain-level ground truth, and empirically studies the observability limits of software supply chain attacks.
- Characterizing and Repairing Obsolete Android GUI Tests under UI Evolution
- Shiwen Song, Yiheng Xiong, Wenbo Guo, Manqi Sun, Jiaolong Kong, Xiaofei Xie
- ACM SIGSOFT International Symposium on Software Testing and Analysis (ISSTA), 2026
- This paper characterizes obsolete Android GUI tests caused by UI evolution and proposes an automated approach to repair them.
- Cutting the Gordian Knot: Detecting Malicious PyPI Packages via a Knowledge-Mining Framework
- Wenbo Guo, Chengwei Liu, Ming Kang, Yiran Zhang, Jiahui Wu, Zhengzi Xu, Vinay Sachidananda, Yang Liu
- USENIX Security Symposium (USENIX Security), 2026
- This paper proposes a knowledge-mining framework that transforms false positives from existing tools into knowledge to drive malicious package detection, discovering over 500 new malicious packages in the PyPI ecosystem.
- Bridging Expert Reasoning and LLM Detection: A Knowledge-Driven Framework for Malicious Packages
- Wenbo Guo, Shiwen Song, Jiaxun Guo, Zhengzi Xu, Chengwei Liu, Haoran Ou, Mengmeng Ge, Yang Liu
- The Web Conference (WWW), 2026
- This paper proposes a knowledge-driven framework that extracts and summarizes knowledge from historical expert malware analysis reports to drive LLM-based malicious package detection.
- Casting a SPELL: Sentence Pairing Exploration for LLM Limitation-breaking
- Yifan Huang, Xiaojun Jia, Wenbo Guo, Yuqiang Sun, Yihao Huang, Chong Wang, Yang Liu
- ACM International Conference on the Foundations of Software Engineering (FSE), 2026
- This paper proposes SPELL, a testing framework that evaluates security alignment weaknesses in LLM code generation through time-division sentence pairing strategies, achieving high attack success rates across major code models.
- IntelliRadar: A Comprehensive Platform to Pinpoint Malicious Packages from Cyber Intelligence
- Wenbo Guo, Chengwei Liu, Limin Wang, Yiran Zhang, Jiahui Wu, Zhengzi Xu, Yang Liu
- 48th International Conference on Software Engineering (ICSE), 2026
- This paper presents IntelliRadar, a comprehensive platform designed to identify and pinpoint malicious packages using cyber intelligence.
- BinStruct: Binary Structure Recovery Combining Static Analysis and Semantics
- Yiran Zhang, Zhengzi Xu, Zhe Lang, Chengyue Liu, Yuqiang Sun, Wenbo Guo, Chengwei Liu, Weisong Sun, Yang Liu
- 40th IEEE/ACM International Conference on Automated Software Engineering (ASE), 2025
- This paper proposes BinStruct, a novel framework that combines static analysis and semantics for binary structure recovery.
- Evolaris: A Roadmap to Self-Evolving Software Intelligence Management
- Chengwei Liu, Wenbo Guo, Yuxin Zhang, Limin Wang, Sen Chen, Lei Bu, Yang Liu
- 29th International Conference on Engineering of Complex Computer Systems (ICECCS), 2025 (Position Paper)
- This paper presents Evolaris, a roadmap to self-evolving software intelligence management.
- SpearBot: Leveraging Large Language Models in a Generative-Critique Framework for Spear-Phishing Email Generation
- Qinglin Qi, Yun Luo, Yijia Xu, Wenbo Guo, Yong Fang
- Information Fusion, Volume 122, 2025
- This paper proposes SpearBot, a novel framework that leverages large language models in a generative-critique approach for automated spear-phishing email generation.
- Few-shot graph classification on cross-site scripting attacks detection
- Hongyu Pan, Yong Fang, Wenbo Guo, Yijia Xu, Changhui Wang
- Computers & Security, 2024
- This paper is about detecting cross-site scripting attacks using the few-shot learning.
- An Empirical Study of Malicious Code In PyPI Ecosystem
- Wenbo Guo, Zhengzi Xu, Chengwei Liu, Cheng Huang, Yong Fang, Yang Liu
- 2023 38th IEEE/ACM International Conference on Automated Software Engineering (ASE)
- This paper is the empirical study paper for the malicious packages in the PyPI Package Manager.
- We constructed the largest PyPI malicious package dataset and open-sourced the dataset at here
- HyVulDect: a hybrid semantic vulnerability mining system based on graph neural network
- Wenbo Guo, Yong Fang, Cheng Huang, Haoran Ou, Chun Lin, Yongyan Guo
- Computers & Security, 2022
- This paper is about vulnerability detection based on the graph neural network.
- Viopolicy-detector: An automated approach to detecting GDPR suspected compliance violations in websites
- Haoran Ou, Yong Fang, Yongyan Guo, Wenbo Guo, Cheng Huang
- Proceedings of the 25th International Symposium on Research in Attacks, Intrusions and Defenses (RAID 2022)
- This paper is about detecting privacy compliance using machine learning method.
- Intelligent mining vulnerabilities in python code snippets
- Wenbo Guo, Cheng Huang, Weina Niu, Yong Fang
- Journal of Intelligent & Fuzzy Systems, 2021
- This paper is about detecting python vulnerability using the machine learning method.
- HackerRank: Identifying key hackers in underground forums
- Cheng Huang, Yongyan Guo, Wenbo Guo, Ying Li
- International Journal of Distributed Sensor Networks, 2021
- This paper is about detecting key hackers in the underground forums.
- No Pie in the Sky: The Digital Currency Fraud Website Detection
- Haoran Ou, Yongyan Guo, Chaoyi Huang, Zhiying Zhao, Wenbo Guo, Yong Fang, Cheng Huang
- 2021 IEEE 12th EAI International Conference on Digital Forensics & Cyber Crime (ICDF2C)
- This paper is about detecting Ponzi cryptocurrency using machine learning methods.