Assessing Adversarial Vulnerabilities in Fake News Detection: A Comparative Study of GPT-2 and BERT Variants
Abstract
The increasing sophistication of fake news dissemination poses a growing threat to digital information integrity, demanding the deployment of robust and intelligent detection systems. Transformer-based language models particularly BERT, RoBERTa, DistilBERT, and GPT-2have shown promising results in detecting misinformation by leveraging deep contextual understanding. However, their vulnerability to adversarial attacks reveals a critical weakness in their deployment for real-world applications. This study conducts a comprehensive evaluation of these models under both standard and custom adversarial attack scenarios to assess their reliability in detecting manipulated or misleading content. Using the "newsmediabias/fake_news_elections_labelled_data" dataset, we fine-tune each model and subject them to a battery of adversarial techniques, including TextFooler, PWWS, BAE, DeepWordBug, TextBugger, as well as novel attack methods designed specifically for this study: Enhanced Substitution Attack (ESA) and Comprehensive Text Attack (CTA). We analyze model behavior in terms of accuracy degradation, perturbation efficiency, and computational cost. Our findings reveal stark contrasts in model robustness: while RoBERTa maintains the highest performance on clean data, it�along with other models�is significantly compromised under even subtle adversarial manipulations. The study highlights GPT-2's limitations as a generative model repurposed for classification, as it fails catastrophically under most attack conditions. These insights underscore the urgent need for adversarial resilience in fake news detection systems and pave the way for future research focused on integrating robust defense mechanisms into transformer-based architectures
Keywords: Adversarial Attacks, Fake News Detection, BERT, RoBERTa, DistilBERT, GPT-2, TextAttack, Model Robustness, NLP Security
References
- 1) Vaswani, A., Shazeer, N., Parmar, N., et al. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30.
- 2) Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. arXiv:1810.04805.
- 3) Liu, Y., Ott, M., Goyal, N., et al. (2019). RoBERTa: A robustly optimized BERT pretraining approach. arXiv:1907.11692.
- 4) Sanh, V., Debut, L., Chaumond, J., & Wolf, T. (2019). DistilBERT: A distilled version of BERT. arXiv:1910.01108.
- 5) Radford, A., Wu, J., Child, R., et al. (2019). Language models are unsupervised multitask learners (GPT-2). OpenAI.
- 6) Zhou, X., & Zafarani, R. (2018). Fake news: A survey of research, detection methods, and opportunities. arXiv:1812.00315.
- 7) Shu, K., Sliva, A., Wang, S., Tang, J., & Liu, H. (2017). Fake news detection on social media: A data mining perspective. ACM SIGKDD Explorations, 19(1), 22–36.
- 8) Ahmed, H., Traore, I., & Saad, S. (2017). Detecting fake news on Facebook. Stanford CS229.
- 9) Ruchansky, N., Seo, S., & Liu, Y. (2017). CSI: A hybrid deep model for fake news detection. CIKM, 797–806.
- 10) Wang, W. Y. (2017). Liar, liar pants on fire: A new benchmark dataset for fake news detection. ACL, 422–426.
- 11) Jin, D., Jin, Z., Zhou, J. T., & Szolovits, P. (2020). Is BERT really robust? A strong baseline for natural language attack on text classification and entailment. AAAI, 34(05), 8018–8025.
- 12) Morris, J., Lifland, E., Yoo, J., et al. (2020). TextAttack: A framework for adversarial attacks, data augmentation, and adversarial training in NLP. EMNLP (Demos), 119–126.
- 13) Ren, S., Diao, Q., Liang, Y., & Zhang, H. (2019). Generating natural language adversarial examples through probability weighted word saliency. ACL, 1085–1097.
- 14) Gao, J., et al. (2018). Black-box generation of adversarial text sequences to evade deep learning classifiers. IEEE S&P Workshops, 50–56.
- 15) Li, J., Ji, S., Du, T., Li, B., & Wang, T. (2019). TextBugger: Generating adversarial text against real-world applications. NDSS Symposium.
- 16) Jia, R., & Liang, P. (2017). Adversarial examples for evaluating reading comprehension systems. ACL, 2021–2031.
- 17) Wang, B., et al. (2021). Adversarial training for large neural language models. arXiv:2010.12563.
- 18) Michel, P., Levy, O., & Neubig, G. (2019). Are sixteen heads really better than one? NeurIPS, 14014–14024.
- 19) Ribeiro, M. T., Singh, S., & Guestrin, C. (2018). Semantically equivalent adversarial rules for debugging NLP models. ACL, 856–865.
- 20) Zhang, C., et al. (2020). Learning to detect and refute misinformation on social media. EMNLP, 549–560.
- 21) Raza, S., Rahman, M., & Ghuge, S. (2024). Analyzing the impact of fake news on the 2024 election. arXiv:2312.03750.
- 22) Wang, A., et al. (2018). GLUE: A multi-task benchmark and analysis platform for natural language understanding. ICLR.
- 23) Horne, B. D., & Adalı, S. (2017). This just in: Fake news packs a lot in title, uses simpler, repetitive content in text body. ICWSM.
- 24) Zhang, Z., et al. (2018). FakeNewsNet: A data repository with news content, social context and dynamic information. arXiv:1809.01286.
- 25) TextAttack documentation. https://textattack.readthedocs.io
- 26) Brown, T., et al. (2020). Language models are few-shot learners (GPT-3). NeurIPS.
- 27) Wallace, E., Feng, S., Kandpal, N., et al. (2019). Universal adversarial triggers for attacking and analyzing NLP. EMNLP, 2153–2162.
- 28) Schick, T., & Schütze, H. (2021). Exploit cloze questions for few-shot text classification and NLI. EACL.
- 29) Niven, T., & Kao, H. Y. (2019). Probing neural network comprehension of natural language arguments. ACL, 4658–4664.
- 30) Bhargava, P., et al. (2021). Generalizing adversarial attacks to generative models. arXiv:2107.06817.
- 31) Garg, S., & Ramakrishnan, G. (2020). BAE: BERT-based adversarial examples for text classification. EMNLP, 6174–6181.
- 32) Zhang, Z., et al. (2019). Generating fluent adversarial examples for natural languages. ACL, 5564–5574.
- 33) Ebrahimi, J., Rao, A., Lowd, D., & Dou, D. (2018). HotFlip: White-box adversarial examples for text classification. ACL, 31–36.
- 34) Wallace, E., et al. (2020). Imitating text style with adversarially trained generators. EMNLP, 1170–1185.
- 35) Koenders, C., Filla, J., Schneider, N., & Woloszyn, V. (2021). How vulnerable are fake news detection methods to adversarial attacks? arXiv:2107.07970.
- 36) Hu, E., Shen, Y., Wallis, P., et al. (2021). LoRA: Low-rank adaptation of large language models. arXiv:2106.09685.
- 37) Ding, M., et al. (2022). Parameter-efficient transfer learning with LoRA for text classification. NeurIPS.
- 38) Xu, H., et al. (2022). Fine-tuning pretrained transformers efficiently with LoRA. arXiv:2207.05914.
- 39) Dettmers, T., et al. (2023). QLoRA: Efficient fine-tuning of quantized LLMs. arXiv:2305.14314.
- 40) Liu, P., et al. (2021). Pre-train, prompt, and predict: A systematic survey of prompting methods in NLP. arXiv:2107.13586.
- 41) Brundage, M., et al. (2018). The malicious use of artificial intelligence: Forecasting, prevention, and mitigation. arXiv:1802.07228.
- 42) Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., & Choi, Y. (2019). Defending against neural fake news. NeurIPS.
- 43) Gilmary, C., et al. (2022). Responsible AI for fake news detection: Bias, fairness, and accountability. ACM Computing Surveys.
- 44) Buchanan, B., & Miller, T. (2020). Machine learning for policymakers: Fake news and social threats. Brookings Institution.
- 45) Al-Rubaie, M., & Chang, J. M. (2019). Privacy-preserving machine learning: Threats and solutions. IEEE Security & Privacy.
- 46) Zhang, J., et al. (2021). A survey on adversarial attacks and defenses in text. ACM Computing Surveys.
- 47) Yuan, X., He, P., Zhu, Q., & Li, X. (2019). Adversarial examples: Attacks and defenses for deep learning. IEEE Transactions on Neural Networks and Learning Systems.
- 48) Wang, Y., et al. (2020). Survey on NLP attacks and defenses. arXiv:2010.13303.
- 49) Qiu, X., et al. (2020). Pre-trained models for natural language processing: A survey. Science China.
- 50) Mozes, M., et al. (2021). Frequency-guided word substitutions for detecting adversarial attacks. ACL.
- 51) Glavaš, G., & Štajner, S. (2021). Simplicity bias in transformers. Findings of EMNLP, 3766–3775.
- 52) Kumar, S., et al. (2021). Adversarial robustness of pre-trained language models: A survey. arXiv:2108.07258.
Explore Our Related Journals
Looking for the right journal for your next manuscript? Explore our international peer-reviewed journals covering engineering, management, computer science, artificial intelligence and multidisciplinary research.