A Unified Multimodal Framework for Fake News Detection Using BERT and Vision Transformers
Abstract: The proliferation of digital misinformation across modern web platforms poses severe societal challenges. In this paper, we propose an integrated multimodal deep learning framework that harnesses the contextual representation power of Bidirectional Encoder Representations from Transformers (BERT) along with hierarchical Vision Transformers (ViT/Swin). By introducing a gated cross-modal fusion layer, our network adaptively weights linguistic credibility indicators against visual incongruities, effectively suppressing deceptive signal propagation. Extensive empirical evaluations validate superior classification performance and generalization across diverse fake news benchmarks.