Research & Publications
My research interests span Multimodal Deep Learning, Vision Transformers, Transformer-Based Modality Encoders, and Automated Misinformation Detection.
-
01 A Unified Multimodal Framework for Fake News Detection Using BERT and Vision Transformers
Accepted & Presented 90.18% Test Acc 0.9485 ROC-AUCAyush Mishra and Naveen Kumar • International Conference on Intelligent Computing, Cognitive Networks, and Smart Systems (IC2NS2 2026) • Paper ID: 182 • SpringerDeveloped an end-to-end multimodal fake news detection framework that couples high-capacity modality-specific encoders with a streamlined, computationally efficient decision pipeline. The architecture utilizes a BERT–BiGRU network to capture deep contextual semantics alongside sequential narrative dependencies in news text, combined with a Vision Transformer (ViT-B/16) to model global visual context from images.
Mathematical Formulation & Architecture Pipeline- Text Stream (BERT–BiGRU): Contextual embeddings are refined through a Bidirectional GRU to capture temporal and narrative flow:
t = [hfinal ∥ hfinal] ∈ ℝdt
- Image Stream (ViT-B/16): Visual patch features are extracted through the Vision Transformer and projected into a latent subspace:
v = ReLU(Wv · vraw + bv) ∈ ℝdv
- Multimodal Concatenation Fusion & Classification: f = [t ∥ v], ŷ = softmax(Wc(ReLU(Wf · f + bf)) + bc)
- Empirical Performance (Fakeddit Benchmark): Evaluated on 30,900 test samples achieving 90.18% Accuracy, 0.9019 Precision, 0.9018 Recall, 0.9018 F1-Score (Macro F1: 0.9084), and 0.9485 ROC-AUC.
- Text Stream (BERT–BiGRU): Contextual embeddings are refined through a Bidirectional GRU to capture temporal and narrative flow:
Current Research Focus Areas (Forthcoming Work)
- Adaptive Cross-Modal Gated Alignment In Preparation: Designing non-linear gating mechanisms combining disentangled language representations (DeBERTa-v3) with hierarchical shifted-window attention (Swin-B / CLIP-ViT) to dynamically evaluate modality trustworthiness and filter incongruent noise.
- Hierarchical Vision Transformers & Artifact Detection: Exploring multi-scale patch attention mechanisms (Swin Transformer, MaxViT) to localize fine-grained manipulation boundaries and deepfake artifacts in social imagery.
- Explainable Multimodal Decisions: Developing gradient-weighted attribution maps and cross-attention rollout visualizations for verifiable misinformation auditing in production streams.