2025

BanglaMemeEvidence: A Multimodal Benchmark Dataset for Explanatory Evidence Detection in Bengali Memes

Fatema Tuj Johora Faria, Mukaffi Bin Moin, Pronay Debnath, Md. Mahfuzur Rahman, Asif Iftekher Fahim, Faisal Muhammad Shah

Under review, 6th International Conference on Innovations in Computational Intelligence and Computer Vision (ICICV 2026) 2026

BanglaMemeEvidence: A Multimodal Benchmark Dataset for Explanatory Evidence Detection in Bengali Memes

Fatema Tuj Johora Faria, Mukaffi Bin Moin, Pronay Debnath, Md. Mahfuzur Rahman, Asif Iftekher Fahim, Faisal Muhammad Shah

Under review, 6th International Conference on Innovations in Computational Intelligence and Computer Vision (ICICV 2026) 2026

Enhancing Bangla NLP Tasks with LLMs: A Study on Few-Shot Learning, RAG, and Fine-Tuning Techniques

Saidur Rahman Sujon, Ahmadul Karim Chowdhury, Fatema Tuj Johora Faria, Mukaffi Bin Moin, Faisal Muhammad Shah

Under review, 28th International Conference on Computer and Information Technology (ICCIT 2025) 2025

Enhancing Bangla NLP Tasks with LLMs: A Study on Few-Shot Learning, RAG, and Fine-Tuning Techniques

Saidur Rahman Sujon, Ahmadul Karim Chowdhury, Fatema Tuj Johora Faria, Mukaffi Bin Moin, Faisal Muhammad Shah

Under review, 28th International Conference on Computer and Information Technology (ICCIT 2025) 2025

PotatoGANs: Utilizing Generative Adversarial Networks, Instance Segmentation, and Explainable AI for Enhanced Potato Disease Identification and Classification

Fatema Tuj Johora Faria*, Mukaffi Bin Moin*, Mohammad Shafiul Alam*, Ahmed Al Wase, Md. Rabius Sani, Khan Md Hasib (* equal contribution)

11th IEEE International Conference on Sustainable Technology and Engineering (IEEE i-COSTE 2025) 2025

Proposes PotatoGANs, using CycleGAN/Pix2Pix to synthesize realistic diseased-potato images for data augmentation, combined with three Explainable AI methods (Grad-CAM, Grad-CAM++, Score-CAM) across three CNN backbones for interpretable disease classification.

PotatoGANs: Utilizing Generative Adversarial Networks, Instance Segmentation, and Explainable AI for Enhanced Potato Disease Identification and Classification

Fatema Tuj Johora Faria*, Mukaffi Bin Moin*, Mohammad Shafiul Alam*, Ahmed Al Wase, Md. Rabius Sani, Khan Md Hasib (* equal contribution)

11th IEEE International Conference on Sustainable Technology and Engineering (IEEE i-COSTE 2025) 2025

Proposes PotatoGANs, using CycleGAN/Pix2Pix to synthesize realistic diseased-potato images for data augmentation, combined with three Explainable AI methods (Grad-CAM, Grad-CAM++, Score-CAM) across three CNN backbones for interpretable disease classification.

Exploring Explainable AI Techniques for Improved Interpretability in Lung and Colon Cancer Classification

Mukaffi Bin Moin, Fatema Tuj Johora Faria, Swarnajit Saha, Busra Kamal Rafa, Mohammad Shafiul Alam

4th International Conference on Computing and Communication Networks (ICCCNet-2024) 2024

Benchmarks ResNet50, VGG16, and DenseNet121 on histopathological lung/colon images with Grad-CAM, Grad-CAM++, and SHAP for interpretability; DenseNet121 reaches 92.78% accuracy with Grad-CAM++ giving the most precise localization.

Exploring Explainable AI Techniques for Improved Interpretability in Lung and Colon Cancer Classification

Mukaffi Bin Moin, Fatema Tuj Johora Faria, Swarnajit Saha, Busra Kamal Rafa, Mohammad Shafiul Alam

4th International Conference on Computing and Communication Networks (ICCCNet-2024) 2024

Benchmarks ResNet50, VGG16, and DenseNet121 on histopathological lung/colon images with Grad-CAM, Grad-CAM++, and SHAP for interpretability; DenseNet121 reaches 92.78% accuracy with Grad-CAM++ giving the most precise localization.

Unraveling the Dominance of Large Language Models Over Transformer Models for Bangla Natural Language Inference: A Comprehensive Study

Fatema Tuj Johora Faria, Mukaffi Bin Moin, Asif Iftekher Fahim, Pronay Debnath, Faisal Muhammad Shah

4th International Conference on Computing and Communication Networks (ICCCNet-2024) 2024

Compares LLMs (GPT-3.5 Turbo, Gemini 1.5 Pro) against Transformer baselines (BanglaBERT, mBERT) for Bangla natural language inference; GPT-3.5 Turbo reaches 92.15% few-shot accuracy, 6.48 points above BanglaBERT.

Unraveling the Dominance of Large Language Models Over Transformer Models for Bangla Natural Language Inference: A Comprehensive Study

Fatema Tuj Johora Faria, Mukaffi Bin Moin, Asif Iftekher Fahim, Pronay Debnath, Faisal Muhammad Shah

4th International Conference on Computing and Communication Networks (ICCCNet-2024) 2024

Compares LLMs (GPT-3.5 Turbo, Gemini 1.5 Pro) against Transformer baselines (BanglaBERT, mBERT) for Bangla natural language inference; GPT-3.5 Turbo reaches 92.15% few-shot accuracy, 6.48 points above BanglaBERT.

Towards Robust Chain-of-Thought Prompting with Self-Consistency for Remote Sensing VQA: An Empirical Study Across Large Multimodal Models

Fatema Tuj Johora Faria, Laith H. Baniata, Ahyoung Choi, Sangwoo Kang

Mathematics (MDPI), Vol. 13, Issue 18, Article 3046 2025

Evaluates GPT-4o, Grok 3, Gemini 2.5 Pro, and Claude 3.7 Sonnet on remote-sensing VQA using zero-shot, chain-of-thought, and self-consistency prompting (Self-GeoSense), reaching up to 94.69% accuracy on basic judging tasks with Grok 3.

Towards Robust Chain-of-Thought Prompting with Self-Consistency for Remote Sensing VQA: An Empirical Study Across Large Multimodal Models

Fatema Tuj Johora Faria, Laith H. Baniata, Ahyoung Choi, Sangwoo Kang

Mathematics (MDPI), Vol. 13, Issue 18, Article 3046 2025

Evaluates GPT-4o, Grok 3, Gemini 2.5 Pro, and Claude 3.7 Sonnet on remote-sensing VQA using zero-shot, chain-of-thought, and self-consistency prompting (Self-GeoSense), reaching up to 94.69% accuracy on basic judging tasks with Grok 3.

BanglaCalamityMMD: A Comprehensive Benchmark Dataset for Multimodal Disaster Identification in the Low-Resource Bangla Language

Fatema Tuj Johora Faria, Mukaffi Bin Moin, Busra Kamal Rafa, Swarnajit Saha, Md. Mahfuzur Rahman, Khan Md Hasib, M. F. Mridha

International Journal of Disaster Risk Reduction, Vol. 130, Article 105800 2025

Introduces BanglaCalamityMMD, a 7,903-sample multimodal (text+image) benchmark across seven disaster categories, and DisasterMultiFusionNet, which fuses Swin Transformer and mBERT to reach 85.25% accuracy, a 5.35% gain over the best unimodal baseline.

BanglaCalamityMMD: A Comprehensive Benchmark Dataset for Multimodal Disaster Identification in the Low-Resource Bangla Language

Fatema Tuj Johora Faria, Mukaffi Bin Moin, Busra Kamal Rafa, Swarnajit Saha, Md. Mahfuzur Rahman, Khan Md Hasib, M. F. Mridha

International Journal of Disaster Risk Reduction, Vol. 130, Article 105800 2025

Introduces BanglaCalamityMMD, a 7,903-sample multimodal (text+image) benchmark across seven disaster categories, and DisasterMultiFusionNet, which fuses Swin Transformer and mBERT to reach 85.25% accuracy, a 5.35% gain over the best unimodal baseline.

Analyzing Diagnostic Reasoning of Vision–Language Models via Zero-Shot Chain-of-Thought Prompting in Medical Visual Question Answering

Fatema Tuj Johora Faria, Laith H. Baniata, Ahyoung Choi, Sangwoo Kang

Mathematics (MDPI), Vol. 13, Issue 14, Article 2322 2025

Proposes a zero-shot chain-of-thought prompting framework that makes vision-language model reasoning explicit for medical VQA; on PMC-VQA, Gemini 2.5 Pro reaches 72.48% accuracy, ahead of Claude 3.5 Sonnet and GPT-4o Mini.

Analyzing Diagnostic Reasoning of Vision–Language Models via Zero-Shot Chain-of-Thought Prompting in Medical Visual Question Answering

Fatema Tuj Johora Faria, Laith H. Baniata, Ahyoung Choi, Sangwoo Kang

Mathematics (MDPI), Vol. 13, Issue 14, Article 2322 2025

Proposes a zero-shot chain-of-thought prompting framework that makes vision-language model reasoning explicit for medical VQA; on PMC-VQA, Gemini 2.5 Pro reaches 72.48% accuracy, ahead of Claude 3.5 Sonnet and GPT-4o Mini.

MultiBanFakeDetect: Integrating Advanced Fusion Techniques for Multimodal Detection of Bangla Fake News in Under-Resourced Contexts

Fatema Tuj Johora Faria, Mukaffi Bin Moin, Zayeed Hasan, Md. Arafat Alam Khandaker, Niful Islam, Khan Md Hasib, M. F. Mridha

International Journal of Information Management Data Insights, Vol. 5, Issue 2, Article 100347 2025

Introduces the MultiBanFakeDetect dataset and MultiFusionFake, an early-fusion text+image model (DenseNet-169 + mBERT) for Bangla fake news detection, reaching 79.69% accuracy versus 73.13% for a text-only baseline.

MultiBanFakeDetect: Integrating Advanced Fusion Techniques for Multimodal Detection of Bangla Fake News in Under-Resourced Contexts

Fatema Tuj Johora Faria, Mukaffi Bin Moin, Zayeed Hasan, Md. Arafat Alam Khandaker, Niful Islam, Khan Md Hasib, M. F. Mridha

International Journal of Information Management Data Insights, Vol. 5, Issue 2, Article 100347 2025

Introduces the MultiBanFakeDetect dataset and MultiFusionFake, an early-fusion text+image model (DenseNet-169 + mBERT) for Bangla fake news detection, reaching 79.69% accuracy versus 73.13% for a text-only baseline.

SentimentFormer: A Transformer-Based Multimodal Fusion Framework for Enhanced Sentiment Analysis of Memes in Under-Resourced Bangla Language

Fatema Tuj Johora Faria, Laith H. Baniata, Mohammad H. Baniata, Mohannad A. Khair, Ahmed Ibrahim Bani Ata, Chayut Bunterngchit, Sangwoo Kang

Electronics (MDPI), Vol. 14, Issue 4, Article 799 2025

Presents SentimentFormer, an intermediate-fusion transformer combining SwiftFormer and mBERT for Bangla meme sentiment analysis on the MemoSen dataset, reaching 79.04% accuracy versus 73.31% (text-only) and 64.72% (image-only).

SentimentFormer: A Transformer-Based Multimodal Fusion Framework for Enhanced Sentiment Analysis of Memes in Under-Resourced Bangla Language

Fatema Tuj Johora Faria, Laith H. Baniata, Mohammad H. Baniata, Mohannad A. Khair, Ahmed Ibrahim Bani Ata, Chayut Bunterngchit, Sangwoo Kang

Electronics (MDPI), Vol. 14, Issue 4, Article 799 2025

Presents SentimentFormer, an intermediate-fusion transformer combining SwiftFormer and mBERT for Bangla meme sentiment analysis on the MemoSen dataset, reaching 79.04% accuracy versus 73.31% (text-only) and 64.72% (image-only).

2024

Investigating the Predominance of Large Language Models in Low-Resource Bangla Language Over Transformer Models for Hate Speech Detection: A Comparative Analysis

Fatema Tuj Johora Faria, Laith H. Baniata, Sangwoo Kang

Mathematics (MDPI), Vol. 12, Issue 23, Article 3687 2024

Evaluates GPT-3.5 Turbo and Gemini 1.5 Pro against traditional methods for Bangla hate speech detection across multiple datasets; GPT-3.5 Turbo reaches up to 98.53% accuracy, a 6.28-point gain over prior approaches.

Investigating the Predominance of Large Language Models in Low-Resource Bangla Language Over Transformer Models for Hate Speech Detection: A Comparative Analysis

Fatema Tuj Johora Faria, Laith H. Baniata, Sangwoo Kang

Mathematics (MDPI), Vol. 12, Issue 23, Article 3687 2024

Evaluates GPT-3.5 Turbo and Gemini 1.5 Pro against traditional methods for Bangla hate speech detection across multiple datasets; GPT-3.5 Turbo reaches up to 98.53% accuracy, a 6.28-point gain over prior approaches.

Uddessho: An Extensive Benchmark Dataset for Multimodal Author Intent Classification in Low-Resource Bangla Language

Fatema Tuj Johora Faria, Mukaffi Bin Moin, Md. Mahfuzur Rahman, Md Morshed Alam Shanto, Asif Iftekher Fahim, Md. Moinul Hoque

18th International Conference on Information Technology and Applications (ICITA 2024) 2024

Introduces the Uddessho dataset (3,048 social-media posts) and a multimodal author-intent classification framework combining text and images; multimodal fusion reaches 76.19% accuracy, an 11.66-point gain over text-only.

Uddessho: An Extensive Benchmark Dataset for Multimodal Author Intent Classification in Low-Resource Bangla Language

Fatema Tuj Johora Faria, Mukaffi Bin Moin, Md. Mahfuzur Rahman, Md Morshed Alam Shanto, Asif Iftekher Fahim, Md. Moinul Hoque

18th International Conference on Information Technology and Applications (ICITA 2024) 2024

Introduces the Uddessho dataset (3,048 social-media posts) and a multimodal author-intent classification framework combining text and images; multimodal fusion reaches 76.19% accuracy, an 11.66-point gain over text-only.

Motamot: A Dataset for Revealing the Supremacy of Large Language Models Over Transformer Models in Bengali Political Sentiment Analysis

Fatema Tuj Johora Faria*, Mukaffi Bin Moin*, Rabeya Islam Mumu, Md Mahabubul Alam Abir, Abrar Nawar Alfy, Mohammad Shafiul Alam (* equal contribution)

The IEEE Region 10 Symposium (TENSYMP 2024) 2024

Introduces the Motamot dataset (7,058 instances) for Bangladeshi political sentiment analysis; with few-shot learning, Gemini 1.5 Pro reaches 96.33% accuracy, ahead of GPT-3.5 Turbo (94%) and BanglaBERT (88.10%).

Motamot: A Dataset for Revealing the Supremacy of Large Language Models Over Transformer Models in Bengali Political Sentiment Analysis

Fatema Tuj Johora Faria*, Mukaffi Bin Moin*, Rabeya Islam Mumu, Md Mahabubul Alam Abir, Abrar Nawar Alfy, Mohammad Shafiul Alam (* equal contribution)

The IEEE Region 10 Symposium (TENSYMP 2024) 2024

Introduces the Motamot dataset (7,058 instances) for Bangladeshi political sentiment analysis; with few-shot learning, Gemini 1.5 Pro reaches 96.33% accuracy, ahead of GPT-3.5 Turbo (94%) and BanglaBERT (88.10%).

Explainable Convolutional Neural Networks for Retinal Fundus Classification and Cutting-Edge Segmentation Models for Retinal Blood Vessels from Fundus Images

Fatema Tuj Johora Faria, Mukaffi Bin Moin, Pronay Debnath, Asif Iftekher Fahim, Faisal Muhammad Shah

Under review, Journal of Visual Communication and Image Representation 2024

Benchmarks eight pretrained CNNs with five Explainable AI methods for fundus image classification (ResNet101: 94.17% accuracy) and ten segmentation architectures for retinal vessel segmentation (Swin-Unet: 86.19% mean pixel accuracy).

Explainable Convolutional Neural Networks for Retinal Fundus Classification and Cutting-Edge Segmentation Models for Retinal Blood Vessels from Fundus Images

Fatema Tuj Johora Faria, Mukaffi Bin Moin, Pronay Debnath, Asif Iftekher Fahim, Faisal Muhammad Shah

Under review, Journal of Visual Communication and Image Representation 2024

Benchmarks eight pretrained CNNs with five Explainable AI methods for fundus image classification (ResNet101: 94.17% accuracy) and ten segmentation architectures for retinal vessel segmentation (Swin-Unet: 86.19% mean pixel accuracy).

2023

Vashantor: A Large-Scale Multilingual Benchmark Dataset for Automated Translation of Bangla Regional Dialects to Bangla Language

Fatema Tuj Johora Faria, Mukaffi Bin Moin, Ahmed Al Wase, Mehidi Ahmmed, Md Rabius Sani, Tashreef Muhammad

Under review, Neural Computing and Applications 2023

Introduces Vashantor, 12,000 sentence pairs across five Bangla regional dialects, and benchmarks mBART, NLLB, and GPT-3.5 Turbo for dialect-to-standard-Bangla translation; NLLB reaches a BLEU score of 32.45.

Vashantor: A Large-Scale Multilingual Benchmark Dataset for Automated Translation of Bangla Regional Dialects to Bangla Language

Fatema Tuj Johora Faria, Mukaffi Bin Moin, Ahmed Al Wase, Mehidi Ahmmed, Md Rabius Sani, Tashreef Muhammad

Under review, Neural Computing and Applications 2023

Introduces Vashantor, 12,000 sentence pairs across five Bangla regional dialects, and benchmarks mBART, NLLB, and GPT-3.5 Turbo for dialect-to-standard-Bangla translation; NLLB reaches a BLEU score of 32.45.

Classification of Potato Disease with Digital Image Processing Technique: A Hybrid Deep Learning Framework

Fatema Tuj Johora Faria, Mukaffi Bin Moin, Ahmed Al Wase, Md Rabius Sani, Khan Md Hasib, Mohammad Shafiul Alam

2023 IEEE 13th Annual Computing and Communication Workshop and Conference (CCWC) 2023

Proposes a hybrid CNN framework (ResNet50 + VGG16 with histogram equalization and edge detection) for potato leaf disease classification, reaching 89.34% accuracy on 5,000 annotated images, a 7.21-point gain over traditional ML methods.

Classification of Potato Disease with Digital Image Processing Technique: A Hybrid Deep Learning Framework

Fatema Tuj Johora Faria, Mukaffi Bin Moin, Ahmed Al Wase, Md Rabius Sani, Khan Md Hasib, Mohammad Shafiul Alam

2023 IEEE 13th Annual Computing and Communication Workshop and Conference (CCWC) 2023

Proposes a hybrid CNN framework (ResNet50 + VGG16 with histogram equalization and edge detection) for potato leaf disease classification, reaching 89.34% accuracy on 5,000 annotated images, a 7.21-point gain over traditional ML methods.