Developing an Amh3MD as a multimodal dataset for the detection of information disorders
Published in Computational Sciences
Social media has accelerated the dissemination of information across diverse communities, creating benefits as well as risks. One significant risk is the rapid spread of misinformation, disinformation and mal‑information, which can have serious societal consequences, particularly in low‑resourced language communities where benchmark datasets for automatic detection are scarce. Existing publicly available multimodal benchmark datasets are largely dominated by high‑resource languages and binary classification paradigms, limiting their applicability to linguistically and culturally diverse contexts and to more nuanced distinctions among types of information disorder. To address this gap, this study developed Amh3MD, the first publicly available multimodal benchmark dataset targeted at information disorder detection in Amharic. Amh3MD comprises 8053 annotated multimodal samples collected from prominent national and regional media outlets and from public figures and activists with more than 10 000 followers on social media. Each sample contains textual content paired with a corresponding image or meme, reflecting the multimodal nature of contemporary online information. Following intent‑ and context‑based guidelines, items in the dataset are labelled into four classes - misinformation, disinformation, mal‑information, and normal enabling research into finer‑grained detection tasks beyond binary true/false distinctions. To ensure annotation quality, three domain experts participated in a structured annotation process, yielding substantial inter‑annotator agreement, with an overall agreement of 88.78% and a Cohen’s Kappa of 0.67. In addition to the dataset, the study proposed a multimodal deep learning architecture that combines transformer‑based textual representations with visual feature extraction models, employing feature‑level fusion strategies to integrate modalities. Together, Amh3MD and the proposed modelling approach provide a foundation for future implementation, evaluation, and extension of information disorder detection research in Amharic and other multilingual settings.