Machine Learning Based Classification of Qur’anic Verses into Makki and Madani

Tracking #: 952-1932


Submission Type: 

Research Paper

Abstract: 

The Holy Quran has about six thousand verses from various historical and cultural contexts, making classification difficult in computational linguistics and Quranic studies. This research aims to classify Quranic verses into two types: Meccan and Medinan. Meccan verses typically focus on matters of faith, while Medinan verses address social organisation and community governance. The research utilises a dataset of 6,236 Quranic verses. A hybrid text processing methodology was applied, beginning with data cleaning and the use of the Arabic linguistic root (ISRI Stemmer) to reduce semantic dimensions. This was followed by feature extraction using the TF-IDF statistical weighting system, combined with structural features such as verse length, to ensure the highest classification efficiency. The SelectKBest feature selection algorithm was used to select the top 500 linguistic features, and standard scaling was applied. An ensemble voting classifier was constructed, integrating three different logic algorithms: Random Forest (RF), Decision Trees (DT), and k-Nearest Neighbours (KNN). The system assessment employed the GroupKFold cross-validation test to ensure the reliability of results and prevent data leaking. The results showed that integrating linguistic and structural variables with ensemble models can automatically analyse and classify Qur’anic text with 82.33% accuracy

Manuscript: 

Tags: 

  • Reviewed

Data repository URLs: 

Date of Submission: 

Monday, July 13, 2026

Date of Decision: 

Thursday, July 30, 2026


Nanopublication URLs:

Decision: 

Reject (Pre-Screening)