Abstract:
The Holy Quran has about six thousand verses from various historical and cultural contexts, making classification
difficult in computational linguistics and Quranic studies. This research aims to classify Quranic verses into two
types: Meccan and Medinan. Meccan verses typically focus on matters of faith, while Medinan verses address social
organisation and community governance. The research utilises a dataset of 6,236 Quranic verses. A hybrid text
processing methodology was applied, beginning with data cleaning and the use of the Arabic linguistic root (ISRI
Stemmer) to reduce semantic dimensions. This was followed by feature extraction using the TF-IDF statistical weighting
system, combined with structural features such as verse length, to ensure the highest classification efficiency. The
SelectKBest feature selection algorithm was used to select the top 500 linguistic features, and standard scaling was
applied. An ensemble voting classifier was constructed, integrating three different logic algorithms: Random Forest
(RF), Decision Trees (DT), and k-Nearest Neighbours (KNN). The system assessment employed the GroupKFold cross-validation
test to ensure the reliability of results and prevent data leaking. The results showed that integrating linguistic and
structural variables with ensemble models can automatically analyse and classify Qur’anic text with 82.33% accuracy