Original Article

Vol. 26 (2026): ELECTRICA (Continuous Publication)

TOLID: Turkish Offensive Language Identification Dataset and Transformer-Based Benchmarks

Main Article Content

Mehmet Salih Kurt
Eylem Yücel

Abstract

This study presents the Turkish Offensive Language Identification Dataset (TOLID), a large-scale, high-quality dataset for the automatic detection of offensive language in Turkish social media posts and evaluates the performance of several transformer-based models. Although most previous studies have been constrained by small sample sizes, imbalanced class distributions, or narrow topical focus, TOLID includes a wide range of offensive expressions without imposing restrictions on topics, individuals, or groups. The dataset was annotated by three independent experts using a hierarchical and fine-grained scheme that addresses not only general offensive language but also specific subcategories such as sexist, racist, political, and religious insults. To evaluate the dataset, several transformer-based models adapted for Turkish, including BERTurk (Bidirectional Encoder Representations from Transformers for Turkish), ConvBERTurk (Convolutional Bidirectional Encoder Representations from Transformers for Turkish), and ELECTRA-Turkish (Efficiently Learning an Encoder that Classifies Token Replacements Accurately for Turkish), were trained and tested. Among these, ConvBERTurk achieved the highest scores, reaching 82.76% macro F1 in offensive vs. non-offensive classification and 78.14% in targeted vs. non-targeted offensive classification, outperforming previous research. These results demonstrate that combining a balanced, multi-annotated dataset with advanced deep learning architectures can effectively address the linguistic richness, contextual complexity, and informal nature of Turkish social media text. Additionally, a web-based application was developed to provide a practical interface for analyzing text and visualizing model outputs, extending the study's impact beyond academic research to real-world applications. Overall, this study addresses key limitations of prior research and makes a significant contribution to Turkish natural language processing by providing a meticulously constructed dataset and extensive benchmarks with state-of-the-art models. TOLID establishes a robust foundation for future work on offensive language subtypes, automatic moderation, and toxicity analysis in Turkish social media.


Cite this article as: M. S. Kurt and E. Yücel, “TOLID: Turkish offensive language identification dataset and transformer-based benchmarks,” Electrica, 26, 0404, 2026. doi: 10.5152/electrica.2026.25404.


 

Article Details