DOI : 10.5281/zenodo.23078465
- Open Access

- Authors : Abhijith B S, Adwaith Raj R, Afsal A, Aswin Saj
- Paper ID : IJERTV15IS090622
- Volume & Issue : Volume 15, Issue 09 , September – 2026
- Published (First Online): 01-10-2026
- ISSN (Online) : 2278-0181
- Publisher Name : IJERT
- License:
This work is licensed under a Creative Commons Attribution 4.0 International License
Cyber Shield: A Scalable Real-Time Framework for Automated Cyberbullying Detection using Classical Machine Learning and Asynchronous Web Architecture
Abhijith B S, Adwaith Raj R, Afsal A, Aswin Saj
Department of Computer Science and Engineering
Vidya Academy of Science and Technology Technical Campus, Thiruvananthapuram, Kerala, India
Abstract – The rapid expansion of social media has escalated digital harassment, creating an imperative for low-latency, scalable moderation architectures. This paper introduces Cyber Shield, an end-to-end detection pipeline coupling classical statistical classifiers with an asynchronous web execution tier for live remediation of toxic online discourse. Three supervised classification paradigms Stochastic Gradient Descent Linear Support Vector Machines (SGD-SVM), Logistic Regression (LR), and Naive Bayes (NB)are evaluated across two textual representation strategies: Term Frequency-Inverse Document Frequency (TF-IDF) and Latent Semantic Analysis (LSA) implemented via Truncated Singular Value Decomposition. Benchmarked across a consolidated corpus of 367,245 annotated social entries, the sublinear TF-IDF + Linear SVM configuration yields optimal discrimination, attaining a 92.9% accuracy, 96.0% precision, 92.8% recall, and an F1 metric of 94.4%, requiring only 0.002 seconds per batch evaluation. The serving layer, engineered on an Asynchronous Server Gateway Interface (ASGI) utilizing FastAPI and Uvicorn, sustains 15,00020,000 transactions per second, delivering an order-of-magnitude throughput gain over synchronous WSGI designs. Model interpretability is integrated via Local Interpretable Model-Agnostic Explanations (LIME) and Shapley Additive exPlanations (SHAP), providing granular, token-level transparency for moderation decisions. The empirical evidence confirms that lightweight, optimized linear classifiers over high- dimensional sparse representations offer the most pragmatic convergence of detection efficacy, sub-millisecond execution, compute efficiency, and auditability for web-scale moderation.
Keywords cyberbullying detection; natural language processing; support vector machines; TF-IDF; FastAPI; explainable artificial intelligence; asynchronous inference; content moderation
-
INTRODUCTION
Ubiquitous digital connectivity has transformed collaborative communication while simultaneously lowering barriers to hostility and systematic harassment. Cyberbullyingdefined by persistent, deliberate hostility mediated through networked platforms toward individuals with limited capacity for self-defensetargets protected attributes including ethnicity, gender, sexual identity, physical appearance, and disability [1]. Unlike physical altercations, digital abuse transcends spatial and temporal constraints, providing bad actors with an asymmetrical reach amplified by screen anonymity. Large-scale cross-sectional studies reveal that digital abuse affects upwards of 60% of youth populations and roughly 40% of adult internet users [1].
Contemporary containment frameworks remain vulnerable to operational bottlenecks. Human-in-the-loop review queues cannot digest the millions of incoming streams produced per minute, and prolonged exposure to toxic content inflicts documented psychological harm on human annotators. Conversely, rule-based lexicon systems rely on static string matching, causing elevated false alarms on benign linguistic nuances while failing against basic obfuscations like leetspeak, deliberate spelling variations, and phonetic transliterations [2].
Consequently, research has pivoted toward statistical natural language processing and machine learning. While attention-based deep architectures (e.g., BERT, RoBERTa) demonstrate high benchmark scores, their multi-layered
parameter footprints entail prohibitive compute costs, high latency, and excessive memory overhead during peak serving loads [3]. Production deployments require an operational compromise spanning classification robustness, operational execution speed, and decision transparency.
To address these operational constraints, this study presents Cyber Shield, a high-throughput moderation engine coupling tuned classical classifiers with an asynchronous execution framework. The key contributions of this research are as follows:
-
A systematic comparative analysis of six algorithmic configurations evaluated across an aggregated benchmark corpus containing 367,245 validated social text entries.
-
Quantitative validation that an SGD-optimized Linear SVM operating over sublinear TF-IDF tokens achieves top operational utility (92.9% accuracy, 94.4% F1 score) with an average batch latency of 0.002 seconds.
-
Implementation of a non-blocking asynchronous serving pipeline via FastAPI and Uvicorn, consistently processing 15,000 to 20,000 transactions per second.
-
Integration of dual interpretability layers (SHAP and LIME) to facilitate compliance auditing and token-level algorithmic explainability.
-
-
RELATED WORK
-
Classical Machine Learning for Cyberbullying Detection
Statistical machine learning has long anchored automated lexical processing. Isa et al. [1] established standard baseline metrics leveraging Naive Bayes, Logistic Regression, and linear boundary estimators across n-gram matrices. Naive Bayes operates on strong attribute-independence assumptions, remaining computationally cheap but underperforming in the presence of strong phrase interdependencies. Support Vector Machines define optimal boundary hyperplanes across wide feature spaces, showing inherent resistance to input noise and dimensional sparseness [4]. Logistic Regression yields direct posterior probabilities, making it an established baseline; yet standard bag-of-words approximations discard sequential word structure, complicating the detection of nuanced contextual hostility.
-
Deep Learning and Attention Architectures
To model sequential dependencies, hybrid deep networks have been widely explored. Luo et al. [5] and Daraghmi [6] demonstrated that stacking Convolutional Neural Networks (CNNs) with Bidirectional Long Short-Term Memory (BiLSTM) units captures localized phrase patterns together with temporal semantic cues, achieving F1 values spanning
0.78 to 0.98. Multi-head transformer encoders (e.g., BERT, RoBERTa) leverage cross-token self-attention to resolve polysemy and discern veiled harassment [7]. However, as demonstrated by Verma et al. [7] and Sanh et al. [8], the massive parameter budgets and matrix multiplication requirements of deep transformers lead to latencies up to two orders of magnitude higher than linear models, restricting real- time deployment on standard server hardware.
-
Scalability and Asynchronous Deployment
Deploying inference engines within high-traffic web environments requires robust handling of volatile query volumes. Almomani et al. [9] documented that classical synchronous WSGI wrappers (such as standard Flask configurations) suffer rapid thread depletion and cascading latency degradation during traffic surges. Asynchronous Server Gateway Interface (ASGI) frameworks resolve this bottleneck by utilizing event loops and non-blocking I/O routines, maintaining high connection concurrency with minimal compute overhead.
-
Feature Extraction Paradigms
Model discrimination is heavily governed b the initial token projection strategy. Sublinear TF-IDF balances term occurrences against cross-document distribution, dampening excessively frequent words while prioritizing distinct discriminative tokens [10]. In contrast, Latent Semantic Analysis (LSA) applies Truncated Singular Value Decomposition (SVD) to project broad term frequencies into compact continuous semantic vectors. Although LSA handles synonymy well, the runtime projection matrix multiplication introduces non-negligible processing latency during live serving.
-
-
SYSTEM DESIGN AND METHODOLOGY
The Cyber Shield architecture is organized as a decoupled, modular pipeline coordinating data ingestion, text sanitization, tokenization, model inference, and real-time explanation generation.
-
Dataset Aggregation and Preprocessing
The validation set was constructed by aggregating two distinct open-access social harassment repositories. Dataset 1 includes 115,802 instances (12.8% positive class), and Dataset
2 contains 251,443 instances (14.1% positive class). The merged corpus consists of 367,245 annotated records exhibiting a 13.7% overall positive prevalence and a mean string length of 353.9 characters. To prevent majority-class decision bias, stratified splitting was applied to maintain balanced class ratios across the 80/20 train-test division.
The raw string records pass through a deterministic preprocessing sequence: (1) URL removal via regular expression filters (http\S+|www\.\S+); (2) user handle (@mention) and digit stripping; (3) removal of extraneous punctuation and special symbols while preserving single space delimiters; (4) case normalization to lowercase; and (5) token normalization through whitespace collapsing.
-
Feature Vectorization Architectures
Two feature extraction pipelines are implemented: (1) Sublinear TF-IDF: Configured for unigram and bigram extraction, capped at a vocabulary limit of 15,000 features, applying sublinear scaling (1 + log(TF)) to dampen dominant tokens; (2) LSA Vectorization: Chaining a 20,000-term CountVectorizer with Truncated SVD to project the lexical space into 200 latent semantic dimensions, followed by Euclidean L2 normalization to enforce unit length.
-
Classification Algorithms
Three supervised algorithm families were systematically trained across each feature space: (1) Linear SVM / SGD: A Linear Support Vector Classifier trained via Stochastic Gradient Descent using an elastic net penalty, encapsulated in a 3-fold cross-validated Platt scaling wrapper (CalibratedClassifierCV) to yield reliable class posterior probabilities; (2) Logistic Regression: Regularized with L2 penalty using the liblinear coordinate descent solver; (3) Naive Bayes: Evaluated via Multinomial Naive Bayes over sparse TF-IDF matrices, and Gaussian Naive Bayes over continuous LSA semantic projections.
-
Asynchronous Web Architecture
The backend inference engine is constructed with FastAPI and deployed via the Uvicorn ASGI server. To eliminate redundant runtime initialization, pre-trained model weights and vectorizers are loaded into an in-memory dictionary cache during the application startup event. Incoming HTTP POST requests containing raw text payloads are parsed asynchronously, transformed into feature vectors via the cached vectorizer, and evaluated by the classifier within an asynchronous task loop.
The client interface is implemented as a Single Page Application (SPA) built with React 19 and Vite 6, presenting real-time toxicity scores, inference execution metrics, and interpretability graphs.
-
Explainable AI Integration
To address algorithmic opacity in automated moderation, two complementary interpretability mechanisms were integrated: (1) LIME (Local Interpretable Model-Agnostic Explanations): Computes local linear surrogate approximations around target instances, identifying specific abusive words to provide intuitive justifications for moderation staff [11]; (2) SHAP (Shapley Additive exPlanations): Employs cooperative game theory principles to calculate global Shapley values across the corpus, verifying that feature weights align with authentic hostility indicators rather than spurious dataset-specific correlations [12].
-
-
EXPERIMENTAL RESULTS AND DISCUSSION
-
Classification Performance
Table I reports the performance across classical supervised configurations on the 20% validation split. Table II contextualizes these empirical outcomes against contemporary deep learning benchmarks reported in recent cyberbullying detection literature.
TABLE I. PERFORMANCE OF CLASSICAL ML CONFIGURATIONS
Algorithm
Feature
Acc.
Prec.
Recall
F1
Time(s)
SGD / Linear SVM
TF-IDF
92.9%
96.0%
92.8%
94.4%
0.002
Bagging Classifier
TF-IDF
92.7%
96.7%
92.3%
94.5%
0.210
Decision Tree
TF-IDF
92.5%
95.5%
92.9%
94.2%
0.024
Random Forest
TF-IDF
91.7%
94.8%
92.3%
93.6%
2.890
K-Nearest Neighbours
TF-IDF
85.8%
89.5%
88.7%
89.1%
20.495
Naive Bayes (Gauss.)
LSA
84.0%
79.0%
62.0%
69.5%
<0.010
TABLE II. BENCHMARK COMPARISON WITH DEEP LEARNING
Model Architecture
Representation
Accuracy
F1 Score
Ref.
BiLSTM
Word2Vec
Up to 98.0%
0.78 – 0.97
[5] CNN-BiLSTM- GRU
Dense Embeddings
Up to 99.0%
0.95 – 0.98
[6] BERT (Base)
Attention
92.0 – 94.8%
0.92 – 0.96
[7] RoBERTa
Ensemble
Attention
~96.0%
~0.96
[9] -
Analysis of Classifier Outcomes
The calibrated Linear SVM (SGD) operating on TF-IDF features recorded the highest operational efficacy, achieving 92.9% accuracy and an F1 score of 94.4% with an inference latency of 0.002 seconds per batch. The 15,000-dimensional sparse feature space enables effective linear separation with minimal update overhead.
While the Bagging Classifier achieved a slightly higher precision of 96.7%, its prediction latency of 0.210 seconds represents a hundred-fold delay compared to the linear model. In content moderation ecosystems where minimizing false accusations is critical, Bagging provides strong reliability, but its computational cost restricts high-throughput operational scaling.
Conversely, the K-Nearest Neighbours classifier degraded markedly, achieving only 85.8% accuracy while requiring
20.495 seconds per inference batch. In 15,000-dimensional sparse feature spaces, distance concentration severely weakens Euclidean discriminative power, rendering neighborhood lookups computationally prohibitive. Gaussian Naive Bayes over LSA reached an F1 score of only 69.5%, demonstrating that conditional independence assumptions falter when toxicity stems from multi-word contextual dependencies.
-
Asynchronous API Performance
Stress testing conducted under simulated concurrent loads (1,000 to 50,000 simultaneous requests) demonstrated that the FastAPI + Uvicorn deployment sustained an operational throughput between 15,000 and 20,000 requests per second. Under identical hardware conditions, an equivalent synchronous Flask WSGI server collapsed at 2,150 requests per second due to thread pool starvation. In isolated benchmark suites, the asynchronous framework completed execution in 1.56 seconds versus 27.8 seconds for the synchronous baselinea 17.8x acceleration.
Memory provisioning profiles indicated an operational footprint of approximately 1.1 GB per worker process. Sustaining four concurrent workers required ~4.5 GB of system memory, confirming that the architecture operates comfortably on entry-level commodity cloud instances.
-
Explainability Findings
Instance-level evaluations utilizing LIME showed that the Linear SVM model accurately identified abusive unigrams and bigrams with high fidelity to human judgment. Global feature importance computed via SHAP confirmed that identity- targeted slurs and hostile imperatives received dominant positive attribution weights, while neutral grammatical tokens hovered near zero, validating model resistance to dataset- specific stopword artifacts.
-
-
CONCLUSION AND FUTURE SCOPE
This paper presented Cyber Shield, a scalable, real-time cyberbullying detection system resolving the fundamental tensions among classification accuracy, operational latency, and model transparency. By pairing sublinear TF-IDF
vectorization with a calibrated Linear Support Vector Machine, the framework achieves an accuracy of 92.9% and an F1 score of 94.4% across 367,245 annotated social media samples, operating with a sub-millisecond inference latency of
0.002 seconds. Deployed on an asynchronous FastAPI backend, the system sustains a throughput of 15,00020,000 requests per second, while SHAP and LIME integrations provide verifiable moderation auditing.
Future research directions will encompass: (1) integrating parameter-efficient fine-tuned compact language models (via LoRA and 4-bit quantization) to improve sensitivity to subtle, implicit sarcasm without incurring severe latency penalties; (2) expanding the architecture to process multimodal content by pairing text analysis with lightweight Optical Character Recognition (OCR) for image-based memes; (3) implementing continuous active learning pipelines that assimilate moderator feedback dynamically; and (4) exploring Federated Learning with differential privacy to enable collaborative model training across decentralized edge networks.
ACKNOWLEDGMENT
The authors thank the Department of Computer Science and Engineering for providing the computational infrastructure required for model training and concurrency testing, as well as the open-source community for maintaining the foundational machine learning and web frameworks utilized throughout this research.
REFERENCES
-
I. Isa, N. Omar, and M. Z. A. Nazri, “Cyberbullying detection using machine learning and deep learning techniques: A systematic review,” IEEE Access, vol. 11, pp. 5432154338, 2023.
-
H. Rosa, N. Pereira, R. Ribeiro, P. C. Ferreira, J. P. Carvalho, S. Oliveira, D. Coheur, P. Paulino, A. M. Veiga Simão, and I. Trancoso, “Automatic cyberbullying detection: A systematic review,” Computers in Human Behavior, vol. 93, pp. 333345, Apr. 2019.
-
K. Reynolds, A. Kontostathis, and L. Edwards, “Using machine learning to detect cyberbullying,” in Proc. 10th Int. Conf. Mach. Learn. Appl. (ICMLA), Honolulu, HI, USA, 2011, pp. 241244.
-
V. N. Vapnik, The Nature of Statistical Learning Theory, 2nd ed. New York, NY, USA: Springer-Verlag, 2000.
-
Z. Luo, J. Li, and J. Zhu, “Leveraging CNN-BiLSTM for multi-class cyberbullying detection,” in Proc. IEEE Int. Conf. Big Data (Big Data), Orlando, FL, USA, 2021, pp. 32143221.
-
E. Daraghmi, “From text to insight: An integrated CNN-BiLSTM-GRU model for cyberbullying detection,” IEEE Access, vol. 12, pp. 18234 18249, 2024.
-
K. Verma, S. K. Bharti, and R. K. Babu, “Can attention-based transformers explain or interpret cyberbullying detection?” Expert Systems with Applications, vol. 209, p. 118235, Dec. 2022.
-
V. Sanh, L. Debut, J. Chaumond, and T. Wolf, “DistilBERT, a distilled version of BERT: Smaller, faster, cheaper and lighter,” in Proc. NeurIPS EMC2 Workshop, Vancouver, BC, Canada, 2019, pp. 15.
-
A. Almomani, M. Al-Akhras, and M. Al-Duwairi, “Multi-level cyberbullying detection on social media platforms using deep learning ensembles,” Journal of Information Security and Applications, vol. 81, p. 103712, Mar. 2024.
-
C. D. Manning, P. Raghavan, and H. Schütze, Introduction to Information Retrieval. Cambridge, UK: Cambridge University Press, 2008.
-
M. T. Ribeiro, S. Singh, and C. Guestrin, “”Why should I trust you?”: Explaining the predictions of any classifier,” in Proc. 22nd ACM
SIGKDD Int. Conf. Knowl. Discovery Data Mining (KDD), San Francisco, CA, USA, 2016, pp. 11351144.
-
S. M. Lundberg and S.-I. Lee, “A unified approach to interpreting model predictions,” in Advances in Neural Information Processing Systems (NeurIPS 30), Long Beach, CA, USA, 2017, pp. 47654774.
