🏆
Global Publishing Platform
Serving Researchers Since 2012

Cyber Shield: A Scalable Real-Time Framework for Automated Cyberbullying Detection using Classical Machine Learning and Asynchronous Web Architecture

DOI : 10.5281/zenodo.23078465
Download Full-Text PDF Cite this Publication

Text Only Version

Cyber Shield: A Scalable Real-Time Framework for Automated Cyberbullying Detection using Classical Machine Learning and Asynchronous Web Architecture

Abhijith B S, Adwaith Raj R, Afsal A, Aswin Saj

Department of Computer Science and Engineering

Vidya Academy of Science and Technology Technical Campus, Thiruvananthapuram, Kerala, India

Abstract – The rapid expansion of social media has escalated digital harassment, creating an imperative for low-latency, scalable moderation architectures. This paper introduces Cyber Shield, an end-to-end detection pipeline coupling classical statistical classifiers with an asynchronous web execution tier for live remediation of toxic online discourse. Three supervised classification paradigms Stochastic Gradient Descent Linear Support Vector Machines (SGD-SVM), Logistic Regression (LR), and Naive Bayes (NB)are evaluated across two textual representation strategies: Term Frequency-Inverse Document Frequency (TF-IDF) and Latent Semantic Analysis (LSA) implemented via Truncated Singular Value Decomposition. Benchmarked across a consolidated corpus of 367,245 annotated social entries, the sublinear TF-IDF + Linear SVM configuration yields optimal discrimination, attaining a 92.9% accuracy, 96.0% precision, 92.8% recall, and an F1 metric of 94.4%, requiring only 0.002 seconds per batch evaluation. The serving layer, engineered on an Asynchronous Server Gateway Interface (ASGI) utilizing FastAPI and Uvicorn, sustains 15,00020,000 transactions per second, delivering an order-of-magnitude throughput gain over synchronous WSGI designs. Model interpretability is integrated via Local Interpretable Model-Agnostic Explanations (LIME) and Shapley Additive exPlanations (SHAP), providing granular, token-level transparency for moderation decisions. The empirical evidence confirms that lightweight, optimized linear classifiers over high- dimensional sparse representations offer the most pragmatic convergence of detection efficacy, sub-millisecond execution, compute efficiency, and auditability for web-scale moderation.

Keywords cyberbullying detection; natural language processing; support vector machines; TF-IDF; FastAPI; explainable artificial intelligence; asynchronous inference; content moderation

  1. INTRODUCTION

    Ubiquitous digital connectivity has transformed collaborative communication while simultaneously lowering barriers to hostility and systematic harassment. Cyberbullyingdefined by persistent, deliberate hostility mediated through networked platforms toward individuals with limited capacity for self-defensetargets protected attributes including ethnicity, gender, sexual identity, physical appearance, and disability [1]. Unlike physical altercations, digital abuse transcends spatial and temporal constraints, providing bad actors with an asymmetrical reach amplified by screen anonymity. Large-scale cross-sectional studies reveal that digital abuse affects upwards of 60% of youth populations and roughly 40% of adult internet users [1].

    Contemporary containment frameworks remain vulnerable to operational bottlenecks. Human-in-the-loop review queues cannot digest the millions of incoming streams produced per minute, and prolonged exposure to toxic content inflicts documented psychological harm on human annotators. Conversely, rule-based lexicon systems rely on static string matching, causing elevated false alarms on benign linguistic nuances while failing against basic obfuscations like leetspeak, deliberate spelling variations, and phonetic transliterations [2].

    Consequently, research has pivoted toward statistical natural language processing and machine learning. While attention-based deep architectures (e.g., BERT, RoBERTa) demonstrate high benchmark scores, their multi-layered

    parameter footprints entail prohibitive compute costs, high latency, and excessive memory overhead during peak serving loads [3]. Production deployments require an operational compromise spanning classification robustness, operational execution speed, and decision transparency.

    To address these operational constraints, this study presents Cyber Shield, a high-throughput moderation engine coupling tuned classical classifiers with an asynchronous execution framework. The key contributions of this research are as follows:

    1. A systematic comparative analysis of six algorithmic configurations evaluated across an aggregated benchmark corpus containing 367,245 validated social text entries.

    2. Quantitative validation that an SGD-optimized Linear SVM operating over sublinear TF-IDF tokens achieves top operational utility (92.9% accuracy, 94.4% F1 score) with an average batch latency of 0.002 seconds.

    3. Implementation of a non-blocking asynchronous serving pipeline via FastAPI and Uvicorn, consistently processing 15,000 to 20,000 transactions per second.

    4. Integration of dual interpretability layers (SHAP and LIME) to facilitate compliance auditing and token-level algorithmic explainability.

  2. RELATED WORK

      1. Classical Machine Learning for Cyberbullying Detection

        Statistical machine learning has long anchored automated lexical processing. Isa et al. [1] established standard baseline metrics leveraging Naive Bayes, Logistic Regression, and linear boundary estimators across n-gram matrices. Naive Bayes operates on strong attribute-independence assumptions, remaining computationally cheap but underperforming in the presence of strong phrase interdependencies. Support Vector Machines define optimal boundary hyperplanes across wide feature spaces, showing inherent resistance to input noise and dimensional sparseness [4]. Logistic Regression yields direct posterior probabilities, making it an established baseline; yet standard bag-of-words approximations discard sequential word structure, complicating the detection of nuanced contextual hostility.

      2. Deep Learning and Attention Architectures

        To model sequential dependencies, hybrid deep networks have been widely explored. Luo et al. [5] and Daraghmi [6] demonstrated that stacking Convolutional Neural Networks (CNNs) with Bidirectional Long Short-Term Memory (BiLSTM) units captures localized phrase patterns together with temporal semantic cues, achieving F1 values spanning

        0.78 to 0.98. Multi-head transformer encoders (e.g., BERT, RoBERTa) leverage cross-token self-attention to resolve polysemy and discern veiled harassment [7]. However, as demonstrated by Verma et al. [7] and Sanh et al. [8], the massive parameter budgets and matrix multiplication requirements of deep transformers lead to latencies up to two orders of magnitude higher than linear models, restricting real- time deployment on standard server hardware.

      3. Scalability and Asynchronous Deployment

        Deploying inference engines within high-traffic web environments requires robust handling of volatile query volumes. Almomani et al. [9] documented that classical synchronous WSGI wrappers (such as standard Flask configurations) suffer rapid thread depletion and cascading latency degradation during traffic surges. Asynchronous Server Gateway Interface (ASGI) frameworks resolve this bottleneck by utilizing event loops and non-blocking I/O routines, maintaining high connection concurrency with minimal compute overhead.

      4. Feature Extraction Paradigms

    Model discrimination is heavily governed b the initial token projection strategy. Sublinear TF-IDF balances term occurrences against cross-document distribution, dampening excessively frequent words while prioritizing distinct discriminative tokens [10]. In contrast, Latent Semantic Analysis (LSA) applies Truncated Singular Value Decomposition (SVD) to project broad term frequencies into compact continuous semantic vectors. Although LSA handles synonymy well, the runtime projection matrix multiplication introduces non-negligible processing latency during live serving.

  3. SYSTEM DESIGN AND METHODOLOGY

    The Cyber Shield architecture is organized as a decoupled, modular pipeline coordinating data ingestion, text sanitization, tokenization, model inference, and real-time explanation generation.

    1. Dataset Aggregation and Preprocessing

      The validation set was constructed by aggregating two distinct open-access social harassment repositories. Dataset 1 includes 115,802 instances (12.8% positive class), and Dataset

      2 contains 251,443 instances (14.1% positive class). The merged corpus consists of 367,245 annotated records exhibiting a 13.7% overall positive prevalence and a mean string length of 353.9 characters. To prevent majority-class decision bias, stratified splitting was applied to maintain balanced class ratios across the 80/20 train-test division.

      The raw string records pass through a deterministic preprocessing sequence: (1) URL removal via regular expression filters (http\S+|www\.\S+); (2) user handle (@mention) and digit stripping; (3) removal of extraneous punctuation and special symbols while preserving single space delimiters; (4) case normalization to lowercase; and (5) token normalization through whitespace collapsing.

    2. Feature Vectorization Architectures

      Two feature extraction pipelines are implemented: (1) Sublinear TF-IDF: Configured for unigram and bigram extraction, capped at a vocabulary limit of 15,000 features, applying sublinear scaling (1 + log(TF)) to dampen dominant tokens; (2) LSA Vectorization: Chaining a 20,000-term CountVectorizer with Truncated SVD to project the lexical space into 200 latent semantic dimensions, followed by Euclidean L2 normalization to enforce unit length.

    3. Classification Algorithms

      Three supervised algorithm families were systematically trained across each feature space: (1) Linear SVM / SGD: A Linear Support Vector Classifier trained via Stochastic Gradient Descent using an elastic net penalty, encapsulated in a 3-fold cross-validated Platt scaling wrapper (CalibratedClassifierCV) to yield reliable class posterior probabilities; (2) Logistic Regression: Regularized with L2 penalty using the liblinear coordinate descent solver; (3) Naive Bayes: Evaluated via Multinomial Naive Bayes over sparse TF-IDF matrices, and Gaussian Naive Bayes over continuous LSA semantic projections.

    4. Asynchronous Web Architecture

      The backend inference engine is constructed with FastAPI and deployed via the Uvicorn ASGI server. To eliminate redundant runtime initialization, pre-trained model weights and vectorizers are loaded into an in-memory dictionary cache during the application startup event. Incoming HTTP POST requests containing raw text payloads are parsed asynchronously, transformed into feature vectors via the cached vectorizer, and evaluated by the classifier within an asynchronous task loop.

      The client interface is implemented as a Single Page Application (SPA) built with React 19 and Vite 6, presenting real-time toxicity scores, inference execution metrics, and interpretability graphs.

    5. Explainable AI Integration

    To address algorithmic opacity in automated moderation, two complementary interpretability mechanisms were integrated: (1) LIME (Local Interpretable Model-Agnostic Explanations): Computes local linear surrogate approximations around target instances, identifying specific abusive words to provide intuitive justifications for moderation staff [11]; (2) SHAP (Shapley Additive exPlanations): Employs cooperative game theory principles to calculate global Shapley values across the corpus, verifying that feature weights align with authentic hostility indicators rather than spurious dataset-specific correlations [12].

  4. EXPERIMENTAL RESULTS AND DISCUSSION

    1. Classification Performance

      Table I reports the performance across classical supervised configurations on the 20% validation split. Table II contextualizes these empirical outcomes against contemporary deep learning benchmarks reported in recent cyberbullying detection literature.

      TABLE I. PERFORMANCE OF CLASSICAL ML CONFIGURATIONS

      Algorithm

      Feature

      Acc.

      Prec.

      Recall

      F1

      Time(s)

      SGD / Linear SVM

      TF-IDF

      92.9%

      96.0%

      92.8%

      94.4%

      0.002

      Bagging Classifier

      TF-IDF

      92.7%

      96.7%

      92.3%

      94.5%

      0.210

      Decision Tree

      TF-IDF

      92.5%

      95.5%

      92.9%

      94.2%

      0.024

      Random Forest

      TF-IDF

      91.7%

      94.8%

      92.3%

      93.6%

      2.890

      K-Nearest Neighbours

      TF-IDF

      85.8%

      89.5%

      88.7%

      89.1%

      20.495

      Naive Bayes (Gauss.)

      LSA

      84.0%

      79.0%

      62.0%

      69.5%

      <0.010

      TABLE II. BENCHMARK COMPARISON WITH DEEP LEARNING

      Model Architecture

      Representation

      Accuracy

      F1 Score

      Ref.

      BiLSTM

      Word2Vec

      Up to 98.0%

      0.78 – 0.97

      [5]

      CNN-BiLSTM- GRU

      Dense Embeddings

      Up to 99.0%

      0.95 – 0.98

      [6]

      BERT (Base)

      Attention

      92.0 – 94.8%

      0.92 – 0.96

      [7]

      RoBERTa

      Ensemble

      Attention

      ~96.0%

      ~0.96

      [9]
    2. Analysis of Classifier Outcomes

      The calibrated Linear SVM (SGD) operating on TF-IDF features recorded the highest operational efficacy, achieving 92.9% accuracy and an F1 score of 94.4% with an inference latency of 0.002 seconds per batch. The 15,000-dimensional sparse feature space enables effective linear separation with minimal update overhead.

      While the Bagging Classifier achieved a slightly higher precision of 96.7%, its prediction latency of 0.210 seconds represents a hundred-fold delay compared to the linear model. In content moderation ecosystems where minimizing false accusations is critical, Bagging provides strong reliability, but its computational cost restricts high-throughput operational scaling.

      Conversely, the K-Nearest Neighbours classifier degraded markedly, achieving only 85.8% accuracy while requiring

      20.495 seconds per inference batch. In 15,000-dimensional sparse feature spaces, distance concentration severely weakens Euclidean discriminative power, rendering neighborhood lookups computationally prohibitive. Gaussian Naive Bayes over LSA reached an F1 score of only 69.5%, demonstrating that conditional independence assumptions falter when toxicity stems from multi-word contextual dependencies.

    3. Asynchronous API Performance

      Stress testing conducted under simulated concurrent loads (1,000 to 50,000 simultaneous requests) demonstrated that the FastAPI + Uvicorn deployment sustained an operational throughput between 15,000 and 20,000 requests per second. Under identical hardware conditions, an equivalent synchronous Flask WSGI server collapsed at 2,150 requests per second due to thread pool starvation. In isolated benchmark suites, the asynchronous framework completed execution in 1.56 seconds versus 27.8 seconds for the synchronous baselinea 17.8x acceleration.

      Memory provisioning profiles indicated an operational footprint of approximately 1.1 GB per worker process. Sustaining four concurrent workers required ~4.5 GB of system memory, confirming that the architecture operates comfortably on entry-level commodity cloud instances.

    4. Explainability Findings

    Instance-level evaluations utilizing LIME showed that the Linear SVM model accurately identified abusive unigrams and bigrams with high fidelity to human judgment. Global feature importance computed via SHAP confirmed that identity- targeted slurs and hostile imperatives received dominant positive attribution weights, while neutral grammatical tokens hovered near zero, validating model resistance to dataset- specific stopword artifacts.

  5. CONCLUSION AND FUTURE SCOPE

This paper presented Cyber Shield, a scalable, real-time cyberbullying detection system resolving the fundamental tensions among classification accuracy, operational latency, and model transparency. By pairing sublinear TF-IDF

vectorization with a calibrated Linear Support Vector Machine, the framework achieves an accuracy of 92.9% and an F1 score of 94.4% across 367,245 annotated social media samples, operating with a sub-millisecond inference latency of

0.002 seconds. Deployed on an asynchronous FastAPI backend, the system sustains a throughput of 15,00020,000 requests per second, while SHAP and LIME integrations provide verifiable moderation auditing.

Future research directions will encompass: (1) integrating parameter-efficient fine-tuned compact language models (via LoRA and 4-bit quantization) to improve sensitivity to subtle, implicit sarcasm without incurring severe latency penalties; (2) expanding the architecture to process multimodal content by pairing text analysis with lightweight Optical Character Recognition (OCR) for image-based memes; (3) implementing continuous active learning pipelines that assimilate moderator feedback dynamically; and (4) exploring Federated Learning with differential privacy to enable collaborative model training across decentralized edge networks.

ACKNOWLEDGMENT

The authors thank the Department of Computer Science and Engineering for providing the computational infrastructure required for model training and concurrency testing, as well as the open-source community for maintaining the foundational machine learning and web frameworks utilized throughout this research.

REFERENCES

  1. I. Isa, N. Omar, and M. Z. A. Nazri, “Cyberbullying detection using machine learning and deep learning techniques: A systematic review,” IEEE Access, vol. 11, pp. 5432154338, 2023.

  2. H. Rosa, N. Pereira, R. Ribeiro, P. C. Ferreira, J. P. Carvalho, S. Oliveira, D. Coheur, P. Paulino, A. M. Veiga Simão, and I. Trancoso, “Automatic cyberbullying detection: A systematic review,” Computers in Human Behavior, vol. 93, pp. 333345, Apr. 2019.

  3. K. Reynolds, A. Kontostathis, and L. Edwards, “Using machine learning to detect cyberbullying,” in Proc. 10th Int. Conf. Mach. Learn. Appl. (ICMLA), Honolulu, HI, USA, 2011, pp. 241244.

  4. V. N. Vapnik, The Nature of Statistical Learning Theory, 2nd ed. New York, NY, USA: Springer-Verlag, 2000.

  5. Z. Luo, J. Li, and J. Zhu, “Leveraging CNN-BiLSTM for multi-class cyberbullying detection,” in Proc. IEEE Int. Conf. Big Data (Big Data), Orlando, FL, USA, 2021, pp. 32143221.

  6. E. Daraghmi, “From text to insight: An integrated CNN-BiLSTM-GRU model for cyberbullying detection,” IEEE Access, vol. 12, pp. 18234 18249, 2024.

  7. K. Verma, S. K. Bharti, and R. K. Babu, “Can attention-based transformers explain or interpret cyberbullying detection?” Expert Systems with Applications, vol. 209, p. 118235, Dec. 2022.

  8. V. Sanh, L. Debut, J. Chaumond, and T. Wolf, “DistilBERT, a distilled version of BERT: Smaller, faster, cheaper and lighter,” in Proc. NeurIPS EMC2 Workshop, Vancouver, BC, Canada, 2019, pp. 15.

  9. A. Almomani, M. Al-Akhras, and M. Al-Duwairi, “Multi-level cyberbullying detection on social media platforms using deep learning ensembles,” Journal of Information Security and Applications, vol. 81, p. 103712, Mar. 2024.

  10. C. D. Manning, P. Raghavan, and H. Schütze, Introduction to Information Retrieval. Cambridge, UK: Cambridge University Press, 2008.

  11. M. T. Ribeiro, S. Singh, and C. Guestrin, “”Why should I trust you?”: Explaining the predictions of any classifier,” in Proc. 22nd ACM

    SIGKDD Int. Conf. Knowl. Discovery Data Mining (KDD), San Francisco, CA, USA, 2016, pp. 11351144.

  12. S. M. Lundberg and S.-I. Lee, “A unified approach to interpreting model predictions,” in Advances in Neural Information Processing Systems (NeurIPS 30), Long Beach, CA, USA, 2017, pp. 47654774.