Enhancing Social Media Bias Detection With Contextual N-Grams
Keywords:
Bias detection; Social media; N-gram; Log-likelihood; Class imbalanceAbstract
Detecting bias in social media text requires methods that remain reliable under severe class imbalance while preserving interpretability. This study evaluates contextual n-gram bias detection using a unified probabilistic log-likelihood framework with Laplace smoothing. Five variants are compared across four benchmark datasets: Z-Score Baseline, NO_Trigram (unigram baseline), TRIG_NOPOS, TRIG (POS-aware), and PENTA. Results show that simpler lexical representations, especially unigram features, provide the most stable minority-class detection, whereas higher-order contextual encodings increase sparsity and reduce precision-recall balance in skewed settings. Although richer n-gram windows add expressiveness, they amplify estimation noise and do not consistently improve discrimination. On moderately imbalanced datasets, NO_Trigram achieves the strongest overall balance; on highly skewed data, all fixed-window n-gram models show degradation, with complex variants failing to deliver consistent gains over the baseline. These findings demonstrate that increasing n-gram complexity alone is insufficient for robust bias detection and motivate alternative context representations that can capture nuance without sacrificing statistical reliability.





