Subword Representations Close the Out-of- Vocabulary Gap in Minimal-Supervision Part-of-Speech Tagging for Assamese

Authors

  • Biswajit Sarma
  • Rupam Baruah
  • Diganta Baishya

Keywords:

Assamese, part-of-speech tagging, low-resource NLP, minimal supervision, subword representations, character n-grams, FastText, out-of-vocabulary, negative results

Abstract

Minimal-supervision part-of-speech (POS) tagging pipelines for low-resource languages are typically improved by refining the learning algorithm, the supervision signal, or the labelled seed. We report a study in which five such refinements were applied to an established evidence-gated tagging pipeline for Assamese and all five failed, four of them reducing accuracy, before a diagnostic revealed why: on a 44,572-token Assamese corpus with a labelled seed of only 142–228 sentences, 82–86% of all remaining errors fall on word types absent from the pipeline’s accumulated evidence pool. The classifier is already accurate on words it has seen (94.4%) and collapses on words it has not (74.4%); the widely-reported weakness of the adjective category is shown to be largely an artifact of this split, since adjective accuracy is 0.894 on seen types but 0.481 on unseen ones. Motivated by this localization, we add sub word representations—character n-gram indicator features and FastText embedding trained on unlabelled in-domain text—to the pipeline’s backfill classifier. The combination improves tagging accuracy by +1.78 points at a seed of three sentences per distinct sentence length and
+1.46 points at five, improving in all ten paired replicate draws, and consumes no additional annotation. The gain is concentrated exactly where the diagnostic predicted: unseen-type accuracy rises from 0.744 to 0.774 and adjective accuracy on unseen types from 0.492 to 0.533, the first intervention in this line of work to move adjective recall at all. We report the five negative results alongside the positive one, since each delineates a class of technique that does not transfer to this regime, and we show that character n-grams alone—requiring no embedding and no training step—recover roughly 88% of the benefit.

Downloads

Published

2026-10-05

How to Cite

Sarma, B., Baruah, R., & Baishya, D. (2026). Subword Representations Close the Out-of- Vocabulary Gap in Minimal-Supervision Part-of-Speech Tagging for Assamese. International Journal of Artificial Intelligence and Machine Learning, 6(13s), 1205–1212. Retrieved from https://www.svedbergopen.com/index.php/ijaiml/article/view/2822