INFERENCELAB
Back to Engineering Journals
ResearchCohort 2026
RUEmoCorp

Building RUEmoCorp: Annotation Pipelines & Privacy-Preserving Dataset Curation for Roman Urdu

Engineering a 134K-sample multi-class emotion dataset and Transformer benchmark suite for low-resource South Asian NLP.

Published: 2026-04-10 7 min read

Lab Note & Technical Documentation

Overview

RUEmoCorp (Roman Urdu Emotion Corpus) is a large-scale emotion classification resource developed to address the lack of high-quality, openly available datasets for Roman Urdu, the Latin-script form of Urdu widely used in digital communication. The project focuses on building a reproducible data and evaluation foundation for low-resource affective NLP.

Dataset Engineering

RUEmoCorp was constructed as a curated, human-annotated corpus of Roman Urdu text collected from social media and conversational sources. The core release contains approximately 28,000 annotated samples spanning seven emotion categories: anger, disgust, fear, happiness, sadness, surprise, and a neutral/none class. The inclusion of a neutral category is an important engineering and modeling decision because real-world conversational data frequently contains text without an explicit emotional signal.

The corpus preserves natural characteristics of Roman Urdu, including orthographic variation, informal expressions, abbreviations, and code-switching with English. Rather than aggressively normalizing these properties, the pipeline retains linguistic variation that is relevant to real-world model deployment.

Annotation and Validation

Quality control was treated as a first-class component of the dataset engineering process. A stratified benchmark of 700 samples was independently evaluated by four annotators representing multiple Pakistani institutions. The resulting Fleiss' κ of 0.6588 indicates substantial inter-annotator agreement for the seven-class task. This validation layer provides an empirical basis for assessing model behavior beyond automated labeling alone.

Modeling Pipeline

RUEmoCorp serves as the training corpus for roman-urdu-emotion-xlmr-v2, based on XLM-RoBERTa with a custom two-layer classification head. The reported evaluation on the in-distribution test set reaches a macro F1 of approximately 0.9896. The associated work also compares transformer and conventional machine-learning baselines, establishing a reproducible benchmark for Roman Urdu emotion classification.

Engineering Significance

The primary contribution of RUEmoCorp is not only its size but its emphasis on data quality, annotation validation, reproducibility, and low-resource language infrastructure. The project demonstrates an end-to-end workflow spanning corpus construction, privacy-aware preprocessing, expert annotation, agreement analysis, model training, and public release.

From an engineering perspective, RUEmoCorp establishes a foundation on which future Roman Urdu NLP systems can be evaluated consistently. Its release also supports research into multilingual and code-switched affective computing while preserving the linguistic characteristics that make Roman Urdu challenging for conventional NLP pipelines.

Project Repositories & Artifacts

Technologies
NLPRoman UrduDataset CurationHuggingFaceTransformers