Building RUEmoCorp: Annotation Pipelines & Privacy-Preserving Dataset Curation for Roman Urdu
Engineering a 134K-sample multi-class emotion dataset and Transformer benchmark suite for low-resource South Asian NLP.
Lab Note & Technical Documentation
Overview
RUEmoCorp (Roman Urdu Emotion Corpus) is a large-scale emotion classification resource developed to address the lack of high-quality, openly available datasets for Roman Urdu, the Latin-script form of Urdu widely used in digital communication. The project focuses on building a reproducible data and evaluation foundation for low-resource affective NLP.
Dataset Engineering
RUEmoCorp was constructed as a curated, human-annotated corpus of Roman Urdu text collected from social media and conversational sources. The core release contains approximately 28,000 annotated samples spanning seven emotion categories: anger, disgust, fear, happiness, sadness, surprise, and a neutral/none class. The inclusion of a neutral category is an important engineering and modeling decision because real-world conversational data frequently contains text without an explicit emotional signal.
The corpus preserves natural characteristics of Roman Urdu, including orthographic variation, informal expressions, abbreviations, and code-switching with English. Rather than aggressively normalizing these properties, the pipeline retains linguistic variation that is relevant to real-world model deployment.
Annotation and Validation
Quality control was treated as a first-class component of the dataset engineering process. A stratified benchmark of 700 samples was independently evaluated by four annotators representing multiple Pakistani institutions. The resulting Fleiss' κ of 0.6588 indicates substantial inter-annotator agreement for the seven-class task. This validation layer provides an empirical basis for assessing model behavior beyond automated labeling alone.
Modeling Pipeline
RUEmoCorp serves as the training corpus for roman-urdu-emotion-xlmr-v2, based on XLM-RoBERTa with a custom two-layer classification head. The reported evaluation on the in-distribution test set reaches a macro F1 of approximately 0.9896. The associated work also compares transformer and conventional machine-learning baselines, establishing a reproducible benchmark for Roman Urdu emotion classification.
Engineering Significance
The primary contribution of RUEmoCorp is not only its size but its emphasis on data quality, annotation validation, reproducibility, and low-resource language infrastructure. The project demonstrates an end-to-end workflow spanning corpus construction, privacy-aware preprocessing, expert annotation, agreement analysis, model training, and public release.
From an engineering perspective, RUEmoCorp establishes a foundation on which future Roman Urdu NLP systems can be evaluated consistently. Its release also supports research into multilingual and code-switched affective computing while preserving the linguistic characteristics that make Roman Urdu challenging for conventional NLP pipelines.
Project Repositories & Artifacts
Data Annotator & Validator
Lead Engineer & Researcher
Building a Reproducible NLP Dataset Quality Auditor
We built `inference-audit`, a reproducible Python toolkit for systematically auditing NLP dataset quality across five core dimensions. The pipeline combines deterministic checks for label distribution, missing values, near-duplicates, language contamination, and annotation consistency. Extensive testing and real-corpus evaluation exposed performance and reliability issues, leading to targeted fixes and measurable optimizations. The final system emphasizes reproducibility, explicit failure states, citable evidence, and transparent limitations.
Read Report llm-eval-kitBuilding a Deterministic Offline Evaluation Engine for LLM Systems
We redesigned llm-eval-kit into a deterministic, zero-network evaluation pipeline built for reliable continuous evaluation in CI/CD environments. The new architecture introduces a decoupled criteria registry, fail-fast orchestration, hybrid semantic and symbolic verification, and graceful handling of inapplicable evaluation criteria. The implementation reached 93% test coverage across 88 tests, with cross-version CI validation from Python 3.9 to 3.12. The work demonstrates how carefully defined evaluation contracts and offline heuristics can provide reproducible, privacy-preserving model assessment without relying on external LLM-as-a-judge APIs.
Read Report