INFERENCELAB
Back to Engineering Journals
Engineering FellowshipCohort 2026
RUEmoCorp

Building RUEmoCorp: Annotation Pipelines & Privacy-Preserving Dataset Curation for Roman Urdu

Engineering a 134K-sample multi-class emotion dataset and Transformer benchmark suite for low-resource South Asian NLP.

Published: 2026-04-10 7 min read

Lab Note & Technical Documentation

## Background & Problem Statement Roman Urdu is widely spoken across digital channels in South Asia, yet lacks curated multi-class emotion benchmarks. Existing datasets suffer from noisy transcriptions, inconsistent spelling variants, and lack of annotator agreement metrics. ### Key Pipeline Highlights 1. **Fleiss Kappa Verification**: Achieved Inter-Annotator Agreement of Fleiss κ = 0.658 (substantial agreement) across 5 annotator pools. 2. **Differential Privacy Filtration**: Synthetic embeddings generated to strip PII before publishing. 3. **HuggingFace Hub Release**: Instant integration with `datasets.load_dataset('Inferencelab/RUEmoCorp')`.

Project Repositories & Artifacts

Technologies
NLPRoman UrduDataset CurationHuggingFaceTransformers