Back to Engineering Journals
Engineering FellowshipCohort 2026
RUEmoCorpBuilding RUEmoCorp: Annotation Pipelines & Privacy-Preserving Dataset Curation for Roman Urdu
Engineering a 134K-sample multi-class emotion dataset and Transformer benchmark suite for low-resource South Asian NLP.
Published: 2026-04-10 7 min read
Lab Note & Technical Documentation
## Background & Problem Statement
Roman Urdu is widely spoken across digital channels in South Asia, yet lacks curated multi-class emotion benchmarks. Existing datasets suffer from noisy transcriptions, inconsistent spelling variants, and lack of annotator agreement metrics.
### Key Pipeline Highlights
1. **Fleiss Kappa Verification**: Achieved Inter-Annotator Agreement of Fleiss κ = 0.658 (substantial agreement) across 5 annotator pools.
2. **Differential Privacy Filtration**: Synthetic embeddings generated to strip PII before publishing.
3. **HuggingFace Hub Release**: Instant integration with `datasets.load_dataset('Inferencelab/RUEmoCorp')`.
Project Repositories & Artifacts
Contributors
Muhammad Khubaib Ahmad
Lead Engineer
Technologies
NLPRoman UrduDataset CurationHuggingFaceTransformers