Where rigorous research meets real-world impact
Inference Lab publishes across speech intelligence, natural language processing, human-centered sensing, and applied AI — with each study designed from the ground up to be reproducible, citable, and deployable. Our work spans clinical voice diagnostics and cross-corpus emotion recognition to large-scale low-resource corpora, cognitive fatigue detection, and data-driven agricultural forecasting. Every output is developed in collaboration with international academic partners and released with open datasets, auditable model pipelines, and a permanent DOI.
- Accepted2026
Automated Vocal Fatigue Screening in Professional Voice Users: Development and Occupational Validation of an Automated Assessment System
Development and occupational validation of an automated vocal load assessment tool for professional voice users — clinical-grade speech analysis in production.
Journal of Voice · Hanyang University, Republic of Korea - Under Review2026
Disagreement-Weighted Classification and Calibration-Aware Training for Subjective NLP Tasks: A Multi-Seed Empirical Study
Introduces a disagreement-weighted classification framework that explicitly models annotator disagreement to improve learning in subjective NLP task.
Language Resources and Evaluation · National University of Singapore - Under Review2026
Ergonomic Intervention as Joint Cognitive System Reconfiguration: A Qualitative Case Study in Healthcare
Explores ergonomic intervention as a joint cognitive system reconfiguration to improve coordination between healthcare professionals, technologies, and work environments.
KU Leuven (Belgium), King Saud University - Published Preprint2026
RUEmoCorp: A Large-Scale Roman Urdu Emotion Corpus & Benchmark Suite
First large-scale Roman Urdu emotion corpus — 134K labeled samples with Fleiss κ = 0.658 (substantial agreement), multi-institute annotation, fully open-source on HuggingFace and Harvard Dataverse.
Language Resources and Evaluation (Springer)DOI - Published Preprint2025
RUDaSA: Roman Urdu Dataset for Sentiment Analysis — A Large-Scale, Curated Corpus with Privacy-Preserving Embeddings and Competitive Benchmarking of Transformer Models
Large-scale Roman Urdu sentiment corpus built via privacy-preserving embedding pipelines. Benchmarks state-of-the-art Transformer models — addressing a critical gap in low-resource South Asian NLP.
Research Square · PreprintDOI - Published Preprint2025
Data-Centric Roman Urdu NLP: Dataset Curation & Model Benchmarking
Largest high-quality Roman Urdu sentiment dataset via privacy-preserving embedding pipelines — SOTA 0.84 accuracy, 0.83 Macro-F1.
Zenodo · PreprintDOI - Published Preprint2025
Forecast-Based Decision Support System for Mango Malformation
Time-series forecasting and smart-agriculture DSS — demonstrated 50–60% yield improvement through data-driven intervention.
Zenodo · PreprintDOI - In Progress2026
DASER-Net: Disentangled Adversarial Speech Emotion Recognition with Hierarchical Bottleneck Fusion for Cross-Corpus Generalization
Cross-corpus speech emotion recognition study developing disentangled and domain-invariant representations for robust emotion classification across unseen acoustic domains.
Open to collaboration - In Progress2026
Cursor Kinematics as a Privacy-Preserving Indicator of Cognitive Fatigue and Stress in Knowledge Work: An Explainability-Driven Analysis of Within-Person Interaction Logs
Within-person study investigating cursor kinematics as a privacy-preserving indicator of cognitive fatigue and acute stress using explainable machine learning on workplace interaction logs.
Open to collaboration - In Progress2026
A Post-Quantum Hybrid Key Encapsulation Mechanism with Lorenz Chaotic Diffusion for Authenticated Medical Image Encryption
Combines post-quantum cryptography and chaos-based diffusion to provide strong confidentiality, integrity, and resistance against quantum-enabled attacks.
Open to collaboration