INFERENCELAB
Back home
Research output

Where rigorous research meets real-world impact

INFERENCE Lab publishes across speech intelligence, natural language processing, human-centered sensing, and applied AI — with each study designed from the ground up to be reproducible, citable, and deployable. Our work spans clinical voice diagnostics, low-resource corpora, cognitive fatigue detection, and post-quantum cryptography.

Core Focus

Research Areas & Domains

Our lab focuses on fundamental problems where theoretical rigor can be translated into open software, benchmarks, and clinical-grade tools.

Domain 01

Low-Resource Speech Intelligence & Acoustic Biomarkers

Pioneering clinical-grade vocal fatigue screening, speaker verification, and emotion recognition from raw speech. Developing novel neural representations (e.g. ECAPA-TDNN-VHE) that operate robustly with minimal training data.

Key Topics & Pipelines:

Vocal Fatigue QuantificationAcoustic Feature ExtractionDisentangled Speech RepresentationsReal-Time Inference Engines
Domain 02

Low-Resource NLP & Regional Language Corpora

Curating large-scale foundational corpora and benchmarking language models for Roman Urdu and under-represented South Asian dialects. We build data-centric pipelines that enforce high inter-annotator agreement and privacy-preserving embeddings.

Key Topics & Pipelines:

RUEmoCorp & RUDaSA DatasetsSubjective NLP & Disagreement WeightingTransformer CalibrationMulti-Seed Empirical Benchmarks
Domain 03

Human-Centered Sensing & Cognitive Ergonomics

Investigating unobtrusive digital biomarkers—such as cursor kinematics and workstation interaction logs—as indicators of cognitive workload, acute stress, and fatigue in knowledge work and clinical environments.

Key Topics & Pipelines:

Workplace Interaction LogsExplainable ML for ErgonomicsJoint Cognitive Systems ReconfigurationPrivacy-Preserving Telemetry
Domain 04

Deployable AI Systems & Security

Bridging scientific discovery and production engineering. Every research artifact is packaged into pip-installable libraries (PyPI), verifiable registries, and benchmarked cryptographic schemes.

Key Topics & Pipelines:

Image Encryption AlgorithmsDeterministic Model PipelinesOpen-Source Python Libraries
Publications & Preprints

Original Papers (9)

  • Under Review2026

    Disagreement-Weighted Classification and Calibration-Aware Training for Subjective NLP Tasks: A Multi-Seed Empirical Study

    Introduces a disagreement-weighted classification framework that explicitly models annotator disagreement to improve learning in subjective NLP tasks.

    Language Resources and Evaluation · National University of Singapore
  • Under Review2026

    Ergonomic Intervention as Joint Cognitive System Reconfiguration: A Qualitative Case Study in Healthcare

    Explores ergonomic intervention as a joint cognitive system reconfiguration to improve coordination between healthcare professionals, technologies, and work environments.

    KU Leuven (Belgium) · King Saud University · Ergonomics
  • In Progress2026

    A Post-Quantum Hybrid Key Encapsulation Mechanism with Lorenz Chaotic Diffusion for Authenticated Medical Image Encryption

    Combines post-quantum cryptography and chaos-based diffusion to provide strong confidentiality, integrity, and resistance against quantum-enabled attacks on sensitive medical imagery.

    Open to Collaboration · Medical Image Cryptography
  • Under Review2026

    RUEmoCorp: A Large-Scale Roman Urdu Emotion Corpus & Benchmark Suite

    First large-scale Roman Urdu emotion corpus — 134K labeled samples with Fleiss κ = 0.658 (substantial agreement), multi-institute annotation, fully open-source on HuggingFace and Harvard Dataverse.

    Language Resources and Evaluation (Springer)DOI
  • Published Preprint2026

    RUDaSA: Roman Urdu Dataset for Sentiment Analysis — A Large-Scale, Curated Corpus with Privacy-Preserving Embeddings and Competitive Benchmarking of Transformer Models

    Large-scale Roman Urdu sentiment corpus built via privacy-preserving embedding pipelines. Benchmarks state-of-the-art Transformer models — addressing a critical gap in low-resource South Asian NLP.

    PreprintDOI
  • Published Preprint2025

    Forecast-Based Decision Support System for Mango Malformation

    Time-series forecasting and smart-agriculture DSS — demonstrated 50–60% yield improvement through data-driven intervention.

    PreprintDOI
  • In Progress2026

    DASER-Net: Disentangled Adversarial Speech Emotion Recognition with Hierarchical Bottleneck Fusion for Cross-Corpus Generalization

    Cross-corpus speech emotion recognition study developing disentangled and domain-invariant representations for robust emotion classification across unseen acoustic domains.

    Open to Collaboration
  • In Progress2026

    Cursor Kinematics as a Privacy-Preserving Indicator of Cognitive Fatigue and Stress in Knowledge Work: An Explainability-Driven Analysis of Within-Person Interaction Logs

    Within-person study investigating cursor kinematics as a privacy-preserving indicator of cognitive fatigue and acute stress using explainable machine learning on workplace interaction logs.

    Open to Collaboration
  • Published2026

    Automated Vocal Fatigue Screening in Professional Voice Users: Development and Occupational Validation of an Automated Assessment System

    Development and occupational validation of an automated vocal load assessment tool for professional voice users — clinical-grade speech analysis in production.

    Journal of Voice · Hanyang University, Republic of KoreaDOI