INFERENCELAB
Engineering
Technical Publications

Engineering Journal

Detailed Lab Notes and engineering reports documenting how our projects are designed, architected, benchmarked, and deployed. Built by Engineering Fellows and lab contributors.

Engineering FellowshipCohort 2026
2026-09-19 6 min read
inference-audit

Building a Reproducible NLP Dataset Quality Auditor

We built inference-audit, a reproducible Python toolkit for systematically auditing NLP dataset quality across five core dimensions. The pipeline combines deterministic checks for label distribution, missing values, near-duplicates, language contamination, and annotation consistency. Extensive testing and real-corpus evaluation exposed performance and reliability issues, leading to targeted fixes and measurable optimizations. The final system emphasizes reproducibility, explicit failure states, citable evidence, and transparent limitations.

Contributors
Khadija Faisal(Operations & Research Associate)
Muhammad Shoaib Altaf(Implementation Engineer)
NLPTrustworthy AILow-Resource NLP
Engineering FellowshipCohort 2026
2026-09-19 6 min read
llm-eval-kit

Building a Deterministic Offline Evaluation Engine for LLM Systems

We redesigned llm-eval-kit into a deterministic, zero-network evaluation pipeline built for reliable continuous evaluation in CI/CD environments. The new architecture introduces a decoupled criteria registry, fail-fast orchestration, hybrid semantic and symbolic verification, and graceful handling of inapplicable evaluation criteria. The implementation reached 93% test coverage across 88 tests, with cross-version CI validation from Python 3.9 to 3.12. The work demonstrates how carefully defined evaluation contracts and offline heuristics can provide reproducible, privacy-preserving model assessment without relying on external LLM-as-a-judge APIs.

Contributors
Mahrukh Baig(Lead Engineer)
Warisha Arshad(Evaluation Engineer)
Muhammad Maaz(Implementation Engineer)
NLPTrustworthy AILLM Engineering
Engineering FellowshipCohort 2026
2026-08-19 6 min read
faker-pk v2.0

Engineering a Relational Data Layer and Consistency-Aware Synthetic Data Generation for faker-pk

Engineering faker-pk into a consistency-aware synthetic data layer where every generated field agrees with the same real-world context.

Contributors
Ayesha Anwar(Data Engineer)
Muhammad Khubaib Ahmad(Supervisor)
Data Engineeringfaker-pklocalized data generation
Engineering FellowshipCohort 2026
2026-05-15 8 min read
auralis-vfs

Engineering a Clinical-Grade Vocal Fatigue Scoring Pipeline from Raw Speech

How we designed, benchmarked, and packaged ECAPA-TDNN-VHE into a pip-installable Python library with sub-second inference latency.

Contributors
Muhammad Khubaib Ahmad(Founder & Director)
Speech AIPyPIECAPA-TDNNPyTorchAudio Processing
ResearchCohort 2026
2026-04-10 7 min read
RUEmoCorp

Building RUEmoCorp: Annotation Pipelines & Privacy-Preserving Dataset Curation for Roman Urdu

Engineering a 134K-sample multi-class emotion dataset and Transformer benchmark suite for low-resource South Asian NLP.

Contributors
Khadija Faisal(Data Annotator & Validator)
Muhammad Khubaib Ahmad(Lead Engineer & Researcher)
NLPRoman UrduDataset CurationHuggingFaceTransformers
Engineering FellowshipCohort 2026
2026-09-19 6 min read
docling-pk

Building Robust OCR Extraction for Pakistani Documents

Built the core image processing, OCR, and field extraction pipeline for docling-pk, supporting Pakistani identity and education documents. Real-world testing exposed failures caused by image preprocessing assumptions and OCR layout variability, leading to more robust extraction strategies. The system now handles rotated documents, layout-dependent fields, and explicit extraction failures rather than relying on ideal document structure. The work reinforced a practical lesson: real document variability must drive OCR design, not assumptions from standard preprocessing techniques.

Contributors
Muhammad Abdul Moiz(Lead Engineer)
Computer VisionNLP
Engineering FellowshipCohort 2026
2026-08-19 6 min read
voiceMonitor v1.1.0

Building a Vocal Load Impulse Response Model for voiceMonitor

Engineering voiceMonitor to track acute and chronic vocal strain, estimate recovery, and adapt to each speaker's baseline.

Contributors
Sarib Azim(Lead Engineer)
Muhammad Khubaib Ahmad(Supervisor)
Speech AIVocal HealthAudio ProcessingVocal Fatigue Detection