AI/ML Engineer · 9 years · Clinical & Health-Tech Systems

I build the AI infrastructure health-tech teams can actually run in production.

Real-time clinical speech transcription, medical RAG chatbots, structured EMR extraction, hybrid pharmaceutical search — end-to-end, from first architecture decision to the audit log that proves it's working.

What I take on

Scoped engagements, not open-ended retainers. Each one ends with a system you (or your team) can run without me in the room.

  • Clinical & Healthcare AI Systems

    Speech-to-text for clinical documentation, medical RAG chatbots, patient-facing tools built with safety chains and audit logging from day one — not bolted on after a compliance review.

  • RAG & LLM Application Engineering

    Retrieval pipelines, structured extraction from unstructured text, multi-pass LLM architectures with proper evaluation — designed to be measured, not just demoed.

  • Production ML Infrastructure

    GPU-backed inference services, multi-model serving via vLLM, observability and structured logging, containerization — the parts that turn a working notebook into a system your team can trust.

  • Data & NLP Pipelines

    Hybrid search (FAISS + BM25), domain-specific entity extraction, enrichment pipelines that carry real business context through the full request lifecycle.

  • AI System Audits & Remediation

    Inherited an AI system that works in the demo but not in production? I run a full production-readiness review and hand you a scoped remediation plan — then implement it.

Selected work

Systems built for clinical and healthcare use. Details generalized where required for confidentiality.

  • Production Hardening

    Medical RAG Patient Safety Chatbot

    Patient-facing safety chatbot built on a five-stage pipeline — preprocessing, intent classification, safety policy, FAISS + cross-encoder retrieval, and structured audit logging — designed so every answer is traceable back to its source and its safety check.

    FastAPIFAISSCross-EncoderGPT-4.1
  • Active Development

    DrugSearch — Pharmaceutical Hybrid Search

    Hybrid FAISS + BM25 search over pharmaceutical data, enriching prescriptions with hospital-aware context so multi-site hospital networks get results scoped to the correct facility and hospital type.

    FastAPIFAISSBM25MySQLPydantic
  • In Validation

    Clinical Note / EMR Extraction Pipeline

    A four-pass, concurrent LLM pipeline that turns raw doctor-patient consultation transcripts into ~18 categories of structured, EMR-ready data — complaints, history, exam findings, medications, follow-up — validated across two open-weight model families on GPU-served inference.

    vLLMLlama 3.1MedGemmaStructured Output
  • Internal Tool

    vLLM Multi-Model Console

    A demo console for evaluating multiple production LLM deployments side-by-side — streaming responses, structured JSON output, and tool calling, all behind a single reverse-proxied endpoint.

    vLLMFastAPIJavaScript

Background

Nine years of applied AI/ML — most recently independent; before that, inside government and enterprise R&D. Employer names are withheld below; the work is described at a general level.

  • 7 years

    Government & Enterprise R&D — NLP, Computer Vision & Speech

    Multilingual NLP and computer vision systems for public-sector and enterprise use: OCR pipelines reaching 94% character-level accuracy, ASR/TTS systems for Indic languages, sentence-similarity models fine-tuned via Siamese networks, and multimodal chatbots across three languages.

  • Predictive ML

    Predictive Health-Risk Alerting

    Built real-time and batch inference pipelines to flag health-risk deviations from multimodal sensor and behavioral data, using ensemble models (XGBoost, Random Forest) trained partly on synthetic (CTGAN) data to address class imbalance, deployed via cloud-based training infrastructure with automated batch scheduling.

  • Clinical data

    Clinical Predictive Modeling

    Worked directly with structured clinical datasets to build early-detection models for a pregnancy-related condition, under data governance and auditability requirements for clinical data.

  • 97% accuracy

    Health Screening System

    Designed and evaluated multiple ML models for an early-detection screening application; the top-performing model reached 97% accuracy.

How I work

  1. 01

    Discovery & audit

    I review what exists first — requirements, or your current system. This is usually where production blockers surface, before anything new gets designed.

  2. 02

    Scoping

    A bounded plan: what's built first, what's explicitly deferred, and how we'll know each part is actually working.

  3. 03

    Build

    Production-grade code from day one — structured logging, error handling, and tests included, not a prototype that needs a rewrite to ship.

  4. 04

    Handoff & monitoring

    Audit logging, observability, and documentation, so the system doesn't depend on me staying in the room.

Have a system that needs to reach production?

Tell me what you're building and where it's stuck. I'll tell you honestly whether I'm the right fit.

rishav@rmlk.dev