Hey, I'm

Martin Ngoh

|

I build data and AI products, from a police-accountability publication to an AI lease screener.

  • ABOUT

    Senior data scientist building and shipping machine learning and LLM systems in production, for federal clients and my own products. Looking for a senior data scientist role on ML and LLM systems.

    5+ Years Experience
    3 Live Apps
    $1M+ Saved Per Year
    4M+ Records Analyzed
    2024 — Present

    Senior Data Scientist

    Accenture Federal Services

    • LLM document pipeline scaled from 100 to 10,000+ documents a day
    • XGBoost service that turned 4-hour validation sessions into 30-second calls
    2021 — 2024

    Data Scientist

    Deloitte Consulting

    • Fraud detection improved 8% by tuning a model ensemble
    • Model-training turnaround cut from 3 days to 2 hours
    • Analysis that informed a 3-year, $30M product initiative
    2022

    M.S. Business Analytics

    Georgetown University

    2020

    B.S. Supply Chain Management & Analytics

    Virginia Commonwealth University

    Skills

    Languages
    Python SQL PySpark R SAS
    ML & AI
    Claude API LLM pipelines RAG XGBoost scikit-learn PyTorch Prophet Bayesian inference
    MLOps & Data
    AWS SageMaker Docker MLflow FastAPI PostgreSQL Railway

    PROJECTS

    Products and research

    Residents Count home page listing its reports

    Residents Count

    A public-safety research publication turning police records into rates per resident, city by city. Latest: DC violence fell after the Guard arrived, but no more where troops stood.

    How I built it 2.5M FBI victim records for 59 cities against Census denominators; pre-registered tests, placebo posts and a synthetic control from 590 agencies.

    Python pandas pyfixest Chart.js
    Causal Inference Synthetic Control Pre-registration

    live

    Justice Lens home page

    Justice Lens

    Built and ran a live platform testing whether DC policing is applied equitably: nine statistical analyses on 1.4M+ public police records. Its findings now run as Residents Count's Policing in DC series.

    How I built it Daily ingest of DC Open Data into PostGIS, veil-of-darkness and Bayesian threshold tests fitted with MCMC, ~500 automated tests.

    Python FastAPI PostGIS
    Bayesian MCMC Hypothesis Testing Geospatial
    Charts from a findings page built with disparity-kit

    disparity-kit

    The toolkit behind the assault analyses: from a raw police file to a published, caveated findings page. Seven cities and a 59-city run built with it; earlier results reproduce exactly.

    How I built it Ten Claude Code skills over a Python kit: data audit, Census denominators by tract, rival-explanation tests, a Poisson adjustment ladder, replication in a second source, a racial-bias check on the data and the writeup, and a page builder.

    Python pandas statsmodels Claude Code
    Agent Skills Poisson GLM Reproducibility
    Lease Screener home page

    Lease Screener

    Reads a lease, flags clauses against the DC and Virginia tenant-law texts, and answers questions about it. Built and shipped as a product; kept live as a working demo.

    How I built it Lease chunked and indexed with BM25, reranked with a cross-encoder, answered by Claude Haiku with a Sonnet pass checking each report against the tenant-law texts.

    Python FastAPI Claude API Railway
    RAG BM25 + Reranking LLM-as-Judge

    live

    More on GitHub →