Data Engineer & ML/AI Practitioner

Data & AI systems built the way research taught me — rigorous, reproducible, shipped.

I'm Onyekachi (Kachi) Emenike — a data engineer and ML practitioner with a PhD in Computational Mathematics. I design ELT pipelines and applied ML/AI systems for teams that need answers they can trust, not just dashboards.

PhD, Computational Mathematics 4 peer-reviewed publications Frankfurt, Germany EU work authorisation
Selected work

Case studies

Four systems, chosen for what they had to solve — not just the stack behind them.

01

Maternal Health AI Assistant

Grounding an AI assistant in real medical guidance instead of model memory.

Business impact · Customer trust & risk

In a health or consumer product, an assistant like this cuts support costs by deflecting routine questions — but the real value is risk: every answer traces back to a real medical guideline instead of model invention, which is what actually gets an AI health feature past legal and compliance review, not stuck in a pilot that never ships.

Problem

New mothers need fast, trustworthy answers on infant nutrition and breastfeeding — but medical guidance is scattered, and generic chatbots hallucinate on health questions.

Approach

Built an agentic RAG system: a LangChain agent decides whether to answer from a FAISS-indexed WHO/AAP/CDC guideline store, or call the USDA FoodData Central API live for nutrient data.

Outcome

Deployed live on Streamlit. Every answer traces back to a real guideline or live nutrition data — not model guesswork.

LangChain · GPT-4o · FAISS · HuggingFace · Streamlit

02

NourishMama — Cloud-Native ELT & Warehouse

Turning scattered nutrition data into direct, age-specific food guidance.

Business impact · Decision quality & cost

Bad pipelines are a hidden cost center — teams make pricing, product, and marketing decisions on numbers nobody validated, and errors compound silently until they surface as bad decisions. A tested, layered warehouse is what makes a company's dashboards trustworthy enough to act on, and it's infrastructure most early-stage teams don't have anyone dedicated to building — the exact gap this project shows I can close alone, not just query someone else's warehouse.

Problem

First-time mothers lack one data-driven place to find nutrient-rich, age-appropriate foods for themselves and babies under one.

Approach

Built a Bruin-orchestrated ELT pipeline on GCP — Terraform for infra, a layered BigQuery warehouse — ingesting 8,000+ nutrition records into analysis-ready marts.

Outcome

A live Streamlit dashboard non-technical users filter by baby age and audience, translating raw data into direct food recommendations.

Python · GCP · Terraform · Bruin · Streamlit

03

House Price Predictor — Rhineland-Palatinate

A regional pricing model that explains its own predictions.

Business impact · Revenue & model risk

In real estate, lending, or insurance, a pricing model that can't explain itself is a liability, not an asset — regulated industries need to justify decisions, and unexplainable models get blocked in model risk review regardless of accuracy. Pairing strong predictive performance with SHAP explainability is what actually lets a model reach production in those industries, not just win a leaderboard.

Problem

Estimating fair property prices in a specific German region needs a model trained on local listings — national models miss regional pricing dynamics.

Approach

Built a local SQLite pipeline from 3,800+ regional listings, ran full EDA and feature engineering, and trained an XGBoost regressor with MLflow experiment tracking.

Outcome

R² of 0.75 on held-out data. Shipped with SHAP explainability and an LLM chat assistant, so users see why a price was predicted, not just the number.

Python · XGBoost · SQLite · MLflow · SHAP · Streamlit

04

MLOps Pipeline — Demand Forecasting

A forecasting pipeline that watches its own accuracy.

Business impact · Cost & operational risk

Forecasting errors cost real money in both directions — overstock ties up capital, understock loses sales — and the danger isn't the first bad forecast, it's a model that quietly degrades for weeks before anyone notices. Automated drift monitoring turns that silent failure into an early, actionable alert — the operational habit that keeps a forecasting system trustworthy long after launch, not just clean on demo day.

Problem

A forecasting model is only useful in production if it's monitored — silent data drift quietly erodes accuracy over time.

Approach

Engineered a full pipeline from ingestion to deployment: Prefect for orchestration, Docker/Kubernetes for serving, GitHub Actions for CI/CD, Evidently for drift monitoring.

Outcome

A self-monitoring pipeline that flags data and model drift automatically, rather than relying on manual review.

Python · MLflow · Prefect · Docker · Kubernetes · GitHub Actions · Evidently
About

I'm a data engineer and analytics professional with a PhD in Computational Mathematics and hands-on experience designing layered data warehouse architectures, complex SQL transformations, and production-grade ELT pipelines — from raw ingestion through staging to analytics-ready marts.

"I spent years building numerical algorithms that had to be provably correct. That's the standard I hold data pipelines to now."

My background is in computational research: several years building high-performance numerical algorithms and simulation software in international academic collaborations across the US, Austria, and Belgium. I currently lecture in Data Analytics & Machine Learning at the graduate level in Berlin, and write on Medium about applied data science and ML.

Location
Frankfurt, Germany · EU Work Authorisation
Focus areas
Data engineering, ML/AI systems, analytics engineering
Currently
Lecturer — Data Analytics & ML (MSc), Berlin School of Business & Innovation
Available for
Full-time roles, contract & consulting work
Education
PhD Computational Mathematics, Johannes Kepler University
Languages
English (Fluent), German (A2, actively improving)
Capabilities

Tools & stack

Data modelling
SQL, layered warehouse design (raw/staging/marts), dbt, data quality & testing
Programming
Python, SQL, R
Cloud & platforms
GCP (BigQuery, GCS), DuckDB, cloud data warehouses
Orchestration & infra
Airflow, Prefect, Bruin, Docker, Kubernetes, Terraform, GitHub Actions (CI/CD)
Visualisation
Streamlit, Plotly, Power BI
MLOps
MLflow, Evidently
Experience

Where I've worked

2024 — 2026

Lecturer — Data Analytics & Machine Learning (MSc)

Berlin School of Business & Innovation · Berlin, Germany

Designed and delivered graduate curriculum in applied statistics, ML, and data pipeline engineering; mentored students on reproducible workflow design and communicating results to non-technical stakeholders.

2022 — 2023

Senior Research Associate / Visiting Assistant Professor

Illinois Institute of Technology (IIT) · Chicago, USA

Developed modular, reusable Python components for numerical simulation workflows across an international collaboration; taught undergraduate Calculus to 40+ students.

2018 — 2022

Research Scientist

Johann Radon Institute (RICAM) · Linz, Austria

Built high-performance numerical algorithms in Python and MATLAB with collaborators across Belgium, Austria, and the US, resulting in 4 peer-reviewed journal publications.

Credentials

Education & certifications

Education

MSc Financial Engineering (Part-time), WorldQuant University
2024 — Ongoing
PhD Computational Mathematics, Johannes Kepler University, Austria
2019 — 2022
MSc Mathematics, African University of Science & Technology, Nigeria
2016 — 2017
BSc Industrial Mathematics & Computer Science, Ebonyi State University, Nigeria
2010 — 2014

Certifications & awards

Data Engineering Zoomcamp — DataTalksClub
2026
First Place, Data Science & ML Hackathon — MIT Institute for IDSS
2024
MLOps Zoomcamp — DataTalksClub
2024
Special Prize Zonta Award — Johannes Kepler University, Linz
2021
Research

Doctoral research & publications

Before data engineering, five years of peer-reviewed research in numerical analysis — the same rigor now applied to production systems.

2022

Quasi-Monte Carlo Methods: Component-by-Component Algorithms

PhD Thesis · Johannes Kepler University, Linz, Austria

Developed and analyzed adaptive component-by-component construction algorithms for lattice and polynomial lattice rules used in high-dimensional numerical integration — the same class of problem underlying Monte Carlo option pricing and risk simulation.

Read: why this matters for pricing & risk desks →

Full publication list on Google Scholar ↗  ·  ORCID ↗

Contact

Let's talk.

kachi.emenike12@gmail.com

Open to full-time roles, contract work, and collaborations in data engineering and AI/ML. I usually reply within a day or two.