About
Data Scientist with 4+ years building production machine learning (ML), forecasting, NLP, and GenAI/RAG systems across financial services and healthcare. Combines statistical rigor, feature engineering, and MLOps automation to move models from experimentation into monitored production. Translates ambiguous business problems into measurable model and decision-system improvements for executive stakeholders.
Skills & Tools
Programming Languages
ML Frameworks & Libraries
Big Data & Cloud
MLOps & Model Governance
Databases & Visualization
Statistical & ML Methods
Experience
- Own end-to-end data science delivery for portfolio operations, forecasting, and risk analytics, converting ambiguous stakeholder goals into features, experiments, validation plans, and production model outputs used across teams.
- Build forecasting, risk-scoring, and anomaly detection models in Python with Scikit-learn, XGBoost, and LightGBM, improving forecast accuracy by 11% and alert precision by 13%, validated with time-based backtesting.
- Engineer feature pipelines on Spark, AWS Glue, Airflow, and Redshift processing 800GB+ of structured and semi-structured data monthly, cutting manual preparation and reconciliation effort by 22% per team time tracking.
- Design reusable feature layers spanning account, transaction, market, document, and operational signals, standardizing training data across models and removing repeated one-off feature builds from each new model cycle.
- Productionize model training, batch scoring, and inference APIs with MLflow, SageMaker, FastAPI, Docker, Kubernetes, and GitHub Actions, reducing model refresh cycles from 4 days to under 2 days, tracked in deployment logs.
- Develop RAG retrieval workflows with LangChain, LangGraph, OpenAI GPT, and FAISS, using metadata filtering and reranking to cut analyst document discovery from 15 minutes to under 6 minutes, measured via request-level usage logs.
- Implement backtesting and evaluation frameworks comparing model versions, decision rules, and prompts on lift, stability latency, cost, and error-slice metrics, giving stakeholders a consistent basis for promotion and rollback.
- Apply SHAP, LIME, feature importance, and drift detection to explain model behavior and support governance documentation, reducing validation turnaround by roughly 30% based on review cycle times tracked with compliance partners.
- Modeled healthcare utilization, cost forecasting, and member risk stratification in Python and SQL with Scikit-learn, XGBoost, and LightGBM, giving care management teams ranked intervention lists built from claims data.
- Engineered claims, enrollment, provider, and clinical-text features with PySpark, Hadoop, and ETL pipelines, improving high-risk member identification by 16%, validated with time-based holdout splits on historical outcomes.
- Led A/B test design with power analysis, confidence intervals, and segment-lift measurement for outreach programs, giving product and clinical stakeholders a defensible framework for rollout and expansion decisions.
- Built NLP pipelines using BERT, spaCy, TF-IDF, and named entity recognition (NER) to classify member feedback and clinical note themes, cutting manual review effort by roughly 25% based on reviewer workload tracking.
- Automated model retraining, validation, and registry promotion with Airflow, MLflow, and GitLab CI/CD, shortening recurring model refresh cycles from 5 days to 3 days as recorded in scheduled pipeline execution histories.
- Created anomaly detection and clustering workflows surfacing unusual utilization patterns, care gaps, and member cohorts, then presented calibration curves and error analysis to align model thresholds with business and compliance risk.
Education
Certificates
Let's Connect
Got an idea, a role, or a problem worth solving? Drop me a message.