I build end-to-end machine learning systems — from data pipelines to deployed models — for problems where accuracy has real consequences, like early medical diagnosis and risk prediction. Every project below ships with the metrics that back it up.
Diagnostic classification of tumors from clinical measurements, where missed positives carry the highest cost.
Engineered the 10 clinically-critical mean features from the diagnostic dataset, tuned the classifier for F1/recall rather than raw accuracy alone, and cached inference so the Streamlit dashboard responds in under a second.
Predict diabetes risk from health indicators despite a dataset skewed heavily toward negative cases.
Built a custom imputation strategy for missing clinical values, then tuned scale_pos_weight in an XGBoost ensemble to correct for class imbalance without sacrificing diagnostic sensitivity.
Classify tumor data into diagnostic categories where feature overlap between classes makes separation difficult.
Combined statistical exploratory analysis with engineered features to sharpen the decision boundary, then benchmarked ensemble and neural approaches against a held-out validation split before selecting the most stable model.
Predict player draft status — drafted vs. undrafted — from unstructured athletic scouting report text.
Extracted textual indicators from scouting reports and paired them with engineered tabular metrics, then deployed ensemble models (XGBoost / Random Forest) evaluated under cross-validation for a fair accuracy estimate.
Identify which telecom customers are likely to churn from noisy, high-cardinality behavioral and billing data.
Built the feature pipeline directly from relational customer data, trained a LightGBM classifier, and validated it with stratified K-fold cross-validation to keep the churn/non-churn ratio honest across folds.
From problem framing and data collection through training, validation, and honest reporting of results — the same process behind every project above.
Scalable pipelines that turn messy, raw, or relational data into clean, analysis-ready datasets — with validation routines built in, not bolted on.
Getting a trained model out of a notebook and into something usable — serialized, served behind an API or dashboard, and monitored after launch.
Exploratory analysis and statistical visualization that surfaces what's actually in the data, before a single model gets trained on it.
If you're working on a problem where prediction quality actually matters — I'd like to hear about it. Based in Mansoura, Egypt, working with teams anywhere.