Machine Learning & Data Engineering

Models that turn raw data into reliable decisions.

I build end-to-end machine learning systems — from data pipelines to deployed models — for problems where accuracy has real consequences, like early medical diagnosis and risk prediction. Every project below ships with the metrics that back it up.

90%
best model accuracy
Mansoura, Egypt

Five models, five evaluation reports

Every number below is the tested accuracy of a shipped model — not a claim.
01 — Impact
90%
Breast cancer classification accuracy
Scikit-learn · Streamlit
88%
Diabetes risk prediction accuracy
XGBoost · imputation
86%
Tumor detection & classification accuracy
PyTorch · feature engineering
85%
NLP draft-status classification accuracy
NLP · XGBoost
71%
Telecom churn prediction accuracy
LightGBM · Stratified K-Fold
Also: 95% on AWS SageMaker's professional ML-engineer assessment.

Featured projects

Each one follows the same discipline: define the problem precisely, choose an architecture that fits the data, then validate honestly.
02 — Work
Breast Cancer Diagnostic Tool
PythonScikit-learnStreamlit
Problem

Diagnostic classification of tumors from clinical measurements, where missed positives carry the highest cost.

Approach

Engineered the 10 clinically-critical mean features from the diagnostic dataset, tuned the classifier for F1/recall rather than raw accuracy alone, and cached inference so the Streamlit dashboard responds in under a second.

90%
classification accuracy
Optimized for recall — sub-second inference
Diabetes Analytics & Prediction System
PythonXGBoostStreamlit
Problem

Predict diabetes risk from health indicators despite a dataset skewed heavily toward negative cases.

Approach

Built a custom imputation strategy for missing clinical values, then tuned scale_pos_weight in an XGBoost ensemble to correct for class imbalance without sacrificing diagnostic sensitivity.

88%
prediction accuracy
Class-imbalance corrected via scale_pos_weight
Tumor Detection & Classification
PythonPyTorchScikit-learn
Problem

Classify tumor data into diagnostic categories where feature overlap between classes makes separation difficult.

Approach

Combined statistical exploratory analysis with engineered features to sharpen the decision boundary, then benchmarked ensemble and neural approaches against a held-out validation split before selecting the most stable model.

86%
classification accuracy
Validated on a held-out split
Sports NLP Draft Prediction System
PythonNLPXGBoostPandas
Problem

Predict player draft status — drafted vs. undrafted — from unstructured athletic scouting report text.

Approach

Extracted textual indicators from scouting reports and paired them with engineered tabular metrics, then deployed ensemble models (XGBoost / Random Forest) evaluated under cross-validation for a fair accuracy estimate.

85%
classification accuracy
Cross-validated on scouting-report text
Telecom Customer Churn Prediction
PythonLightGBMPandasSQL
Problem

Identify which telecom customers are likely to churn from noisy, high-cardinality behavioral and billing data.

Approach

Built the feature pipeline directly from relational customer data, trained a LightGBM classifier, and validated it with stratified K-fold cross-validation to keep the churn/non-churn ratio honest across folds.

71%
prediction accuracy
Stratified K-Fold on imbalanced churn data

What I can help with

Scoped engagements for teams that need a model built, a pipeline fixed, or a prototype validated.
03 — Services
01

End-to-end ML model development

From problem framing and data collection through training, validation, and honest reporting of results — the same process behind every project above.

Scikit-learnXGBoostPyTorch
02

Data engineering & ETL pipelines

Scalable pipelines that turn messy, raw, or relational data into clean, analysis-ready datasets — with validation routines built in, not bolted on.

SQLPandasETL
03

Model deployment & MLOps

Getting a trained model out of a notebook and into something usable — serialized, served behind an API or dashboard, and monitored after launch.

AWS SageMakerDockerFastAPI
04

EDA & data-driven insight

Exploratory analysis and statistical visualization that surfaces what's actually in the data, before a single model gets trained on it.

EDAStreamlitVisualization

Technical skills

The stack behind every project above, from raw data to a deployed endpoint.
04 — Stack

Machine learning & deep learning

  • XGBoost / LightGBM
  • Ensemble learning
  • Scikit-learn
  • PyTorch
  • Feature engineering
  • Hyperparameter tuning

Data engineering

  • Pandas / NumPy
  • SQL & relational databases
  • ETL pipelines
  • Data preprocessing
  • Data imputation
  • Data warehousing

Cloud & MLOps

  • AWS SageMaker
  • Model serialization (Pickle)
  • Endpoint deployment
  • Performance monitoring
  • Docker

Languages & frameworks

  • Python (expert)
  • C++
  • FastAPI
  • Streamlit
  • Git
  • Jupyter Notebooks

Experience & education

Practical training across data engineering, applied ML, and cloud deployment.
05 — Background
Jul 2026 — Present
Data Engineering Trainee
Digital Egypt Pioneers Initiative (DEPI)
Designing scalable ETL pipelines, optimizing SQL queries and relational schemas, and running automated data-validation routines for enterprise data.
Aug 2026
AI Hackathon Finalist
Creativa Innovation Hubs (ITIDA / TIEC / Orange)
Reached the advanced stage of a national AI hackathon, building and pitching an end-to-end ML/DL application with a cross-functional team under a strict deadline.
Apr 2026 — Aug 2026
AI & Data Science Scholar
University of Tokyo — Matsuo Lab (GCI)
Built end-to-end ML pipelines with Scikit-learn and Pandas, trained ensemble models (XGBoost, LightGBM, Random Forest), and ran statistical EDA across the full model lifecycle.
Feb 2026 — May 2026
Machine Learning & AWS Scholar
Manara
Built, trained, and deployed ML pipelines on AWS SageMaker; scored 95% on the professional technical assessment while applying MLOps practices for serialization and deployment.

Bachelor of Computer Science

Thebes Academy
2023 – 2027 · Cumulative grade: Very Good

Certifications

AWS: Becoming a Machine Learning EngineerManara, 2026
GCI World — AI & Data Science ProgramU. Tokyo, 2026
Deep Learning with PyTorch (Vision & NLP)ITI, 2026
Practical ML for Data ScientistsITI, 2026
ML with Python & Python for Data ScienceIBM, 2026
Data Science Essentials with PythonCisco, 2026
Arabic — native
English — professional

Open to ML engineering roles and collaborations.

If you're working on a problem where prediction quality actually matters — I'd like to hear about it. Based in Mansoura, Egypt, working with teams anywhere.