Available for work · Toronto, ON

I engineer data into intelligent systems

// Data · AI · Machine Learning · Cloud

I'm a Data & ML engineer with 3+ years turning raw, messy data into scalable pipelines, production ready models, and cloud native architecture, bridging the gap between engineering and data science.

0
Years of experience
0
Shipped projects
0
Certifications
0
AWS · Azure · GCP

01 Toolbox

Technologies I work with

A complete data toolkit, from ingestion and orchestration through modeling, GenAI, and deployment.

Languages & Query

  • Python
  • SQL
  • Scala
  • TypeScript
  • Java
  • C++

Data Engineering

  • Spark
  • PySpark
  • Kafka
  • Airflow
  • Hadoop
  • Hive
  • Flink

Cloud & Warehousing

  • AWS
  • Azure
  • GCP
  • Databricks
  • Snowflake
  • Redshift
  • BigQuery

ML & Deep Learning

  • PyTorch
  • TensorFlow
  • Keras
  • Scikit-Learn
  • XGBoost
  • CNN · RNN · LSTM

GenAI & LLMs

  • LLMs
  • RAG Pipelines
  • Agentic AI
  • LangGraph
  • Transformers
  • Prompt Engineering

MLOps & DevOps

  • Docker
  • Kubernetes
  • Terraform
  • MLflow
  • GitHub Actions
  • CI/CD

Analytics & BI

  • Power BI
  • Tableau
  • Looker
  • Grafana
  • Pandas
  • NumPy

Automation & Platform

  • Power Automate
  • Copilot Studio
  • Jira
  • SharePoint

02 Experience

Where I've made an impact

  1. Associate · Data & AI

    Sep 2025 — Present

    PwC Canada

    Modernizing a Tier 1 Canadian bank's management reporting platform: migrating legacy Informatica ETL to a cloud ready AWS architecture on Redshift and Spark, and shipping governed GenAI solutions on Azure.

    • Designed a 3 layer data architecture with end to end Airflow DAGs, PySpark transformations, PII masking, and data quality checks
    • Cut manual migration effort by ~98% with an LLM powered accelerator generating SQL, DDLs, and source to target mappings
    • Built a GenAI reporting agent on Azure AI Foundry, saving 17+ managers 10+ hrs/week
  2. Machine Learning Engineer

    Jan 2022 — Dec 2022

    BrainyBeam Technologies

    Owned end to end ML pipelines, from data cleaning and feature engineering to experiment tracking and containerized deployment, for production models.

    • Shipped a telecom churn model (XGBoost, MLflow, Dockerized Flask) reaching 0.87 AUC ROC with GitHub Actions CI/CD
    • Cut deployment latency by 40% and improved training efficiency by 25%
    • Engineered 18+ features, reducing overfitting by 30%
  3. Data Analyst

    Jun 2021 — Nov 2021

    Ecubix

    Directed SQL based ETL workflows and analytics on high volume datasets, applying advanced statistical methods to uncover trends, detect anomalies, and support evidence based planning.

    • Drove a 20% increase in customer acquisition through actionable insights
    • Enabled a 15% revenue boost via timely, data driven interventions
  4. Python Developer Intern

    Jan 2020 — Dec 2020

    Bascom Bridge Education

    Developed automation scripts for high volume data processing, optimizing Python solutions for complex workflows.

    • Reduced operational overhead by 32% through automation
    • Accelerated project delivery by 10%

03 Selected work

Projects

End to end data engineering pipelines and applied machine learning, with the metrics that mattered.

Data Engineering+50% efficiency

Real Time Stock Market Streaming

Real time financial streaming pipeline on Kafka and AWS (EC2, S3, Glue, Athena), boosting analysis efficiency by 50% on high throughput market data.

  • Kafka
  • AWS
  • Glue
  • Athena
Data Engineering+50% efficiency

Football Analytics Pipeline

End to end analytics pipeline on Apache Airflow and Azure Data Lake, automating extraction, cleaning, and transformation with Databricks and Power BI.

  • Airflow
  • Azure Data Lake
  • Databricks
  • Power BI
Data Engineering+60% efficiency

Uber Trip Analytics

Streamlined ETL and analytics on GCP with BigQuery, Mage, and Looker Studio, delivering insight into trip patterns and fare distribution.

  • GCP
  • Mage
  • BigQuery
  • Looker
Data Engineering+40% throughput

Reddit Data Engineering Pipeline

Scalable ETL on AWS (S3, Glue, Athena, Redshift) orchestrated with Airflow, boosting data processing efficiency by 40%.

  • Airflow
  • AWS
  • Redshift
  • PostgreSQL
GenAIAgentic RAG

GenAI Agentic Customer Support

Multi agent RAG system on GCP using LangGraph with supervisor orchestration and full observability across the agent workflow.

  • LangGraph
  • Vertex AI
  • GCP
  • Docker
Machine Learning95% accuracy

Cell Segmentation with YOLOv8

95% accurate cell segmentation system on YOLOv8 and Azure, with a Docker and GitHub Actions CI/CD pipeline that cut deployment time by 40% for healthcare diagnostics.

  • YOLOv8
  • Azure
  • Docker
  • GitHub Actions
Machine Learning+60% diagnostics

Kidney Disease Classification

High accuracy kidney disease classifier in TensorFlow with a reproducible MLflow and DVC pipeline, improving diagnostic capability by 60%.

  • TensorFlow
  • MLflow
  • DVC
  • CNN
Deep LearningCNN + LSTM

Image Caption Generator

Image captioning model combining CNN and LSTM to produce coherent natural language descriptions on the Flickr_8K dataset.

  • CNN
  • LSTM
  • Keras
  • NLP

Let's build something

Have a data problem worth solving?

I'm open to full time roles and freelance projects across data, AI, and cloud. Drop me a line, I reply fast.