Aurangzaib Ramzan

Senior Data Scientist

Python-focused engineer and data scientist who has built and scaled production systems, shipped CI/CD-backed data products, and improved reliability at high volume. With 11 years of experience in technology and software-oriented product/data platform environments, I excel in Python development, backend and cloud-scaled system delivery, and collaborative problem-solving. Proficient in deploying data pipelines and implementing solutions on cloud infrastructures, I am committed to enhancing operational efficiency and data accessibility through analytical skills and teamwork.

Experience

Jan 2019 — Present

Senior Data Scientist

ComScore

  • Managed the end-to-end delivery of a core data product built on Python and Snowflake, implementing over
  • 50 logic enhancements while utilizing CI/CD practices to ensure robustness and reliability across 10+
  • internal teams.
  • Developed a Difference-in-Differences analysis framework packaged into an automated production pipeline,
  • enabling seamless integration with existing data workflows across 20 countries. This allowed for data-driven
  • decisions that significantly enhanced the reliability of findings.
  • Built an XGBoost churn model to identify at-risk panelists, operationalizing it within the production
  • environment and reducing month-over-month variance by 5%, leading to 20% better resource allocation
  • through proactive retention strategies.
  • Collaborated with Product and Engineering teams to maintain an Isolation Forest anomaly detection model
  • using Python, streamlining deployment and version control through Git, which reduced fraud incidents by
  • 5% and yielded operational cost savings.

Apr 2021 — Jul 2022

Data Scientist (Contract)

Cloudeagle.AI

  • Spearheaded an end-to-end product search platform for an early-stage startup using Python and AWS,
  • scaling the system to handle over 100,000 requests per minute while maintaining 99.9% uptime, significantly
  • improving user experience and system reliability.
  • Enhanced search relevance and discovery by deploying and integrating TensorFlow models within a
  • production API, version-controlled via Git. This optimization, using vector indexes and quantization, resulted
  • in a 15% reduction in latency, which improved user engagement across the platform.
  • Led a cross-functional team to deploy real-time data pipelines utilizing Kafka and Spark Streaming, ensuring
  • high availability and sub-second latency for product search queries at scale leading to enhanced data
  • handling and user satisfaction.

Apr 2015 — May 2019

Statistical Analyst

ComScore

  • Developed a Naïve Bayes classifier in Python using scikit-learn to infer demographics from web behavior
  • signals. This deployment improved targeting accuracy in marketing campaigns by enhancing sample sizes
  • for weighting.
  • Applied regression-based strata analysis with PySpark to identify high-signal features, reducing model
  • runtime by 2 hours through vectorized code optimization that enhanced operational efficiency in backend
  • processes across the organization.
  • Established standardized primary and guardrail metrics through Grafana dashboards, automating alert
  • systems for proactive data quality monitoring. This initiative improved trust in reported metrics and
  • minimized recurring QA efforts by 20%.
  • Created scalable PySpark & SQL pipelines that unify billions of web-log records daily, implementing
  • production deployment and Git versioning. This boosted data consistency, enabling same-day reporting and
  • significantly enhancing data delivery speed for analytics.

Education

Jan 2020

WorldQuant University

Master of Science in Financial Engineering in Financial Engineering · Financial Engineering

Jan 2010 — Dec 2014

Whittier College

Bachelor of Arts in Mathematics in Mathematics · Mathematics

Jan 2006 — Dec 2010

Emory University

Bachelor of Science in Mathematics & Economics in Mathematics & Economics · Mathematics & Economics

Skills

  • Python
  • JavaScript
  • Java
  • C++
  • Go
  • AWS
  • GCP
  • SQL
  • NoSQL
  • Microservices
  • CI/CD
  • Docker
  • Kubernetes
  • React
  • Node.js