Search Tech Journey

Find topics, journeys and posts

· about

Sai Dinesh Reddy
Maluchuru.

Data & ML Engineer at Microsoft — OneNote Copilot and OneDrive SharePoint. Previously Amazon and JPMorgan Chase. I architect large-scale data platforms, ML pipelines, and cost-optimization systems. I write this blog to make what I learn easier to find, and easier to reason about.

Hyderabad, IndiaIIT Madras · Dual Degree · Electrical EngIIT-JEE 2012 · AIR 728
work
Microsoft
current
Jul 2022 — Present
Hyderabad

Data & ML Engineer 2

  • ODSP · scale

    200 TB+ daily raw logs and 300 B+ single-day Blob transactions; pipelines that cut compute and storage by 70%.

  • OneNote Copilot · eval

    Designed the LLM-as-a-Judge framework for video content — rubrics, scoring pipeline, prod rollout with ML + Applied Science teams.

  • OneNote Copilot · DE team

    Stood up the Data Engineering division from scratch — 200+ pipelines and ML workflows with monitoring + alerting.

  • Blob cost analytics

    Ran ODSP's biggest cost line — Highcharts dashboards surfaced $M-level savings and drove exec budget decisions.

  • Windows analytics

    Business-critical system used by 28+ Windows teams; drove Passkey adoption +30%.

  • Store personalization

    Cut ML training costs 80% via Resilience LPG framework optimizations; experimentation days → hours.

  • LLM advocacy

    Trained 50+ engineers on EDA with Claude and GPT-4; championed LLM adoption across ODSP.

Amazon
Jul 2019 — Jul 2022
Hyderabad

Data Engineer

  • Package lifecycle pipeline

    Three foundational datasets — Node Processing, Package Attributes, Package Items — that became the source of truth for logistics analytics worldwide.

  • Delivery Estimate Accuracy

    Root-cause attribution across millions of daily packages in India — improved promise accuracy, reduced customer-facing misses.

  • Leave-at-Door pilot

    Instrumented + analysed the India pilot; dashboards drove the nationwide rollout decision.

  • Fintech + tax compliance

    Compliance and tax pipelines across regions with 99.9% accuracy and full auditability.

  • Platform + SLA

    80+ downstream teams on Hoot / Redshift / Airflow / ECS; SLA adherence +25% on mission-critical AMZL datasets.

JPMorgan Chase
Jun 2017 — Jul 2019
Hyderabad

Associate

  • Oracle → Spark migration

    Led end-to-end migration of the batch data platform — 3× throughput, resolved data skew, cut infra cost.

  • Cost-Based Allocation Model

    Cross-Loans/Cards/Chase-Merchant P&L attribution model for leadership.

  • Tableau BI

    Interactive dashboards for Mortgage Banking, Cards, Merchant Services — 300+ daily users, -40% manual reporting.

  • Alteryx + ML enablement

    Automated data-wrangling POC (-60% data-prep time); led team-wide ML fundamentals sessions.

General Electric
May 2015 — Jul 2015
Bangalore

Summer Intern

  • Vehicle fault detection

    ML for engine fault detection using multi-sensor parameter streams; text mining to establish reliable ground truth labels.

skills
Languages
PythonSQLKQLSpark SQLHiveQLBashR
Cloud + big data
Azure Data FactorySynapseCosmos DBKustoAWS ECS / FargateS3RedshiftSparkHadoopHiveSqoop
AI + ML
ClaudeGPT-4LLM-as-a-JudgePrompt engineeringSupervised MLUnsupervised MLMLOps
Data + viz
Power BITableauHighchartsAlteryxAirflowDJS
Core strengths
Petabyte-scale pipelinesCost optimizationOracle → Spark migrationsEDA automationCapacity planning
education
2012 — 2017
Chennai

Indian Institute of Technology, Madras

Dual Degree (B.Tech + M.Tech) · Electrical Engineering
  • · IIT-JEE 2012: All India Rank 728 out of 560,000+ candidates (top 0.13%).
  • · M.Tech thesis: vehicle classification via image processing and ML — 80% accuracy.
  • · Consistently rated Exceeds Expectations at Microsoft; recognized for outstanding engineering contributions.
writing
this blog

A six-month learning plan broken into 12 modules and 130 self-contained sessions — the syllabus I wish someone had handed me on day one. Independent essays land in the same feed as they come. Sessions are short (~15 minutes), one idea per session, prev/next chained so you can walk the plan straight through or drop in anywhere.

  • AI + LLMsM09 · 10 sessions
  • MLOps + productionM10 · 10 sessions
  • Data engineeringM05 · 10 sessions
  • System designM11 · 10 sessions
  • Python + DSAM01 · M03 · 21 sessions
  • Databases + SQLM04 · 10 sessions

Full DAG at /learning-path. Recent posts at /.

house rules
  • · Zero background assumed. If a post needs a concept, it's introduced first.
  • · Diagram first. If it can't be drawn, it isn't yet understood.
  • · Real production numbers. Never fabricated. Missing numbers stay missing.
  • · Then the code. Small, runnable, commented.
  • · Sessions, not essays. One idea, one coffee.
contact

Fastest reply on email. Open to interesting conversations on data + AI infra, LLM eval, cost optimization, and mentoring.


· one session at a time