Skip to content
View donthula9908's full-sized avatar

Block or report donthula9908

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
donthula9908/README.md

Naveen Donthula

Senior Data Engineer | Cloud data platforms, lakehouse architecture, and reliable pipelines

I design and operate data platforms across Azure, Databricks, Microsoft Fabric, AWS, Snowflake, and dbt. My work spans ingestion, transformation, data quality, observability, governance, CI/CD, and analytics delivery.

LinkedIn Email

Selected impact

  • Improved financial-pipeline throughput by approximately 9x through SQL execution-plan and ordering optimization.
  • Built and maintained 80+ Azure Data Factory pipelines across development, staging, and production environments.
  • Designed PySpark row-level reconciliation using exceptAll to validate Databricks-to-Synapse migrations.
  • Automated incident-management workflows and documented SOX controls for daily financial reconciliation.

Featured engineering work

Project What it demonstrates Core technologies
Azure end-to-end data engineering Medallion architecture from ingestion through analytics serving ADF, ADLS Gen2, Databricks, Synapse
Databricks Lakehouse Governed batch and streaming patterns on Delta Lake PySpark, DLT, Structured Streaming, Unity Catalog
Snowflake and dbt analytics Modular ELT models, testing patterns, and warehouse automation Snowflake, dbt Core, SQL
Apache Airflow data pipelines Production-oriented orchestration and incremental ingestion Airflow, dbt, Databricks, Docker
AI pipeline anomaly detection Detecting volume, schema, and pipeline-behavior anomalies Databricks, PySpark, Isolation Forest
Power BI analytics dashboards Semantic modeling and governed enterprise reporting Power BI, DAX, Fabric, DirectLake

Engineering focus

  • Platform architecture: lakehouse, medallion, batch, streaming, dimensional modeling
  • Data engineering: Python, SQL, PySpark, Delta Lake, dbt, Airflow, Kafka
  • Cloud platforms: Azure, Databricks, Microsoft Fabric, AWS, Snowflake
  • Reliability and governance: data quality, reconciliation, observability, lineage, Unity Catalog, SOX controls
  • Delivery: Azure DevOps, GitHub Actions, ARM templates, Databricks Asset Bundles, Docker

Certifications

  • Microsoft Azure Data Engineer Associate (DP-203)
  • Microsoft Fabric Analytics Engineer Associate (DP-600)
  • Databricks Certified Data Engineer Associate
  • Databricks Certified Associate Developer for Apache Spark
  • AWS Certified Data Engineer – Associate (DEA-C01)
  • SnowPro Core Certification

How I approach data platforms

Sources -> Reliable ingestion -> Governed storage -> Tested transformations
        -> Observable pipelines -> Trusted data products -> Business outcomes

I care about systems that are maintainable after launch: explicit contracts, repeatable deployments, actionable monitoring, clear ownership, and documentation that helps the next engineer succeed.

Pinned Loading

  1. apache-airflow-data-pipelines apache-airflow-data-pipelines Public

    Production Apache Airflow DAGs — incremental ingestion, dbt orchestration, Databricks job triggers, custom operators

    Python

  2. azure-ai-data-engineering azure-ai-data-engineering Public

    AI patterns for Data Engineers — Azure OpenAI RAG over Unity Catalog, Text-to-SQL agent, LangChain + Azure AI Foundry

    Python

  3. azure-data-engineering-projects azure-data-engineering-projects Public

    End-to-end Azure Data Engineering — ADF, ADLS Gen2, Databricks, Synapse Analytics, Medallion Architecture

    Python

  4. data-engineer-roadmap data-engineer-roadmap Public

    Practical Data Engineer Roadmap 2025 — Azure, Databricks, Fabric, AWS, Snowflake, dbt, Airflow, AI — built from 6+ years experience

  5. databricks-lakehouse-projects databricks-lakehouse-projects Public

    Enterprise Databricks Lakehouse — PySpark, Delta Lake, Delta Live Tables, Unity Catalog, Structured Streaming

    Python

  6. netflix-data-engineering-pipeline netflix-data-engineering-pipeline Public

    End-to-end Netflix-style streaming analytics pipeline on Azure + Databricks — viewer sessions, content KPIs, churn signals

    Python