Sujal Varshney
Data Engineer
Remote (global)
LinkedIn profile available to registered employers
Self-assessed English levels
Metadata-Driven Lakehouse & Data Quality Framework
● Architected Lakeflow Declarative Pipelines (DLT) on Databricks implementing end-to-end Bronze-Silver-Gold Medallion Architecture for unified batch and streaming workloads. ● Built a metadata-driven ingestion framework supporting full-load and incremental/CDC processing across heterogeneous sources (CSV, JSON, SQL, REST APIs), reducing pipeline development time by 35% through reusable, config-based logic. ● Implemented a custom data quality and validation framework, isolating invalid records into quarantine tables and enabling automated logging and alerting for failed validation checks, flagging ~8-10% of incoming records for quality violations. ● Delivered production-ready Gold-layer datasets and Databricks Streaming Tables for downstream analytics consumption, eliminating dependency on ADF orchestration in favor of Databricks-native scheduling — reducing orchestration maintenance overhead by ~15% and cutting scheduling-related incidents by ~30%.
