What is ACID Transactions?
ACID represents four properties that help make database transactions reliable and consistent. The easiest way to understand ACID is through […]
Understanding Data Skew and Salting in PySpark
When working with large datasets in Apache Spark or Azure Databricks, you may encounter a situation where a job looks […]
Databricks / Spark
Term Short explanation Databricks Cloud platform commonly used for large-scale data processing using Spark. Apache Spark Distributed processing engine that […]
Azure & Pipeline Terms
Term Short interview explanation Azure Data Factory (ADF) Cloud service for data ingestion, movement, and pipeline orchestration. Linked Service Connection […]
What is metadata-driven pipeline?
A metadata-driven pipeline means you build one reusable pipeline that can process many tables/files based on configuration, instead of creating […]