درباره این اپیزود
Welcome back to Episode 7! Today, let’s look at Data Factory inside Fabric and how we bring data into the system.
Now, not every team wants to write lines of Python or Spark code just to move files around. Fabric includes simple, visual tools for this, broken into two main parts:
1. Data Pipelines: Think of a pipeline like a project manager for your data. If you need to pull daily sales records from an office SQL server, check if the file
exists, run a Spark script to process it, and send a Teams or email alert if anything fails—you build that visually inside a pipeline. It comes with ready-made
connectors for databases, APIs, and cloud tools, so you never have to write custom connection code from scratch.
2. Dataflows Gen2: This is essentially Power Query in your web browser. If a business analyst needs to clean messy data—like splitting full names into first
and last, removing blank rows, or combining two sheets—they don't need Python. They can just point, click, and transform data right on their screen.
The biggest upgrade with Dataflows Gen2 is where that clean data goes.
In older setups, cleaned Power Query data was locked inside Power BI. With Gen2 in Fabric, the moment you finish cleaning a table, it writes directly into OneLake as
open Delta-Parquet files.
That means an analyst can clean a dataset with simple mouse clicks in Power Query, and a data engineer can immediately pick up that exact same table in Python or
SQL without any extra export steps.
In Episode 8, we’ll move away from scheduled daily batches and look at handling live, real-time streaming data.
About the Host
Rajnish Pandey is a Senior Data Engineer with over decade of experience in the data industry. Through QueryZens, he shares practical insights, real-world experiences, and conversations around Data Engineering and modern data platforms.
Hosted by Rajnish Pandey | QueryZens
Making Data Engineering easier to understand, one conversation at a time.