Orchestrating ML Pipelines with Azure Data Factory
Leveraging Azure Data Factory for Scalable and Efficient Machine Learning Workflows
Lesson preparation & details
Level: beginner
By the end, you should be able to
- Training fails after writing half a model file. How should the downstream inference activity know not to consume it?
- Explain the version and execution boundaries before applying the examples
Bring with you
- Basic programming and HTTP; follow the chapter or cloud-track sequence
Editorial review: · What review means
In this article · 23 sections
Review and execution boundary
Azure Data Factory v2 Custom activity on Azure Batch; not the similarly named Fabric Azure Batch activity.
Reviewed on 7 October 2026 against the official source snapshots linked below. The review is bounded editorial correction, not certification of every dependency, security property or cloud deployment. Historical setup commands and optional exercises were not executed. No cloud resources, third-party packages or external side effects were created. Old screenshots and unavailable private assets remain in the private recovery archive, not prerequisites for this lesson.
Using ADF pipelines for orchestrating ML training & inference
Introduction
An ML workflow needs more than a training script: data must arrive in storage, compute must run the script, and the resulting model must be available for inference. This MNIST proof of concept separates those responsibilities: Blob Storage holds the artifacts, Azure Batch runs Python tasks, and Azure Data Factory (ADF) is the proposed orchestrator. The walkthrough records the resource setup and Batch job; the ADF-triggered run remains a next step rather than a completed end-to-end test.
Project Setup
Creating a GitHub Repository
To begin, we'll create a new repository on GitHub to store our project files.
Setting Up Azure Resources
Next, we'll create a new resource group in Azure specifically for this project.
Resource Group Name: adf-ml-project
Integrating with Azure Data Factory
Connecting GitHub to ADF
To leverage version control and collaborative features, we'll connect our GitHub repository to Azure Data Factory.
- In the Azure Data Factory interface, navigate to the "Author & Monitor" section.
- Click on "Set up code repository" to begin the integration process.
Authorizing ADF Access
During the setup, you'll be prompted to authorize Azure Data Factory to access your GitHub account. This step is crucial for enabling seamless integration between ADF and your code repository.
Configuring Repository Settings
After authorization, choose which repository, development branch and root folder will hold the authored ADF definitions. The configuration screenshot records those choices; distinguish them from the publish branch used later for deployment templates.
Key points to note:
- Select your repository from the dropdown.
- Choose the branch you want to use for development.
- ADF will automatically create a "publish" branch for ARM templates.
- Select a repository root folder for ADF definitions. This does not convert arbitrary scripts into ADF JSON.
I will be using the MNIST dataset for this ML model : https://www.kaggle.com/code/prashant111/mnist-deep-neural-network-with-keras In this version we will build everything from UI of ADF itself , but we will document all steps one by one.
Creating a blob storage in the resource group which we will use for storing data and storing trained model
The repository stores the workflow definitions; Blob Storage has a different role. It holds the MNIST inputs and the trained model so that separate tasks can exchange artifacts without depending on one machine's local files. Create the storage account before configuring those tasks.
Using LRS type redundancy as Locally Redundant Storage is good for our usecase but for production systems it is better to have RA-GRS type redundancy as this is a GEO-redundant storage and copies 3 times locally and 3 times in the other region.
Created a container adftrain and folders for storing training data, inference data and trained model.
Adding access to the storage account via linked service
Add the Managed Identity of the ADF workspace and provide the least-privilege Blob data-plane role required by the linked service as below.
Once added then add this account to the linked service in the ADF workspace. Validate with the Test Connection at the bottom to check if the connection is successful.
Key points to note:
- With the above steps we created a ADF workspace, Storage account and Linked Service to access storage account from ADF service.
- Going forward we can keep publishing ARM template for any changes so that the repo : https://github.com/dinesh-coderepo/adf-ml-project ,
branch : adf_publishis updated with the latest changes. - Publishing to adf_publish involves some checks from ADF and then the ARM template is published.
- Incase of any issues we can revert to the previous commit in the repo.
- Ideally in production we should not publish the code changes directly to the repo without testing in the dev environment.
- We can have a CI/CD pipeline to monitor any changes to adf_publish branch and it detects any changes it will deploy these to target environment such as staging and production
- The deployment uses the ARM template to create update or delete resources in target ADF environment.
- We will cover this entire setup in an another blog post.
Data Flow Graph
Read this graph as the intended lifecycle of the MNIST artifacts, not as a record of an executed ADF pipeline. Blob Storage appears at both the input and output of training: the training task reads data from it and writes the model back for inference.
In this graph:
- A : External Source (in our case, keras.datasets)
- B : Local Processing (preprocessing if needed)
- C : Blob Storage (where we store our data and trained model)
- D : ML Model (training process)
- E : Prediction Service (for inference)
Uploading data to blob storage
As part of next steps first we will get the mnist data from keras datasets and then upload it to blob storage. We are trying to mimic a real time scenario where we first get the data from some external source and then upload it to blob storage, in this case we are are getting the data from keras.datasets module then uploading it to blob storage.
Creating a Batch Account to run the scripts we would need
Storage alone cannot execute the download script. Azure Batch supplies that compute: the account manages the Batch resources, the pool supplies worker nodes, and a job groups the tasks to run on them. The archived screenshots follow that setup from account to pool to task.
-
Attaching the storage account to the batch service, with user managed identity authentication.
-
Created a User managed identity, assigned the required Blob data-plane role at the smallest appropriate scope and added the identity to the batch account. After this we can link the storage account to the batch account.
-
Creating a Batch Pool and added the commands to install the python packaged neccesary for the job
-
After creating the pool we create a job with a task to run the script to download the data from keras.datasets and upload it to blob storage.
with this setup we can replicate the same for training and inference by just changing the script in the task.
- Follow below steps to configure on running the batch job from ADF pipeline, same configuration can be used for training and inference steps.
- Note : Keeping in mind the costs and this is a POC , not triggering from ADF for this test setup.
- In future I will be using this framework to implement a real world application in Upcoming projects.
Below flow can illustrate the flow for dumping data using batch and ADF
The first graph followed the data; this one follows the proposed control path from an ADF trigger to the upload task. Treat the pool node as the compute used by the task, not as proof that creating a job automatically provisions a pool. As noted above, this ADF-triggered path was not run in the cost-limited POC.
Running Azure Batch Jobs in ADF Pipeline
To integrate our Azure Batch job into the ADF pipeline, we'll follow these steps:
- Create a Linked Service for Azure Batch
- Create a Pipeline with Azure Batch Activity
- Publish and Run the Pipeline
Create a Linked Service for Azure Batch
First, we need to create a Linked Service to connect ADF to our Azure Batch account:
- In ADF Studio, go to the "Manage" tab
- Select "Linked services" and click "New"
- Search for and select "Azure Batch"
- Configure the Linked Service with your Batch account details
Create a Pipeline with Azure Batch Activity
Now, let's create a pipeline that uses the Azure Batch Activity:
- In ADF Studio, go to the "Author" tab
- Create a new pipeline
- Drag and drop the "Custom" activity from the Activities pane
- Configure the Azure Batch activity with your script details and Linked Services
Publish and Run the Pipeline
After setting up the pipeline:
- Click "Publish All" to save your changes
- Click "Add trigger" > "Trigger now" to run the pipeline manually
Key points to note:
Before extending this setup to training and inference, verify a complete run: the upload task should leave data in the expected container, training should produce a saved model, and inference should load that artifact. A successful linked-service connection alone does not establish those outcomes. Also check the identities' data-plane permissions: management-plane Contributor access alone does not grant Blob data access.
- Ensure your Python script is uploaded to the specified Blob storage container.
- The Azure Batch activity in ADF allows you to run your data processing and ML tasks at scale.
- You can add additional activities before or after the Batch activity for more complex workflows.
- This setup provides a scalable way to integrate your data processing and ML tasks into your ADF pipelines, allowing for better orchestration and management of your entire ML workflow.
Corrected contracts and failure analysis
ADF orchestrates dependencies and activity status; Batch executes custom code in an existing configured pool; Blob Storage carries immutable inputs and versioned model artifacts. The ADF activity is Custom activity, not a generic button that runs an already-created arbitrary Batch job. Match the linked-service authentication mode to the connector documentation rather than assuming every service supports the same identity shape.
Contributor grants management operations, not Blob data-plane access. Scope Storage Blob Data Reader or Contributor to the needed container and principal; evaluate Batch task identity separately from ADF identity. Git integration stores authored ADF definitions, not arbitrary Python transformed into JSON because a root folder is named src. Publishing templates is not a transactional rollback of data or already-executed tasks.
Use a run ID and input/model digest in artifact paths. A retry must either overwrite an identical output safely or detect that its intended version already exists. Validate saved-model readability before publishing a completion marker; do not let inference discover a partially written model.
Boundary exercise with solution
Training fails after writing half a model file. How should the downstream inference activity know not to consume it?
Solution and reasoning
Write to a run-specific temporary location, verify the artifact, then publish a manifest/completion marker with its digest. Depend on successful activity completion and the verified manifest, not blob existence alone. Retrying must reuse the same input identity or create a new explicit run version.
Source-backed review notes
- Use custom activities in a pipeline - Azure Data Factory & Azure Synapse | Microsoft Learn — accessed 2026-10-07. Exact supporting passage: “To move data to/from a data store that the service does not support, or to transform/process data in a way that isn't supported by the service, you can create a Custom activity with your own data movement or transformation logic and use the activity in a pipeline. The custom activity runs your customized code logic on an Azure Batch pool of virtual machines.”
- Assign an Azure Role for Blob Data Access - Azure Storage | Microsoft Learn — accessed 2026-10-07. Exact supporting passage: “To access blob data in the Azure portal by using Microsoft Entra credentials, a user must have the Azure Resource Manager Reader role, at a minimum, in addition to a data access role such as the Storage Blob Data Reader or Storage Blob Data Contributor role. See Data access from the Azure portal.”
Pause / Recall / Apply
Can you explain it without the page?
Close the example. Reconstruct the core idea, then change one assumption. Mark complete when you’re ready; you can always undo it.
Stored in this browser only. No account, no sync. Clearing browser data removes your record.