Re-platforming a Hadoop application onto Azure
Moving an on-premises Spark, Kafka, HDFS and MySQL application onto Azure's managed data services.
Each tier, mapped to a managed service
IllustrativeSimplified tier mapping for the re-platformed application.
On-premises Hadoop
- HDP Spark
- HDFS
- Kafka
- MySQL
Azure
- Synapse Studio
- Blob Storage
- Event Hubs
- Azure SQL
Select a change
Managed Spark with no cluster to run
Comparison as text
- HDP Spark → Synapse Studio: Managed Spark with no cluster to run
- HDFS → Blob Storage: Object storage replaces the HDFS layer
- Kafka → Event Hubs: Managed event streaming
- MySQL → Azure SQL: Managed relational database
Major challenges
- 4 tiers
Re-platform, don't lift
Spark, HDFS, Kafka and MySQL each mapped to an Azure managed service.
- Hybrid
Data still arriving on-premises
NiFi flows land batch and streaming data in Azure ADLS.
- Releases
A repeatable path to Azure
CI with Azure DevOps and Jenkins, release automation with Azure DevOps.
My contribution
Worked on the platform side of the migration, plus NiFi ingestion into ADLS and Azure DevOps CI/CD on the same estate.
Tools used
- Azure Synapse
- Azure Blob Storage / ADLS
- Azure SQL
- Azure Event Hubs
- Azure DevOps
- NiFi
- Spark
- Kafka
Outcomes
- Application tiers re-platformed from on-premises Hadoop onto Azure managed services.
ArchitectureAnonymised component diagram
On-premises (before)
- HDP Spark
- HDFS
- Kafka
- MySQL
Azure (after)
- Synapse Studio
- Blob Storage
- Event Hubs
- Azure SQL
Read the flows as text
- HDP Spark → Synapse Studio
- HDFS → Blob Storage
- Kafka → Event Hubs
- MySQL → Azure SQL
Full case studyProblem, decisions, implementation, rollout
Problem and constraints
An internal three-tier application ran on the on-premises Hadoop stack: HDP Spark for processing, HDFS for storage, Kafka for events and MySQL for the application database. The goal was to run it on Azure managed services rather than lift the Hadoop servers as they were.
My role and the team’s
A team migration; I worked on its platform side.
- The migration itself. Moving the application’s HDP Spark, HDFS, Kafka and MySQL services to Azure Synapse Studio, Blob Storage, Event Hubs and Azure SQL.
- Related Azure ingestion. Configuring NiFi batch and streaming pipelines that ingest from HDFS, Kafka, SFTP and SQL sources into Hive, Azure ADLS and Elasticsearch.
- Related delivery automation. Continuous integration with Azure DevOps and Jenkins (agents, build jobs, plug-ins, distributed builds), and continuous deployment with Azure DevOps release automation.
- Azure services in my toolset at the time: Synapse, Blob Storage, Azure SQL, Event Hubs, ExpressRoute and Bicep.
Architecture: tier by tier
| On-premises | Azure | What changes |
|---|---|---|
| HDP Spark | Synapse Studio (Spark) | Managed Spark; no cluster to run |
| HDFS | Blob Storage | Object storage replaces the HDFS layer |
| Kafka | Event Hubs | Managed event streaming |
| MySQL | Azure SQL | Managed relational database |
Key decisions and trade-offs
Re-platforming onto managed services removes the Hadoop cluster this application depended on, in exchange for tighter coupling to Azure’s services. Lifting the Hadoop VMs would have kept everything portable, but also kept the work of running the cluster.
Verified outcomes
The application was re-platformed from on-premises Hadoop onto Azure managed services. In 2024 I applied the same patterns in a modular Terraform landing zone for an Azure analytics demo (ADLS Gen2, Synapse, Data Factory, Databricks, HDInsight Spark and Kafka, private endpoints). (Demo work.)
Domains Cloud & infrastructure · Data platforms · Streaming & ingestion · CI/CD & automation
Code Internal employer application, so there is no public code. A separate Terraform landing zone I wrote for a 2024 demo is not linked publicly yet.