Migrating one of India's largest data lakes to Cloudera CDP
Moving a live telecom data lake from Hortonworks to Cloudera without stopping the business that depended on it.
Hortonworks to Cloudera, component by component
IllustrativeSimplified view of the platform before and after the migration.
HDP (before)
- HDFS · single namespace
- NiFi on HDF
- Kafka · ~150 nodes
- Hive
- Kerberos · Ranger
CDP (after)
- HDFS federation · erasure coding · Ozone
- NiFi on CFM
- Kafka on CDP
- Hive · Data Analytics Studio · Workload Manager
- Kerberos · SSSD · Ranger · Knox · Atlas
Select a change
Namespace federation, erasure coding and Ozone configured and tested
Configure new capabilities. Federation, Ozone, Workload Manager and Data Analytics Studio, tested on CDP.
Rebuild flows. NiFi clusters on CFM for batch and streaming.
Test pipelines end to end. Production pipelines, use cases and applications validated on the new stack.
Operate. Monitoring, alerting and hardening for a platform of about 800 servers.
Comparison as text
- HDFS · single namespace → HDFS federation · erasure coding · Ozone: Namespace federation, erasure coding and Ozone configured and tested
- NiFi on HDF → NiFi on CFM: Batch and streaming flows rebuilt with sized repositories
- Kafka · ~150 nodes → Kafka on CDP: Streaming estate carried across
- Hive → Hive · Data Analytics Studio · Workload Manager: Query monitoring and workload management added
- Kerberos · Ranger → Kerberos · SSSD · Ranger · Knox · Atlas: Security rebuilt as part of the platform
- Configure new capabilities: Federation, Ozone, Workload Manager and Data Analytics Studio, tested on CDP.
- Rebuild flows: NiFi clusters on CFM for batch and streaming.
- Test pipelines end to end: Production pipelines, use cases and applications validated on the new stack.
- Operate: Monitoring, alerting and hardening for a platform of about 800 servers.
Major challenges
- ~800 servers
Migrate a live lake
Move a platform of about 800 servers and 50 PB capacity without stopping the business on it.
- ~150 nodes
Carry the streaming estate
A Kafka setup of around 150 nodes in a single cluster had to keep flowing.
- Millions of files
Namespace at scale
NameNode federation and erasure coding to handle namespace size and storage cost.
- End to end
Prove consumers still work
Production pipelines were validated on CDP, not just services.
My contribution
Lead contributor in the platform team. Configured and tested the new CDP capabilities, rebuilt NiFi flows and tested production pipelines end to end.
Tools used
- Hortonworks HDP
- Cloudera CDP
- HDFS
- Ozone
- NiFi
- Kafka
- Ranger
- Kerberos
Outcomes
- Platform moved from HDP to CDP, with production pipelines tested end to end on the new stack.
- Recognised with the RIL Samman award for contributions to the migration.
ArchitectureAnonymised component diagram
Sources
- Network and business systems
Ingestion
- NiFi flowsHDF, then CFM
- KafkaAbout 150 nodes
Platform
- HDP cluster (before)
- CDP cluster (after)Federation, erasure coding, Ozone
Consumers
- Analytics pipelines
- Consuming teams
Read the flows as text
- Network and business systems → NiFi flows
- Network and business systems → Kafka
- NiFi flows → HDP cluster (before)
- Kafka → HDP cluster (before)
- HDP cluster (before) → CDP cluster (after) (migrate and test)
- CDP cluster (after) → Analytics pipelines
- Analytics pipelines → Consuming teams
Full case studyProblem, decisions, implementation, rollout
Problem and constraints
The Jio Big Data Lake was one of India’s largest on-premises data platforms. As recorded in my career records, it had approximately 800 servers and 50 PB of HDFS capacity and took in data from across a national telecom network. Those figures describe the platform’s scale, which a whole team ran. They are not something I built single-handedly.
It ran on Hortonworks HDP, which was heading for end of life after the Cloudera merger, so it had to move to Cloudera CDP while it kept serving the business.
- Scale. Millions of files on the NameNode and continuous ingestion left no quiet window for a big-bang cutover.
- Many consumers. Pipelines owned by many teams had to keep working.
- Security parity. Kerberos, LDAP/SSSD, TLS and Ranger policies all had to carry over.
My role and the team’s
I was part of the platform team that ran the lake and a lead contributor to the migration. My recorded contributions:
- Configuring and testing NameNode federation, Workload Manager, the Ozone file system and Data Analytics Studio on CDP.
- NiFi clusters (HDF, CFM, CDP) for batch and streaming flows, with content and provenance repositories.
- End-to-end testing of production pipelines, use cases and applications on the new stack.
- Exploring and documenting new CDP capabilities for the team.
The migration as a whole was a team effort.
Key decisions and trade-offs
- Re-architect, don’t just lift. We used the move to adopt NameNode federation, HDFS erasure coding and Ozone. This was more change at once, but it avoided a second migration later.
- Test the pipelines, not just the services. “Service is green” isn’t the same as “consumers still work”, so validation was done at pipeline level.
- Security first. Kerberos, SSSD with OU-level configuration, TLS, Ranger, Knox and Atlas tagging were part of the build.
Implementation
- Federation and erasure coding to handle the namespace size and storage cost at this scale.
- NiFi flows rebuilt on CFM, with repositories sized for the volume.
- Monitoring with Check-MK, Grafana and Prometheus, plus Dr. Elephant for Spark jobs, Kafka-HQ, ElastAlert and a custom KDC monitor.
Testing and rollout
Production pipelines were tested end to end on CDP before consumers moved. Platform operations carried on throughout, with hardening and alerting keeping a platform this size healthy.
Verified outcomes
The lake moved to CDP with its production pipelines validated, and my contribution was recognised with the RIL Samman award from the platform head.
Domains Data platforms · Streaming & ingestion · Security & governance · Observability & reliability · CI/CD & automation
Code Internal employer platform, so there is no public code.