ProductionEmployer · Reliance Jio2018 – 2022

Migrating one of India's largest data lakes to Cloudera CDP

Moving a live telecom data lake from Hortonworks to Cloudera without stopping the business that depended on it.

HDPCDPKafka

Hortonworks to Cloudera, component by component

IllustrativeSimplified view of the platform before and after the migration.

HDP (before)

  • HDFS · single namespace
  • NiFi on HDF
  • Kafka · ~150 nodes
  • Hive
  • Kerberos · Ranger

CDP (after)

  • HDFS federation · erasure coding · Ozone
  • NiFi on CFM
  • Kafka on CDP
  • Hive · Data Analytics Studio · Workload Manager
  • Kerberos · SSSD · Ranger · Knox · Atlas

Select a change

Namespace federation, erasure coding and Ozone configured and tested

Configure new capabilities. Federation, Ozone, Workload Manager and Data Analytics Studio, tested on CDP.

Comparison as text
  • HDFS · single namespace → HDFS federation · erasure coding · Ozone: Namespace federation, erasure coding and Ozone configured and tested
  • NiFi on HDF → NiFi on CFM: Batch and streaming flows rebuilt with sized repositories
  • Kafka · ~150 nodes → Kafka on CDP: Streaming estate carried across
  • Hive → Hive · Data Analytics Studio · Workload Manager: Query monitoring and workload management added
  • Kerberos · Ranger → Kerberos · SSSD · Ranger · Knox · Atlas: Security rebuilt as part of the platform
  1. Configure new capabilities: Federation, Ozone, Workload Manager and Data Analytics Studio, tested on CDP.
  2. Rebuild flows: NiFi clusters on CFM for batch and streaming.
  3. Test pipelines end to end: Production pipelines, use cases and applications validated on the new stack.
  4. Operate: Monitoring, alerting and hardening for a platform of about 800 servers.

Major challenges

  1. ~800 servers

    Migrate a live lake

    Move a platform of about 800 servers and 50 PB capacity without stopping the business on it.

  2. ~150 nodes

    Carry the streaming estate

    A Kafka setup of around 150 nodes in a single cluster had to keep flowing.

  3. Millions of files

    Namespace at scale

    NameNode federation and erasure coding to handle namespace size and storage cost.

  4. End to end

    Prove consumers still work

    Production pipelines were validated on CDP, not just services.

My contribution

Lead contributor in the platform team. Configured and tested the new CDP capabilities, rebuilt NiFi flows and tested production pipelines end to end.

Tools used

  • Hortonworks HDP
  • Cloudera CDP
  • HDFS
  • Ozone
  • NiFi
  • Kafka
  • Ranger
  • Kerberos

Outcomes

  • Platform moved from HDP to CDP, with production pipelines tested end to end on the new stack.
  • Recognised with the RIL Samman award for contributions to the migration.
ArchitectureAnonymised component diagram

Sources

  • Network and business systems

Ingestion

  • NiFi flowsHDF, then CFM
  • KafkaAbout 150 nodes

Platform

  • HDP cluster (before)
  • CDP cluster (after)Federation, erasure coding, Ozone

Consumers

  • Analytics pipelines
  • Consuming teams
Simplified view of the migration. Security (Kerberos, SSSD, Ranger, Knox, Atlas) was rebuilt on the new platform rather than bolted on afterwards.
Read the flows as text
  • Network and business systems → NiFi flows
  • Network and business systems → Kafka
  • NiFi flows → HDP cluster (before)
  • Kafka → HDP cluster (before)
  • HDP cluster (before) → CDP cluster (after) (migrate and test)
  • CDP cluster (after) → Analytics pipelines
  • Analytics pipelines → Consuming teams
Full case studyProblem, decisions, implementation, rollout

Problem and constraints

The Jio Big Data Lake was one of India’s largest on-premises data platforms. As recorded in my career records, it had approximately 800 servers and 50 PB of HDFS capacity and took in data from across a national telecom network. Those figures describe the platform’s scale, which a whole team ran. They are not something I built single-handedly.

It ran on Hortonworks HDP, which was heading for end of life after the Cloudera merger, so it had to move to Cloudera CDP while it kept serving the business.

  • Scale. Millions of files on the NameNode and continuous ingestion left no quiet window for a big-bang cutover.
  • Many consumers. Pipelines owned by many teams had to keep working.
  • Security parity. Kerberos, LDAP/SSSD, TLS and Ranger policies all had to carry over.

My role and the team’s

I was part of the platform team that ran the lake and a lead contributor to the migration. My recorded contributions:

  • Configuring and testing NameNode federation, Workload Manager, the Ozone file system and Data Analytics Studio on CDP.
  • NiFi clusters (HDF, CFM, CDP) for batch and streaming flows, with content and provenance repositories.
  • End-to-end testing of production pipelines, use cases and applications on the new stack.
  • Exploring and documenting new CDP capabilities for the team.

The migration as a whole was a team effort.

Key decisions and trade-offs

  • Re-architect, don’t just lift. We used the move to adopt NameNode federation, HDFS erasure coding and Ozone. This was more change at once, but it avoided a second migration later.
  • Test the pipelines, not just the services. “Service is green” isn’t the same as “consumers still work”, so validation was done at pipeline level.
  • Security first. Kerberos, SSSD with OU-level configuration, TLS, Ranger, Knox and Atlas tagging were part of the build.

Implementation

  • Federation and erasure coding to handle the namespace size and storage cost at this scale.
  • NiFi flows rebuilt on CFM, with repositories sized for the volume.
  • Monitoring with Check-MK, Grafana and Prometheus, plus Dr. Elephant for Spark jobs, Kafka-HQ, ElastAlert and a custom KDC monitor.

Testing and rollout

Production pipelines were tested end to end on CDP before consumers moved. Platform operations carried on throughout, with hardening and alerting keeping a platform this size healthy.

Verified outcomes

The lake moved to CDP with its production pipelines validated, and my contribution was recognised with the RIL Samman award from the platform head.

Search the portfolio

Try:

↑ ↓ to move · Enter to open · Esc to close