Platforms

Nine connected domains, and how a data platform built from them actually works.

Engineering constellation

Select a domain to see my work there and its connections, or trace a project across domains.

Data platforms

The centre of my work: Cloudera and Hadoop platforms I've administered, migrated, secured and handed over since 2015.

More detail
  • Cloudera CDP Production

    Migrated a telecom data lake from HDP to CDP, ran the bank's analytics platform on it, and now install, secure, upgrade and hand over CDP clusters for clients. Cloudera CDP Private Cloud Base certified.

  • Hadoop (HDFS, YARN, Hive, Ozone) Production

    Ten years administering Hadoop. At Jio that meant NameNode federation, HDFS erasure coding and Ozone. Earlier work covered Hortonworks clusters, and more recently a documented open-source Hadoop build with Kerberos.

  • ← Streaming & ingestion: Kafka and NiFi land events and files in the lake.
  • → Processing & workflows: Spark jobs read raw lake data and write curated layers back.
  • ← CI/CD & automation: Ansible checks and configures every host before and after a cluster install.
  • → Cloud & infrastructure: Hadoop workloads re-platformed onto Azure managed services.
  • ← Security & governance: Kerberos, Ranger and TLS protect data and access at every layer.

From source to serving

Step through a representative platform: what moves, what triggers the next stage, and where the checks happen.

Illustrative

A representative, anonymised system rather than an exact production deployment.

SourcesDatabases · filesIngestionNiFi · Kafka · CDCProcessingSpark in containersStorageRaw → curatedServingSQL · apps · MLCI/CDScan · build · publishOrchestrationOozie · AirflowObservabilityMetrics · alertsSecurityKerberos · TLS · Ranger · auditSourcesDatabases · filesIngestionNiFi · Kafka · CDCProcessingSpark in containersStorageRaw → curatedServingSQL · apps · MLCI/CDScan · build · publishOrchestrationOozie · AirflowObservabilityMetrics · alertsSecurityKerberos · TLS · Ranger · audit

Step 1 of 9

All steps as text
  1. Data arrives. Databases, event streams and files produce new data all day.
  2. Landed untouched. NiFi, Kafka and change capture copy data into the raw layer and record each landing. Check: Integrity check: hash and schema on arrival.
  3. Arrival triggers work. The orchestrator starts processing only when upstream data has landed, and retries transient failures with backoff.
  4. Code arrives through gates. Job code is quality-gated, security-scanned and built into a versioned image before the cluster runs it. Check: Quality and security gates.
  5. Spark transforms. Containerised Spark jobs clean, join and validate the raw data. Check: Validation before anything is promoted.
  6. Layered storage. Data moves raw → cleaned → curated in HDFS or Ozone with Hive tables; only validated data reaches curated.
  7. Served to consumers. Analysts, applications and ML workloads read curated tables. Check: Access policy checked on every query.
  8. Secured at every layer. Kerberos authentication, TLS in transit, Ranger policies and alerts on privileged logins.
  9. Watched end to end. Metrics, logs and alerts across ingestion, jobs and serving, with runbooks and rollback plans for recovery.

Search the portfolio

Try:

↑ ↓ to move · Enter to open · Esc to close