Containers and CI/CD for a bank's analytics platform
Running Spark jobs in containers on YARN, shipping them through quality and security gates, and creating CI runners only when needed.
Commit to cluster, through every gate
IllustrativeA walk through the delivery path. No live systems.
Commit
A change to a Spark job or model workflow is pushed. An on-demand Podman runner starts for this run only.
Quality gate
Static analysis must pass the gate before anything is built.
Security scan
Findings above the threshold stop the run.
Container build
The job and its dependencies become one image.
Registry
The image is versioned; earlier versions stay available.
Run
The image runs as a container on YARN in Cloudera or on Kubernetes. ML model workflows take the same route.
Security finding
The run stops at the scan. Nothing is built or published, and the cluster keeps the previous version until the fix is pushed.
All steps as text
- Commit (GitHub): A change to a Spark job or model workflow is pushed. An on-demand Podman runner starts for this run only.
- Quality gate (SonarQube): Static analysis must pass the gate before anything is built.
- Security scan (Checkmarx): Findings above the threshold stop the run.
- Container build (Podman / Docker): The job and its dependencies become one image.
- Registry (Nexus): The image is versioned; earlier versions stay available.
- Run (YARN · Kubernetes): The image runs as a container on YARN in Cloudera or on Kubernetes. ML model workflows take the same route.
If security finding: The run stops at the scan. Nothing is built or published, and the cluster keeps the previous version until the fix is pushed.
Major challenges
- Host-bound
Jobs tied to servers
Spark and Python jobs depended on host libraries; containers on YARN gave each job its own dependencies.
- 3 gates
Regulated delivery
SonarQube, Checkmarx and Nexus became mandatory stages for every deployment.
- On demand
No idle build servers
Self-hosted runners start with Podman only when a job needs one.
- ML too
Models on the same path
Containerised ML model workflows deployed on CDP Private Cloud the same way.
My contribution
Built the container framework for Spark jobs on YARN and the gated GitHub Actions pipelines, including on-demand Podman runners.
Tools used
- Cloudera CDP
- Docker / Podman
- YARN
- Kubernetes
- GitHub Actions
- SonarQube
- Checkmarx
- Nexus
- Ansible
Outcomes
- Server-based analytics workflows moved onto a container framework on CDP.
- Every big-data deployment passes quality, security and artefact gates.
- Self-hosted CI runners created on demand instead of kept as permanent build machines.
- Containerised ML model workflows deployed on CDP Private Cloud the same way.
ArchitectureAnonymised component diagram
Commit
- Commit to GitHub
- On-demand Podman runnerContainer in container
Gates
- SonarQube quality gate
- Checkmarx security scan
Build and publish
- Container image
- Nexus artefact repository
Run
- Docker on YARN (CDP)
- Kubernetes
- ML model workflowsCDP Private Cloud
Read the flows as text
- Commit to GitHub → On-demand Podman runner
- On-demand Podman runner → SonarQube quality gate
- On-demand Podman runner → Checkmarx security scan
- SonarQube quality gate → Container image
- Checkmarx security scan → Container image
- Container image → Nexus artefact repository
- Nexus artefact repository → Docker on YARN (CDP)
- Nexus artefact repository → Kubernetes
- Nexus artefact repository → ML model workflows
Full case studyProblem, decisions, implementation, rollout
Problem and constraints
The advanced-analytics team ran Spark and Python workloads on shared servers. Each job depended on whatever libraries were installed on the host, deployments were manual, and there was no consistent quality or security gate before code reached the cluster.
- Regulated environment. Security scanning and an audit trail were mandatory.
- Keep CDP. The bank’s Cloudera platform was where workloads ran, so the solution had to work with YARN.
- Shared infrastructure. Build machines couldn’t be left as long-lived, hand-configured servers.
My role and the team’s
I built the container framework and the CI/CD pipelines. I also led the analytics team’s side of the move from Hortonworks to Cloudera CDP. The analytics team owned their jobs and models; I gave them the platform and the delivery path.
Key decisions and trade-offs
- Docker on YARN rather than a new cluster. Running containers through YARN kept workloads on the platform, security model and capacity the bank already operated. Kubernetes was used alongside it where it fit, rather than replacing YARN outright.
- Gates in the pipeline, not in review. Quality (SonarQube) and security (Checkmarx) became mandatory pipeline stages, so every deployment was checked the same way.
- Ephemeral runners. Podman-based runners start when a job needs one, so there are no permanent build VMs to patch and drift.
Implementation
- A framework for running Spark and big-data jobs inside Docker/Podman containers scheduled by YARN on CDP.
- Server-based workflows moved onto it, each with its own image and dependencies.
- GitHub Actions pipelines: build, SonarQube, Checkmarx, then publish versioned artefacts to Nexus.
- Self-hosted GitHub Actions runners started on demand with Podman (container in container).
- Ansible for Linux and Windows configuration management.
Testing and rollout
Server-based workflows were migrated onto the container framework and released through the same gated pipeline. Versioned artefacts in Nexus keep earlier builds available.
Verified outcomes
Analytics workloads became reproducible containers instead of host-dependent scripts, every deployment passed the same quality, security and artefact gates, and ML model workflows ran on CDP Private Cloud through the same path.
Domains CI/CD & automation · Containers & orchestration · Processing & workflows · AI/ML enablement · Security & governance
Code Internal bank platform, so there is no public code.