ProductionClient assignment via BBI.ai · Government security organisation2025 – 2026

Turning evidence files into governed, queryable data

A government data platform that made investigative evidence searchable. It went from an empty, air-gapped site to three production use cases.

EvidenceRegistryRaw → curatedWrite-back

Follow one document through the platform

Illustrative · synthetic dataSynthetic document. The real data, schemas and entity types are confidential.

Synthetic scanned document · DEMO-0427

What changed here? · Source

Before

  • Scanned PDF stored inside an operational database
  • Not searchable or linked to anything

Action

The source registry picks it up and copies it to the raw layer with a SHA-256 hash.

After

  • raw / doc = DEMO-0427
  • sha256 9f2c…e1 recorded

All steps as text
  1. Source: The source registry picks it up and copies it to the raw layer with a SHA-256 hash. Result: raw / doc = DEMO-0427; sha256 9f2c…e1 recorded
  2. OCR: OCR turns scanned pages into machine-readable text. Result: “Reference DEMO-1042 … dated 12 May 2026 … total 1,250.00”
  3. Extraction: Deterministic rules pull out entities, so the same input always gives the same result. Result: reference: DEMO-1042; date: 2026-05-12; amount: 1250.00
  4. Curated layer: Values are validated and written to the curated layer, keyed back to the source and its hash. Result: curated.documents · DEMO-0427; source reference and hash kept
  5. Write-back: Results are published back to the source database on a schedule. Result: Existing applications read the new fields with no changes on their side

Major challenges

  1. Air-gapped

    Build with no internet

    Every Python and OS dependency was packaged offline so builds stay repeatable inside the air gap.

  2. 3 use cases

    Three evidence formats

    Documents, laboratory packages and device extracts, each with its own pipeline, all taken to production.

  3. No ML

    Explainable extraction

    Machine learning was out of scope, so OCR plus deterministic rules give the same output for the same input.

  4. 100% match

    Reconcile with the source

    The structured use case reconciles table for table with its source database before sign-off.

  5. 4 revisions

    Formal sign-off gates

    Requirements went through four revisions with the client, then acceptance and production gates.

My contribution

Technical lead and build lead for all three use cases. Wrote one use case's requirements, owned the sign-offs and ran the handover.

Tools used

  • Cloudera CDP
  • PySpark
  • Hive
  • Oozie
  • Oracle
  • Informatica
  • Kerberos

Outcomes

  • Three use cases running in production as scheduled workflows.
  • Delivered beyond the agreed scope, including OCR for scanned documents and extra source tables.
  • Full reconciliation between the platform and the source database for the structured use case.
  • Platform, runbooks and guides handed over to the client's own team.
ArchitectureAnonymised component diagram

Sources

  • Document attachmentsStored in an operational database
  • Laboratory packages
  • Structured device extracts

Landing

  • Source registry and ingest jobsNew tables added by configuration
  • Raw layerHDFS and Hive

Processing

  • PySpark jobs on Oozie
  • OCR and deterministic extraction

Curated

  • Cleaned layer
  • Curated layer

Consumers

  • Write-back to source database
  • Existing downstream consumers
Simplified, anonymised data flow. Security (Kerberos, Ranger, TLS) and an offline package mirror apply to every layer.
Read the flows as text
  • Document attachments → Source registry and ingest jobs
  • Laboratory packages → Source registry and ingest jobs
  • Structured device extracts → Source registry and ingest jobs
  • Source registry and ingest jobs → Raw layer
  • Raw layer → PySpark jobs on Oozie
  • Raw layer → OCR and deterministic extraction
  • OCR and deterministic extraction → PySpark jobs on Oozie
  • PySpark jobs on Oozie → Cleaned layer
  • Cleaned layer → Curated layer
  • Curated layer → Write-back to source database
  • Write-back to source database → Existing downstream consumers
Full case studyProblem, decisions, implementation, rollout

Problem and constraints

Investigative evidence arrived in very different shapes: document attachments in an operational database, packages from laboratory instruments, and structured device extracts. Each sat in its own silo, so none of it could be searched or related across cases. The organisation wanted a governed platform that turns this evidence into structured, queryable data, without changing how existing downstream systems consume it.

  • Air-gapped and fully on site. The cluster had no internet access, so every Python and OS dependency had to be packaged and brought in.
  • No machine learning. The requirements explicitly excluded ML, so extraction had to be deterministic and explainable.
  • Integrity. Outputs needed hashing and traceability back to the source.
  • Formal gates. Requirements, acceptance testing and production each needed client sign-off.

My role and the team’s

I was the programme’s technical lead and the build lead for the three use cases. I wrote the requirements for the document-processing use case through four revisions with the client and co-authored the other two. I also owned the acceptance and production sign-offs and ran the handover sessions.

The parallel data warehouse and data-quality streams were built by colleagues; I coordinated them through the same gates.

Key decisions and trade-offs

  • Layered lake on Cloudera, scheduled by Oozie. Oozie already ships with the Cloudera platform, so the air-gapped site didn’t need another orchestrator installed, secured and handed over.
  • Configuration over code for sources. A source registry means a new table is added by configuration rather than a code change. It paid off when the client asked for additional source tables.
  • Write back, don’t replace. Curated results are written back to the existing database, so downstream consumers needed no changes. This moved the integration risk onto us and away from the client’s systems.
  • Deterministic extraction. The requirements excluded machine learning, so extraction uses OCR plus rules whose results can be explained and reproduced.

Implementation

  • PySpark jobs move data through raw, cleaned and curated layers in HDFS and Hive.
  • Each use case is an Oozie workflow with its own schedule and rerun path.
  • An offline Python wheelhouse and OS package bundle make builds repeatable inside the air gap.
  • Hashes and source references are kept with every curated record.

Testing, rollout and handover

  • Acceptance testing against agreed scenarios, signed off by the client before anything went to production.
  • For the structured use case, reconciliation checks compared every table between the platform and the source database before sign-off.
  • Production rollout was gated, with first scheduled runs watched before handover.
  • Handover sessions, installation and user guides, and runbooks for the client’s operations team.

Verified outcomes

All three use cases went live as scheduled production workflows. Scope grew without slipping the gates: OCR for scanned documents, additional source tables and profile clustering were delivered beyond the original agreement. The structured use case reconciles fully with its source.

Domains Data platforms · Processing & workflows · Streaming & ingestion · Security & governance · Observability & reliability

Trace it in the constellation

Code No public code. The client, its data and its internal architecture are confidential.

Search the portfolio

Try:

↑ ↓ to move · Enter to open · Esc to close