Skip to main content
    Scalexa — Senior Engineering & AI Solutions
    Data Engineering

    Cityblock Health - Linking Two Unlinkable Databases

    Project Overview

    Cityblock Health experienced difficulties with linking two unique databases - their data was separated and had no logical pipeline back into one unified repository. Scalexa routed both into something futureproof, without losing a decimal of data, all while keeping the project under its budgeted time frame.

    Cityblock Health - Linking Two Unlinkable Databases project screenshot

    The Challenge

    How does one get an Amazon Redshift database, full with management modules written in Python, into an ETL pipeline that is shared with a Google Cloud storage bank that will flow through Apache Beam?

    • Two completely separate database systems with no integration path
    • Amazon Redshift database with Python-based management modules
    • Google Cloud storage requiring Apache Beam data flow
    • Critical need to preserve 100% data integrity during migration
    • Strict budget and timeline constraints

    Our Solution

    Routing an ETL pipeline towards a BigQuery database that has ML and AI accents turned out to be the way forward. Both previously unprotected and separated databases were joined and hosted in a security and manageable way.

    • Designed unified ETL pipeline architecture connecting both systems
    • Implemented BigQuery as the central data warehouse
    • Added ML and AI capabilities for enhanced data insights
    • Created secure, futureproof data infrastructure
    • Eliminated data isolation and silos completely

    Technical Implementation

    We built a comprehensive data engineering solution that seamlessly bridges Amazon Redshift and Google Cloud ecosystems while maintaining data integrity and security.

    • Custom Python connectors for Amazon Redshift data extraction
    • Apache Beam pipelines for scalable data processing
    • Google BigQuery integration for unified data storage
    • Real-time data synchronization between systems
    • Comprehensive data validation and quality checks
    • Automated monitoring and alerting systems

    Technologies Used

    We leveraged a modern data engineering stack to deliver a robust and scalable solution.

    • Amazon Redshift for source data warehouse
    • Google BigQuery for unified data repository
    • Apache Beam for ETL pipeline orchestration
    • Python for custom data transformations
    • Google Cloud Platform APIs
    • ML/AI integrations for enhanced analytics

    Key Results

    • Successfully unified two previously incompatible database systems
    • Zero data loss during migration - preserved every decimal
    • Delivered project under budget and ahead of schedule
    • Created futureproof infrastructure preventing data isolation
    • Enhanced security for previously unprotected databases
    • Enabled ML and AI capabilities for advanced analytics

    Why It Matters

    The client now sits with a futureproof tool that prevents data isolation. By bridging the gap between Amazon Redshift and Google BigQuery through a carefully architected ETL pipeline, we transformed Cityblock Health's fragmented data landscape into a unified, secure, and intelligent data platform ready for the future.

    Get Started

    Ready to Get Started?

    Book a free 30-minute discovery session with our senior engineers to identify quick wins and show you what's possible.

    View Our Work