Project Overview
Baanx, a leading crypto-finance platform enabling users to buy, spend, and borrow against digital assets, needed to unlock actionable insights from their operational data. Their previous data management investments hadn't delivered—query applications couldn't extract insights efficiently, and analytics remained inaccessible. We designed and delivered an end-to-end ETL pipeline using AWS Glue to synchronise data from AWS RDS to Amazon Redshift, enabling powerful analytics and anomaly detection through Tableau.

The Challenge
Baanx had invested in data management but the outcomes fell short. Their existing setup couldn't extract actionable insights in an efficient or simple manner. Data was siloed in AWS RDS with no pathway to analytical tooling, preventing the team from running the queries needed for business reporting, anomaly detection, and trend analysis.
- Existing data infrastructure failed to meet business reporting requirements
- No efficient pathway from operational databases to analytical tools like Tableau
- Inability to detect anomalies or perform trend analysis on transactional data
- Manual data processes creating bottlenecks in decision-making
Solution Architecture
We implemented a robust ETL pipeline using AWS Glue as the core orchestration engine. The solution extracts data from AWS RDS, transforms it for analytical querying, and loads it into Amazon Redshift—making the data immediately accessible via Tableau for dashboards and ad-hoc analysis.
- AWS Glue Crawlers to populate the Data Catalog with metadata table definitions
- Custom AWS Glue scripts for data transformation and schema mapping
- Scheduled batch synchronisation with manual trigger capability for on-demand refreshes
- Data modelling optimised for simple querying through Tableau and other BI tools
Technical Implementation
The project was delivered in three focused milestones. First, we set up the infrastructure, tooling, IAM permissions, and Glue profiles. Next, we built the synchronisation layer with scheduled batch jobs and manual trigger support. Finally, we transformed the data models to support straightforward querying in Tableau, ensuring analysts could self-serve without engineering support.
- Infrastructure setup: AWS Glue jobs, IAM roles, S3 staging buckets, and Redshift cluster configuration—all provisioned and managed through Terraform for reproducible, version-controlled infrastructure
- Batch sync engine with configurable schedules and on-demand manual triggers
- Data transformation layer converting RDS schemas into Redshift-optimised star schemas
- Extended scope planning for change data capture (CDC) trails and computed business metrics
Technologies
Results & Impact
- Delivered the complete pipeline in under 3 weeks across three milestones
- Enabled self-service analytics through Tableau for business and operations teams
- Scheduled and on-demand data sync eliminated manual data extraction processes
- Data transformation layer reduced query complexity, making insights accessible to non-technical stakeholders
- Laid groundwork for extended scope: historical change tracking and computed business metrics
- Provided Baanx with anomaly detection capabilities across their crypto-finance transaction data
Why It Matters
By building a clean, reliable data pipeline from RDS to Redshift, we gave Baanx the analytical foundation they'd been missing. Their teams can now query transaction data, detect anomalies, and generate business reports through Tableau without engineering intervention. The modular architecture also positions them for the extended scope—historical data trails and computed metrics—as their analytics needs evolve.
Services Behind This Project
The same senior teams that delivered this work: