Processing Payments Big Data with AWS EMR Serverless and S3 Tables

PRACTICAL BIG DATA ENGINEERING

Processing Payments Big Data with AWS EMR Serverless and S3 Tables

How do you run scalable Spark batch processing for payment data without managing clusters, or even Docker images? In this talk, Ioana will share how Worldpay processes around 100 million payment events daily, generating approximately 670 million records for reporting purposes, which are ultimately aggregated into around 200,000 records.

We’ll explore how we run Spark batch jobs at scale on AWS EMR Serverless without relying on Docker images, and our experience migrating workloads and persistence from on-premises infrastructure with relatively minor code changes. The session covers the key challenges we encountered and the solutions we implemented to improve performance and efficiency, including reducing dependency on a single infrastructure team, enabling parallel Spark job execution, supporting multi-region data persistence, and improving failover and resilience. You will also learn about our adoption of S3 Tables as an alternative persistence layer to relational databases, and how they fit into our data processing architecture. During the session, you will gain practical insight into our approach through demonstrations of some of our most important strategies in action.

Book Now