Senior Data Reliability Engineer AWS
Quick Summary
Own production data pipeline and platform reliability, resolving failures, delays, data quality issues, and performance problems. Troubleshoot distributed data-processing workloads across Spark/EMR,
Our vision for the future is based on the idea that transforming financial lives starts by giving our people the freedom to transform their own. We have a flexible work environment, and fluid career paths. We not only encourage but celebrate internal mobility. We also recognize the importance of purpose, well-being, and work-life balance. Within Empower and our communities, we work hard to create a welcoming and inclusive environment, and our associates dedicate thousands of hours to volunteering for causes that matter most to them.
Chart your own path and grow your career while helping more customers achieve financial freedom. Empower Yourself.
Responsibilities
~2 min read- →Own production data pipeline and platform reliability, resolving failures, delays, data quality issues, and performance problems.
- →Troubleshoot distributed data-processing workloads across Spark/EMR, ingestion pipelines, and data warehouses such as Redshift and Snowflake.
- →Lead or support incident response, perform root cause analysis, restore data safely, and implement lasting fixes.
- →Define and monitor data SLAs for freshness, latency, and completeness and improve monitoring, alerting, and observability.
- →Diagnose data issues including late-arriving data, incomplete loads, duplicate data, schema changes, and reconciliation failures.
- →Perform safe backfills, reprocessing, and production data recovery while preventing duplicate or inconsistent data.
- →Develop automation and tooling using Python and SQL to reduce operational toil, improve troubleshooting, and increase platform resilience.
- →Identify and address data-platform performance and cost issues.
- →Partner with Data Engineering teams to improve pipeline design, data quality, reliability, and production readiness.
- →Support disaster recovery, backup validation, and recovery workflows.
- →Create and maintain runbooks, SOPs, and operational documentation.
- →Participate in an on-call rotation for production data systems.
- →5+ years of hands-on experience building, operating, or supporting production data platforms, with significant recent experience in AWS environments.
- →Hands-on experience designing, building, deploying, and operating production-grade data pipelines.
- →Prior experience building data pipelines and seeing them through production, including exposure to real-world failures and operational challenges.
- →Strong hands-on Python and SQL experience in production data environments.
- →Hands-on experience troubleshooting distributed data-processing systems such as Spark/EMR.
- →Experience working with AWS data services such as EMR, S3, Glue, Redshift, DynamoDB, Lambda, or similar data services.
- →Experience operating or troubleshooting cloud data warehouses such as Redshift, Snowflake, or similar platforms.
- →Strong understanding of end-to-end production data architecture, including ingestion, processing, orchestration, storage/warehousing, data quality, monitoring, and downstream consumption.
- →Experience handling production incidents and performing root cause analysis.
- →Experience with data validation, reconciliation, backfills, reprocessing, and late or incomplete data.
- →Strong problem-solving skills and the ability to work through ambiguous production data issues.
- →Strong communication during incidents with both technical and non-technical stakeholders.
- Experience troubleshooting complex Spark/EMR production workloads.
- Experience improving observability and alerting specifically for data pipelines and data platforms.
- Experience with streaming or event-driven data systems such as Kafka, Kinesis, or CDC patterns.
- Experience with Snowflake performance, workload, reliability, or cost optimization.
- Experience designing automated controls for data freshness, completeness, reconciliation, and anomaly detection.
- Experience with disaster recovery, backup validation, and resiliency testing for data platforms.
- Experience developing automation or intelligent tooling to improve data-platform reliability, troubleshooting, performance, or cost efficiency.
- Participate in an on-call rotation for production data systems, with occasional change windows outside normal business hours to support safe releases and resiliency activities.
- This job operates in a professional office environment.
- Candidates should be prepared to discuss and demonstrate the projects, technologies, architecture decisions, and production experience represented on their résumé. Technical assessments are intended to evaluate the candidate's individual hands-on experience and must be completed independently.
- Final candidates may be subject to identity, employment, education, certification, reference, and background verification in accordance with company policy and applicable law
What We Offer
~3 min readWe offer an array of diverse and inclusive benefits regardless of where you are in your career. We believe that providing our employees with the means to lead healthy balanced lives results in the best possible work performance.
Location & Eligibility
Listing Details
- Posted
- September 29, 2026
- First seen
- September 29, 2026
- Last seen
- October 5, 2026
Posting Health
- Days active
- 5
- Repost count
- 0
- Trust Level
- 54%
- Scored at
- October 5, 2026
Signal breakdown
Similar Data Reliability Engineer jobs
View all →Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.