Senior Engineer Site Reliability - Data Operations
Quick Summary
Own and improve the reliability, availability, performance, and operational health of production data platforms and pipelines. Monitor and support pipeline execution, dependencies, failures, delays,
Our vision for the future is based on the idea that transforming financial lives starts by giving our people the freedom to transform their own. We have a flexible work environment, and fluid career paths. We not only encourage but celebrate internal mobility. We also recognize the importance of purpose, well-being, and work-life balance. Within Empower and our communities, we work hard to create a welcoming and inclusive environment, and our associates dedicate thousands of hours to volunteering for causes that matter most to them.
Chart your own path and grow your career while helping more customers achieve financial freedom. Empower Yourself.
The Senior Site Reliability Engineer - Data Platforms will improve the reliability, observability, and operational maturity of cloud-based data platforms and production data pipelines. This SRE-first role focuses on production reliability and operational ownership, with primary data-pipeline development and core infrastructure engineering handled in partnership with the respective engineering teams. You will work closely with Data Engineering, Cloud, Platform, Application, Security, and FinOps teams to improve production health, incident response, automation, observability, resilience, and operational efficiency.
Responsibilities
~2 min read- →Own and improve the reliability, availability, performance, and operational health of production data platforms and pipelines.
- →Monitor and support pipeline execution, dependencies, failures, delays, recovery, and downstream impact.
- →Troubleshoot complex production issues across AWS, data platforms, pipelines, and supporting services.
- →Design and improve observability using enterprise platforms such as Datadog and Splunk, including dashboards, alerting, logging, and service-health indicators.
- →Improve signal quality, reduce alert noise, and strengthen incident detection and diagnosis.
- →Lead or participate in production incident response, service restoration, root-cause analysis, and corrective actions.
- →Use Python to automate operational tasks, monitoring, health checks, remediation, and repetitive support activities.
- →Apply SRE practices including service-level objectives, incident management, operational readiness, capacity planning, resilience, disaster recovery, and reduction of operational toil.
- →Use SQL for production troubleshooting, validation, and investigation.
- →Understand and troubleshoot infrastructure managed through Terraform and Infrastructure as Code, partnering with Cloud or Platform Engineering when deeper infrastructure changes are required.
- →Identify recurring operational issues and drive sustainable engineering improvements.
- →Contribute to performance, capacity, and cost-efficiency improvements across AWS and data-platform services.
- →Develop and improve operational runbooks and recovery procedures.
- →Participate in the PagerDuty on-call rotation, including occasional weekend coverage.
- →Understand how the platform and its pipelines behave in production, recognize emerging reliability risks, diagnose issues quickly, restore service effectively, and implement lasting improvements.
- →Use observability and automation to move operations from reactive support toward proactive reliability engineering, reducing recurring incidents and manual operational effort while improving the overall resilience of the data platform.
- →5+ years of hands-on AWS experience supporting production environments.
- →Demonstrated experience in Site Reliability Engineering, Production Engineering, Platform Engineering, or Cloud Reliability Engineering.
- →Strong practical understanding and implementation of SRE principles and production operations.
- →Experience supporting Amazon Redshift or similar enterprise data platforms in production.
- →Experience owning or supporting the reliability and operations of production data pipelines or data-intensive services.
- →Hands-on experience with enterprise observability platforms such as Datadog and Splunk.
- →Strong Python programming skills for automation and operational tooling.
- →Working knowledge of SQL for production investigation and troubleshooting.
- →Strong understanding of Terraform and Infrastructure as Code, with the ability to read, review, and troubleshoot existing IaC.
- →Experience with incident management, root-cause analysis, and implementing preventive and corrective actions.
- →Strong troubleshooting and problem-solving skills in complex production environments.
- Hands-on production experience with Snowflake.
- Experience supporting multiple cloud-based data platforms.
- Experience designing observability for large-scale, distributed, or data-intensive systems.
- Experience with AIOps, intelligent automation, or agentic operations applied to production operations.
- Experience using AI-assisted approaches for incident detection, diagnosis, alert enrichment, runbook execution, operational automation, or guarded remediation.
- Familiarity with AWS-native monitoring services such as Amazon CloudWatch.
- Experience with data-pipeline orchestration, dependency management, failure recovery, and production support.
- Experience improving cloud-resource utilization and cost visibility.
- Experience helping establish or mature SRE practices within an engineering organization.
This job operates in a professional office environment.
This job description is not intended to be an exhaustive list of all duties, responsibilities and qualifications of the job. The employer has the right to revise this job description at any time. You will be evaluated in part based on your performance of the responsibilities and/or tasks listed in this job description. You may be required to perform other duties that are not included on this job description. The job description is not a contract for employment, and either you or the employer may terminate employment at any time, for any reason, as per terms and conditions of your employment contract.
What We Offer
~3 min readWe offer an array of diverse and inclusive benefits regardless of where you are in your career. We believe that providing our employees with the means to lead healthy balanced lives results in the best possible work performance.
Location & Eligibility
Listing Details
- Posted
- October 5, 2026
- First seen
- October 5, 2026
- Last seen
- October 7, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 65%
- Scored at
- October 6, 2026
Signal breakdown
Similar Site Reliability jobs
View all →Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.