Site Reliability Engineer
Quick Summary
Reliability Engineering & Platform Operations Monitor and maintain production cloud environments to ensure platform availability, performance, scalability, and reliability. Build, configure,
Bachelor's degree in Computer Science, Computer Engineering, or a related technical field. Minimum 3 years of experience in Site Reliability Engineering (SRE), Cloud Operations, DevOps,
About the Role
~1 min readAccela provides cutting-edge technology that enables government agencies to engage and serve their communities. The foundation of our technology is the Accela Civic Platform, a cloud-based platform that supports a broad ecosystem of solutions for permitting, licensing, planning, and public sector operations.
As a Site Reliability Engineer (SRE), you will be responsible for maintaining, monitoring, and improving the reliability, availability, performance, and scalability of Accela's cloud platform and SaaS services. You will work across cloud infrastructure, applications, and operational processes to ensure a seamless customer experience through proactive monitoring, incident response, automation, and continuous improvement.
This role combines cloud operations, production support, and reliability engineering, with a strong emphasis on hands-on administration of Microsoft Azure services, observability platforms, incident management, and operational excellence.
Responsibilities
~2 min readReliability Engineering & Platform Operations
- →Monitor and maintain production cloud environments to ensure platform availability, performance, scalability, and reliability.
- →Build, configure, and optimize monitoring, logging, and alerting capabilities using Datadog.
- →Develop and maintain dashboards that provide real-time visibility into platform health, resource utilization, and service performance.
- →Configure and tune alerts to proactively identify and address service degradation or operational issues.
- →Support and optimize Azure-based infrastructure components, including Azure Kubernetes Service (AKS), Azure SQL, Azure Storage, and Azure Front Door.
- →Execute production releases, operational changes, and platform maintenance activities.
Incident Management & Continuous Improvement
- →Participate in and lead incident response activities to restore service and minimize customer impact.
- →Diagnose and resolve production incidents across infrastructure, platform, and application components.
- →Conduct root cause analysis (RCA) and implement corrective and preventative actions to reduce recurring issues.
- →Develop, maintain, and continuously improve operational documentation, runbooks, and incident response procedures.
- →Identify and implement automation opportunities that improve operational efficiency and system reliability.
Customer Support & Operational Excellence
- →Provide Level 3 (L3) support for customer-reported incidents and service requests.
- →Collaborate with Engineering, Professional Services, and Operations teams to investigate and resolve complex technical issues.
- →Meet established service level agreements (SLAs) for ticket response and resolution.
- →Support customer provisioning, environment maintenance, platform upgrades, and cloud migration activities as needed.
Data & Environment Operations
- →Support production-to-non-production data refresh activities.
- →Assist with database-related operational tasks, troubleshooting, and data movement activities.
- →Validate completion of operational activities and maintain appropriate documentation and records.
After-Hours Support
- →Support critical production activities outside normal business hours when required, including customer go-lives, cloud migrations, and other high-impact operational events.
Requirements
~1 min read- Bachelor's degree in Computer Science, Computer Engineering, or a related technical field.
- Minimum 3 years of experience in Site Reliability Engineering (SRE), Cloud Operations, DevOps, Infrastructure Operations, or a related role.
- Minimum 3 years of hands-on experience supporting production cloud environments.
- Experience administering and supporting Microsoft Azure cloud services.
- Hands-on experience with:
- Azure Kubernetes Service (AKS)
- Azure SQL
- Azure Storage
- Azure Front Door
- Experience implementing and supporting monitoring, logging, and observability solutions using Datadog.
- Experience leading incident response activities, performing root cause analysis, and developing operational runbooks.
- Strong understanding of infrastructure and application performance monitoring, including compute, memory, storage, networking, and availability metrics.
- Experience supporting SaaS platforms in production environments.
- Strong troubleshooting, analytical, and problem-solving skills.
- Excellent written and verbal communication skills with the ability to collaborate effectively across technical and non-technical teams.
- Experience with one or more programming or scripting languages such as Python, JavaScript/Node.js, Java, PowerShell, or .NET.
- Experience with Infrastructure as Code (IaC) tools such as Terraform.
- Experience with Kubernetes administration and containerized workloads.
- Experience with CI/CD pipelines and release automation.
- Familiarity with database operations, ETL processes, and data movement activities.
- Experience supporting local government technology solutions.
- Experience with the Accela Civic Platform.
What We Offer
~1 min read#LI-Remote
Location & Eligibility
Listing Details
- Posted
- October 6, 2026
- First seen
- October 6, 2026
- Last seen
- October 6, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 76%
- Scored at
- October 6, 2026
Signal breakdown
Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.