Senior Site-Reliability Engineer
Quick Summary
Bring at least 5 years of experience in Systems Administration, DevOps, Site-Reliability Engineering, or a closely related infrastructure role.
This is a senior infrastructure engineering role focused on the reliability, availability, and performance of Windows-based production environments. You will bridge development and operations to build highly available services and maintain strong operational standards across hybrid infrastructure. The role combines infrastructure automation, configuration management, observability, incident response, and continuous improvement. You will work closely with development teams to strengthen deployment and release processes while establishing measurable reliability objectives. The position offers an opportunity to work with modern cloud, automation, and monitoring technologies in a mission-critical environment. The role is remote and requires the ability to obtain the appropriate security clearance.
-
Design, implement, and maintain scalable infrastructure using Infrastructure as Code (IaC) practices across production environments.
-
Develop and maintain automation scripts using PowerShell, Python, Ruby, and other scripting languages for operating system provisioning, configuration management, and recurring operational tasks.
-
Implement and manage configuration management solutions such as Terraform, Puppet, and/or Chef across hybrid infrastructure environments.
-
Monitor system health, performance, reliability, and availability using established observability tools and operational best practices.
-
Establish, maintain, and enforce Service Level Agreements (SLAs), Service Level Objectives (SLOs), and error budgets for production services.
-
Participate in an on-call rotation, respond to production incidents, restore services rapidly, and conduct thorough root cause analysis.
-
Partner with development teams to improve deployment pipelines, release processes, and overall service reliability.
-
Create and maintain operational documentation, including procedures, runbooks, and architectural decisions.
-
Lead or contribute to post-mortem reviews and implement corrective actions designed to prevent recurring incidents.
-
Troubleshoot complex infrastructure and application issues across multiple technology layers while maintaining a strong focus on operational excellence.
Requirements
~2 min read-
Bring at least 5 years of experience in Systems Administration, DevOps, Site-Reliability Engineering, or a closely related infrastructure role.
-
Have strong hands-on expertise with Windows Server environments, including Windows Server 2016 or later, Active Directory, IIS, and Microsoft SQL.
-
Demonstrate strong cloud infrastructure skills, with AWS experience preferred.
-
Possess advanced scripting capabilities, including the development of reusable modules and integrations with REST APIs.
-
Have hands-on experience with Terraform for infrastructure provisioning and Puppet or Chef for configuration management.
-
Be experienced with monitoring and observability platforms such as Prometheus, Grafana, Datadog, or New Relic.
-
Have a solid understanding of networking fundamentals, including DNS, TCP/IP, load balancing, and VPN technologies.
-
Demonstrate strong analytical and problem-solving skills, with the ability to troubleshoot complex issues spanning multiple technology layers.
-
A Bachelor's degree in Computer Science, Information Technology, or a related discipline is preferred, although equivalent professional experience may be considered.
-
Relevant certifications such as AWS Solutions Architect, Microsoft certifications, or HashiCorp Certified: Terraform Associate are desirable.
-
Experience with containerization technologies such as Docker and Kubernetes, as well as CI/CD tools including GitLab Pipelines, Jenkins, or GitHub Actions, is a plus.
-
Knowledge of security best practices, compliance frameworks, and log aggregation and analysis tools such as the ELK Stack or Splunk is desirable.
-
Be able to obtain the required security clearance for the position.
What We Offer
~1 min readLocation & Eligibility
Listing Details
- Posted
- October 8, 2026
- First seen
- October 8, 2026
- Last seen
- October 8, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 80%
- Scored at
- October 8, 2026
Signal breakdown
Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.