3h ago
New
USD 105000-140000/yr

Senior Site-Reliability Engineer

United StatesUnited StatesRemoteFull-timesenior
EngineeringDevops Engineer
1 views0 saves0 applied

Quick Summary

Requirements Summary

Bring at least 5 years of experience in Systems Administration, DevOps, Site-Reliability Engineering, or a closely related infrastructure role.

Technical Tools
EngineeringDevops Engineer

This is a senior infrastructure engineering role focused on the reliability, availability, and performance of Windows-based production environments. You will bridge development and operations to build highly available services and maintain strong operational standards across hybrid infrastructure. The role combines infrastructure automation, configuration management, observability, incident response, and continuous improvement. You will work closely with development teams to strengthen deployment and release processes while establishing measurable reliability objectives. The position offers an opportunity to work with modern cloud, automation, and monitoring technologies in a mission-critical environment. The role is remote and requires the ability to obtain the appropriate security clearance.

  • Design, implement, and maintain scalable infrastructure using Infrastructure as Code (IaC) practices across production environments.

  • Develop and maintain automation scripts using PowerShell, Python, Ruby, and other scripting languages for operating system provisioning, configuration management, and recurring operational tasks.

  • Implement and manage configuration management solutions such as Terraform, Puppet, and/or Chef across hybrid infrastructure environments.

  • Monitor system health, performance, reliability, and availability using established observability tools and operational best practices.

  • Establish, maintain, and enforce Service Level Agreements (SLAs), Service Level Objectives (SLOs), and error budgets for production services.

  • Participate in an on-call rotation, respond to production incidents, restore services rapidly, and conduct thorough root cause analysis.

  • Partner with development teams to improve deployment pipelines, release processes, and overall service reliability.

  • Create and maintain operational documentation, including procedures, runbooks, and architectural decisions.

  • Lead or contribute to post-mortem reviews and implement corrective actions designed to prevent recurring incidents.

  • Troubleshoot complex infrastructure and application issues across multiple technology layers while maintaining a strong focus on operational excellence.

Requirements

~2 min read
  • Bring at least 5 years of experience in Systems Administration, DevOps, Site-Reliability Engineering, or a closely related infrastructure role.

  • Have strong hands-on expertise with Windows Server environments, including Windows Server 2016 or later, Active Directory, IIS, and Microsoft SQL.

  • Demonstrate strong cloud infrastructure skills, with AWS experience preferred.

  • Possess advanced scripting capabilities, including the development of reusable modules and integrations with REST APIs.

  • Have hands-on experience with Terraform for infrastructure provisioning and Puppet or Chef for configuration management.

  • Be experienced with monitoring and observability platforms such as Prometheus, Grafana, Datadog, or New Relic.

  • Have a solid understanding of networking fundamentals, including DNS, TCP/IP, load balancing, and VPN technologies.

  • Demonstrate strong analytical and problem-solving skills, with the ability to troubleshoot complex issues spanning multiple technology layers.

  • A Bachelor's degree in Computer Science, Information Technology, or a related discipline is preferred, although equivalent professional experience may be considered.

  • Relevant certifications such as AWS Solutions Architect, Microsoft certifications, or HashiCorp Certified: Terraform Associate are desirable.

  • Experience with containerization technologies such as Docker and Kubernetes, as well as CI/CD tools including GitLab Pipelines, Jenkins, or GitHub Actions, is a plus.

  • Knowledge of security best practices, compliance frameworks, and log aggregation and analysis tools such as the ELK Stack or Splunk is desirable.

  • Be able to obtain the required security clearance for the position.

What We Offer

~1 min read
✓Annual salary range of $105,000–$140,000, with actual compensation determined by experience, qualifications, skills, geographic location, contract requirements, and business needs.
✓Medical, dental, and vision insurance for eligible employees.
✓Life, AD&D, and disability insurance.
✓Paid time off and 11 company holidays.
✓401(k) retirement plan with company matching.
✓Additional employee benefits and wellness resources, subject to applicable eligibility requirements and plan terms.
✓Remote work arrangement based in the United States.

Location & Eligibility

Where is the job
United States
Remote within one country
Who can apply
US

Listing Details

Posted
October 8, 2026
First seen
October 8, 2026
Last seen
October 8, 2026

Posting Health

Days active
0
Repost count
0
Trust Level
80%
Scored at
October 8, 2026

Signal breakdown

freshnesssource trustcontent trustemployer trust
Newsletter

Stay ahead of the market

Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.

A
B
C
D
Join 12,000+ marketers

No spam. Unsubscribe at any time.

Senior Site-Reliability EngineerUSD 105000-140000