SRE Platform Engineer
Quick Summary
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a SRE Platform Engineer based in United States.
This role focuses on ensuring the reliability, performance, and availability of mission-critical AI and data platforms supporting federal oversight operations. You will serve as an operational guardian for AI assistants, enterprise data platforms, and AI-powered applications used in demanding environments. The position combines observability, incident response, performance engineering, capacity planning, disaster recovery, and cloud operations. Working primarily within secure Azure Government environments, you will help build resilient and highly available platforms. You will collaborate across application, data, infrastructure, and operational teams to troubleshoot complex issues and improve service quality. This is an opportunity to apply advanced SRE practices to large-scale technology supporting important government missions.
- Monitor, maintain, and support production and non-production environments, ensuring availability, performance, service health, and adherence to service level objectives.
- Implement comprehensive observability through alerting, dashboards, health checks, synthetic monitoring, and log analysis using tools such as Azure Monitor, Application Insights, and Log Analytics.
- Lead incident response and troubleshooting, including root cause analysis, defect resolution, dependency updates, integration validation, and emergency change coordination.
- Analyze application, API, AI model, data pipeline, and infrastructure metrics to identify bottlenecks, latency, resource constraints, and opportunities for improved efficiency.
- Support capacity planning, resource sizing, autoscaling, and cost optimization across compute, storage, and AI model consumption.
- Maintain backup and restore processes, disaster recovery procedures, high-availability architectures, and business continuity capabilities.
- Monitor data platforms, including Azure Databricks clusters, data pipelines, storage services, and analytical workloads, with appropriate alerting for failures and performance degradation.
- Support the operational readiness and deployment of new applications and capabilities through pre-production validation, performance testing, runbook creation, and go-live coordination.
- Provide specialized troubleshooting and surge support for complex technical issues, large-scale data collection and analysis, and analytical environment optimization.
- Develop and maintain runbooks, troubleshooting guides, architecture diagrams, incident post-mortems, and knowledge-transfer documentation to support sustainable operations.
Requirements
~1 min read- Bachelor’s degree plus 15 years of relevant experience in SRE, DevOps, platform engineering, systems administration, or a related field; equivalent combinations include a Master’s degree plus 12 years, 21 years without a degree, or an Associate degree plus 17 years.
- Ability to obtain an active DHS/EOD clearance as required.
- Extensive knowledge of SRE principles, including monitoring, observability, incident response, capacity planning, performance optimization, and reliability engineering.
- Strong expertise with Azure cloud services covering compute, storage, networking, monitoring, and PaaS offerings, along with strong operational best practices.
- Hands-on experience with observability and monitoring technologies such as Azure Monitor, Application Insights, Grafana, Prometheus, and ELK.
- Proven ability to troubleshoot complex issues across application, platform, and infrastructure layers, supported by strong analytical and problem-solving skills.
- Experience operating AI/ML platforms, Azure OpenAI or other large language model services, Databricks, Synapse, or high-scale cloud applications is highly desirable.
- Experience with Azure Government or other secure government cloud environments, such as AWS GovCloud, including compliance monitoring and security operations, is an advantage.
- Background supporting federal government, mission-critical, or 24/7 operational environments, including incident response, change management, and operational excellence programs, is preferred.
What We Offer
~2 min readLocation & Eligibility
Listing Details
- Posted
- September 24, 2026
- First seen
- September 27, 2026
- Last seen
- September 29, 2026
Posting Health
- Days active
- 1
- Repost count
- 0
- Trust Level
- 68%
- Scored at
- September 29, 2026
Signal breakdown
Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.