Senior Site Reliability Engineer
Quick Summary
About Us GridCARE is a leading venture-backed startup solving the most critical constraint in AI’s growth trajectory: immediate access to power. As demand for computing skyrockets,
GridCARE is a leading venture-backed startup solving the most critical constraint in AI’s growth trajectory: immediate access to power. As demand for computing skyrockets, access to energy has become the defining bottleneck in the AI infrastructure race. While leading tech companies invest billions in speculative, long-term solutions that may take decades to arrive, GridCARE’s pioneering physics-based generative AI platform unlocks gigawatts of hidden capacity in today’s electric grid — enabling hyperscalers, data center developers, and utilities to power AI infrastructure years sooner than conventional approaches and without costly upgrades.
Founded at Stanford’s Doerr School of Sustainability and backed by leading climate-tech and deep-tech investors, GridCARE has assembled a world-class team spanning power systems, AI, and infrastructure.
At GridCARE, you will:
⚡ Work at the intersection of AI, energy, and infrastructure — the foundation of the next industrial revolution.
🤝 Partner with hyperscalers, developers, and utilities on high-impact, real-world deployments.
🌎 Help shape a more abundant, efficient, and resilient energy future for the digital era.
🚀 Join a company defining a new category — capacity acceleration for AI.
💰 Receive competitive compensation, equity, and benefits in a fast-growth, mission-driven environment.
Learn more about GridCARE:
We're looking for a Senior SRE to own the reliability, scalability, and observability of our production systems. You'll work closely with platform and data engineering to keep high-throughput, data-intensive services running at the availability our customers (utilities, data center operators) require.
Responsibilities
~1 min read- →
Design and operate infrastructure on AWS using Terraform and Kubernetes
- →
Build monitoring, alerting, and observability (Prometheus, Grafana, Datadog, or similar) with meaningful SLOs/SLIs
- →
Automate away toil — deployment pipelines, capacity management, self-healing systems
- →
Partner with engineering on architecture reviews to catch reliability and scalability risks before they ship
- →
Manage database and data pipeline reliability for large-scale, real-time grid data processing
- →
Drive security and compliance best practices across infrastructure
Requirements
~1 min read5+ years in SRE, DevOps, or infrastructure engineering roles
Nice to Have
~1 min readExperience with data-intensive or real-time processing systems
Background in energy, climate tech, or critical infrastructure
Experience scaling infrastructure through hypergrowth
On-Prem Kubernetes Deployment Experience
Windows Server Administration Experience
Track record of running on-call for production systems and leading incident response
Experience with CI/CD pipelines (Github Actions) and infrastructure automation
Experience with Gitops concepts and tooling (ArgoCD/Flux)
Solid understanding of networking, distributed systems, and database reliability
Comfortable operating in a fast-moving startup environment with ambiguity
What We Offer
~1 min read$180,000-$230,000 Total
Location & Eligibility
Listing Details
- Posted
- September 14, 2026
- First seen
- September 25, 2026
- Last seen
- September 28, 2026
Posting Health
- Days active
- 2
- Repost count
- 0
- Trust Level
- 28%
- Scored at
- September 28, 2026
Signal breakdown
Similar Devops Engineer jobs
View all →Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.