Senior Infrastructure SRE
Quick Summary
auto-remediation, self-healing systems,
About the Role
~1 min readPointClickCare builds cloud platforms that power safer, more connected care for millions of patients. You'll design and operate resilient infrastructure services across Azure, AWS, GCP, and on-prem environments, driving reliability, automation, and operational excellence for identity, compute, storage, messaging, and shared services with an SRE-first mindset.
Responsibilities
~1 min read- →Design and implement highly available infrastructure solutions for compute, storage, identity, messaging, and shared services
- →Build and maintain Infrastructure as Code using Terraform or Pulumi; establish best practices and standards
- →Automate operational workflows to eliminate toil: auto-remediation, self-healing systems, capacity planning
- →Define and track SLIs and SLOs for critical services; manage error budgets and reliability targets
- →Participate in the on-call rotation and lead incident response for complex infrastructure issues; conduct blameless post-mortems and drive systemic fixes
- →Develop observability strategies: metrics, logs, distributed tracing, alerting frameworks
- →Apply AI-assisted tooling to reduce toil and speed up investigation: log analysis, alert triage, runbook and post-mortem drafting, automation scaffolding
- →Own reliability for one or more infrastructure domains end to end, driving multi-team initiatives with product engineering from problem definition through adoption
- →Mentor intermediate SREs; review infrastructure changes; establish operational best practices
- 5+ years of hands-on experience operating and designing cloud infrastructure
- Expert-level understanding of the core services of either Azure or AWS
- Working proficiency in at least one additional platform (Azure, AWS, or GCP)
- Experience designing and supporting production infrastructure that spans multiple cloud platforms
- 3+ years of production experience with Infrastructure as Code (Terraform, Pulumi, CloudFormation)
- Ability to design scalable, reusable IaC modules and enforce GitOps workflows
- Experience managing IaC across multiple cloud providers, including module design, state layout, and provider-specific resource differences
- Strong proficiency in at least one programming language (Python, Go, Bash) for production automation
- Demonstrated ability to write tested, maintainable automation and tooling
- Practical application of SRE principles in production environments
- Experience defining and managing SLIs/SLOs, error budgets, toil metrics
- Track record of improving system reliability (e.g., MTTR reduction, availability improvements)
- Demonstrated depth across key infrastructure services:
- Expert-level experience running Kubernetes and containerized workloads in production on managed Kubernetes (AKS, EKS, or equivalent), plus VM-based compute
- Practical experience operating a service mesh in production (Istio preferred; Linkerd or equivalent)
- Strong proficiency with enterprise identity and SSO: SAML, OAuth/OIDC, LDAP, and cloud IAM
- Working knowledge of an enterprise federation platform such as PingFederate, Entra ID, Okta, or ADFS
- Strong proficiency in storage solutions (object, block, and file storage; Kubernetes persistent volumes)
- Working knowledge of designing and operating infrastructure in a regulated environment (HIPAA, SOC 2, PCI, FedRAMP, or equivalent)
- Familiarity with audit evidence, access controls, encryption in transit and at rest, and data residency constraints
- Proven track record of measurably reducing operational toil through automation (e.g. ticket volume, manual runbook executions, hours reclaimed)
- Strong communication and documentation skills; demonstrated ability to influence engineering teams
- 2+ years in healthcare technology or highly regulated SaaS environments (HIPAA, SOC 2, HITRUST)
- Cloud certifications: Azure Solutions Architect Expert, AWS Solutions Architect Professional, GCP Professional Cloud Architect, or equivalent
- Working experience operating Kubernetes at scale across multiple clusters (CKA or CKAD certification a plus)
- Working knowledge of messaging and event-streaming platforms (Kafka, Azure Service Bus, Event Hubs, SQS, Pub/Sub)
- Practical experience building CI/CD pipelines and deployment automation (GitLab, GitHub Actions, ArgoCD)
- Familiarity with AI-assisted engineering and operations tooling: LLM-based coding assistants, AIOps, agentic incident investigation
- Judgment about where AI belongs in an operational workflow, including safe handling of sensitive data in prompts and reviewing generated changes before they reach production contribution to open-source SRE tools or infrastructure projects (GitHub profile, PRs merged)
- Bachelor's degree in Computer Science, Computer Engineering, Information Technology, or related technical field
- OR equivalent practical experience with a proven track record in infrastructure and SRE practices evidence of continuous learning and staying current with SRE and cloud-native trends
- Transparent collaboration: We work in the open using OKRs, cross-functional retrospectives, and public roadmaps so everyone knows priorities and progress
- Blameless culture: We confront problems courageously through structured post-mortems and root cause analyses, focusing on systems improvement not individual blame
- Data-driven decisions: We use metrics, APM, and observability data to make evidence-based choices about reliability and performance investments
- Continuous learning: We learn from incidents through retrospectives and RCAs, sharing knowledge across teams to prevent repeat issues
- Outcome accountability: We're accountable for customer and business results, not just completing tasks, measuring success by impact
- Thoughtful experimentation: We hold strong opinions loosely, testing concepts and running small experiments before scaling solutions
- Iterative delivery: We think big but act small, using Scrum, delivery plans, and frequent milestones to ship incrementally and learn fast
- Inclusive environment: We create space to listen and learn, actively growing our Ally Community to support equity, belonging, and career growth for all
Location & Eligibility
Listing Details
- Posted
- September 11, 2026
- First seen
- September 11, 2026
- Last seen
- September 11, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 81%
- Scored at
- September 11, 2026
Signal breakdown

PointClickCare is the leading electronic health record (EHR) technology partner to North America’s long-term post-acute care (LTPAC) and senior care industry.
View company profilePlease let Pointclickcare know you found this job on Jobera.
3 other jobs at Pointclickcare
View all →Explore open roles at Pointclickcare.
Similar Infrastructure jobs
View all →Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.