Intermediate Site Reliability Engineer
Quick Summary
Hey there! We're ContactMonkey ๐ Our mission is to power measurable employee engagement worldwide, and we're looking for an Intermediate Site Reliability Engineer to join our Engineering team.
Hey there! We're ContactMonkey ๐
Our mission is to power measurable employee engagement worldwide, and we're looking for an Intermediate Site Reliability Engineer to join our Engineering team.
About the Role
~1 min readWe're looking for someone who enjoys running production systems, understands how applications work, and brings practical application security experience.
You'll work closely with our SRE and development teams to maintain our infrastructure, improve deployments, investigate production issues, and address security risks. You'll also contribute to the technical controls that support SOC 2 audits and GDPR compliance.
Our environment includes AWS, Kubernetes on EKS, Terraform, Terragrunt, GitHub Actions, Prometheus, Grafana, and CloudWatch. Our applications use Ruby on Rails, Vue.js, and Node.js, with MySQL, PostgreSQL, and Sidekiq supporting the backend.
You'll take ownership of defined projects and operational improvements, with senior engineers available for guidance and review. There's room to develop deeper expertise in reliability, infrastructure automation, and application security.
- Infrastructure & reliability: Maintain AWS and Kubernetes environments, troubleshoot production issues, and improve availability, performance, and resource usage.
- Terraform & Terragrunt: Build and maintain infrastructure as code, review plans, manage environment configuration, and address infrastructure drift.
- Deployments & developer experience: Improve CI/CD pipelines, deployment automation, release checks, and rollback procedures.
- Monitoring & incidents: Improve monitoring and alerts, join the on-call rotation, and contribute to incident reviews.
- Application security: Work with developers to assess vulnerabilities, review security risks, and validate fixes.
- CI/CD security: Maintain code, dependency, secrets, container, and infrastructure scanning.
- Cloud security: Strengthen IAM, secrets management, network controls, and Kubernetes security.
- SOC 2 & GDPR: Support technical controls, audit evidence, and personal data protection.
- Recovery: Test backups and recovery procedures, maintain runbooks, and support production readiness.
- Collaboration: Participate in code reviews, document changes, and support application, AI, and data engineering teams.
- Around 3โ5 years of experience in SRE, DevOps, platform engineering, cloud operations, or a related engineering role. Equivalent practical experience is welcome.
- Hands-on experience supporting production workloads in AWS.
- Experience writing and maintaining Terraform modules and Terragrunt configuration, including reviewing plans and working with remote state.
- Experience with Docker and Kubernetes, including troubleshooting deployments, services, health checks, and resource limits.
- A solid understanding of Linux, networking, DNS, HTTP, and TLS.
- Ability to write automation in Python, Bash, Ruby, JavaScript, or another suitable language.
- Experience with Git, pull requests, CI/CD pipelines, and deployment workflows.
- Experience using logs, metrics, dashboards, and alerts to investigate production problems.
- Practical application security experience through vulnerability remediation, secure code review, threat modelling, or security tooling.
- Understanding of common web application risks, including broken access control, injection, authentication weaknesses, and sensitive data exposure.
- Familiarity with IAM, least privilege, secrets management, and encryption.
- Working knowledge of SOC 2 controls and GDPR principles relevant to engineering, including access restrictions, data minimization, retention, and deletion.
- Clear communication skills and good judgment about when to work independently, request a review, or escalate an issue.
- Supported Ruby on Rails or Node.js applications in production
- Worked with MySQL, PostgreSQL, Redis or Valkey, and Sidekiq
- Experience with GitHub Actions, Argo CD, Helm, Karpenter, or KEDA
- Built useful monitoring with Prometheus, Grafana, or CloudWatch
- Supported services across multiple AWS regions
- Helped remediate penetration-test findings or contributed technical evidence to a SOC 2 audit
- Participated in backup restoration or disaster recovery exercises.
- Familiar with securing AI integrations, agent workloads, or MCP services
- You hold relevant AWS, Kubernetes, Terraform, or security certifications.
What We Offer
~1 min readThe salary range for this role is $130,000-$150,000 Compensation is based on experience, skills, and our internal compensation framework and equity.
We're happy to discuss compensation throughout the hiring process.
This is a full-time position on our SRE team.
The role includes a shared on-call rotation after onboarding. We'll discuss the schedule, escalation support, and expectations during the interview process.
ContactMonkey helps organizations create, send, and measure internal communications directly within Outlook and Gmail.
Our platform brings together email design, employee engagement tools, and analytics so internal communications teams can understand what reaches their people and what gets a response.
As the product grows, we're investing in reliability, security, and tooling that helps our engineering teams deliver changes confidently.
Diversity is our strength
At ContactMonkey, we're building products for diverse organizations, and we need a diverse team to do that. We strongly encourage applications from everyone regardless of race, religion, colour, national origin, gender, sexual orientation, age, marital status, or disability status.
We are committed to an accessible hiring process. If you need accommodations or adjustments during interviews or beyond, please let us know so we can arrange the support you need.
AI Disclosure
We use AI to take notes during our interviews. Applications and interviews are reviewed by our Talent Acquisition team. Our applicant tracking system uses AI for workflows and hiring process efficiencies.
Location & Eligibility
Listing Details
- Posted
- October 8, 2026
- First seen
- October 8, 2026
- Last seen
- October 8, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 67%
- Scored at
- October 8, 2026
Signal breakdown
4 other jobs at
View all โBrowse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.
