1h ago
New↻ Repost

Site Reliability Engineer II

United StatesUnited States·Dallasmid
EngineeringDevops Engineer
2 views0 saves0 applied

Quick Summary

Key Responsibilities

As a member of our global SRE team, you will take shared full-stack ownership of Onapsis production environments, balancing active operational response with modern software engineering principles.

Requirements Summary

Bachelor's or Master’s degree in Computer Science or related fields or equivalent experience. 2-4 years experience in SRE, DevOps,

Technical Tools
EngineeringDevops Engineer

About the Role

~1 min read

The world’s most critical--and at-risk--business applications have been neglected for far too long. Onapsis eliminates this blind spot by providing cybersecurity solutions dedicated to business-critical applications. Onapsis helps nearly 30% of the Forbes Global 100 understand the threats and risks across their SAP and Oracle landscapes, whether running on-premises, in the cloud, or in a hybrid environment. 

We are looking for a Site Reliability Engineer II to join our global engineering team. In this role, you will apply software engineering principles to operations, ensuring our cloud platform and distributed security products remain highly available, scalable, and resilient. You will be directly accountable for monitoring, inspecting, troubleshooting, and resolving service and product issues while continuously working with engineering partners to improve telemetry and related operation automations. 

Rather than performing manual operational maintenance, you will write code, build automation, and design observability frameworks that eliminate toil and prevent system failures. You will work side-by-side with product development teams to embed reliability into the software lifecycle from day one.

Responsibilities

~1 min read

As a member of our global SRE team, you will take shared full-stack ownership of Onapsis production environments, balancing active operational response with modern software engineering principles. You will develop a deep, end-to-end understanding of our system architecture, technical dependencies, and service behaviors to maximize the performance, scalability, and resilience of our enterprise security platforms.

Your time will be split between managing live production environments and driving engineering initiatives that ensure long-term system stability. When troubleshooting live incidents, you will diagnose complex issues across distributed cloud services and stateful infrastructure. Between operational cycles, you will shift into a software engineering mindset—designing, developing, and maintaining custom automation tooling to eliminate repetitive toil, optimize monitoring telemetry, and increase operational efficiency.

The ideal Site Reliability Engineer II brings a strong foundational toolkit spanning the software development lifecycle (SDLC), Linux systems administration, core networking protocols, and cloud computing (AWS preferred). You excel at viewing operational friction through a software lens, leveraging coding and automation to transform system vulnerabilities and outages into permanently solved engineering problems.

Requirements

~1 min read
  • Bachelor's or Master’s degree in Computer Science or related fields or equivalent experience.
  • 2-4 years experience in SRE, DevOps, or Cloud Engineering role supporting production distributed systems in cloud environments
  • Beginner knowledge of programming skills such as Java and Python
  • Intermediate Knowledge of the SRE Observability practices such as Application Performance Monitoring, Error Budget definition and tracking, and ensuring monitoring, logging and tracing are connected.
  • Understanding of SLI, SLO, SLA methodologies
  • Beginner knowledge of Infrastructure as Code and Configuration Management  using Terraform
  • Beginner knowledge of Containers and related orchestration platforms such as Kubernetes
  • Intermediate knowledge of orchestrating technical operations through scripting and automation using Bash or Python.
  • Experience using AI coding assistants for daily tasks, debugging, and rapid prototyping
  • Experience managing Linux OS (Debian/Ubuntu/OpenSuse), Queue servers (RabbitMQ), Database instances (PostgreSQL) and AWS cloud computing infrastructure such as EC2 instances or EKS
  • Intermediate knowledge in troubleshooting complex software and networking issues
  • Proven ability to quickly learn new technical domains and then train others
  • Great verbal and written communication skills
  • Ability to own end-to-end features, service integrations, database schema design, and operational telemetry. 
  • Secure software development best practices knowledge
  • Knowledge of test-driven development (TDD), CI / CD tooling, and Agile methodologies
  • Knowledge of professional software engineering practices & best practices for the full software development life cycle, including coding standards, code reviews, source control management, build processes, testing, and operations
  • Experience with distributed computing and enterprise-wide systems
  • Experience in communicating with users, other technical teams, and senior management to collect requirements, describe software product features, technical designs, and product strategy

What We Offer

~1 min read
✓A role in shaping the future of protecting the most critical applications that run the world's business and a career that grows as the company grows.
✓A unique culture of high achievement and teamwork.
✓Supportive and humble colleagues are some of the space's top problem solvers and innovators.
✓Financial security through competitive compensation and incentives.
✓A comprehensive benefits plan, including medical, dental, vision, disability, life insurance, and a 401K.
✓Flexible work options: OnaFlex PTO

Onapsis protects the business applications that run the global economy. The Onapsis Platform delivers vulnerability management, change assurance, and continuous compliance for business applications from leading vendors such as SAP, Oracle, and others. The Onapsis Platform is powered by the Onapsis Research Labs, the team responsible for the discovery and mitigation of more than 1,000 zero-day vulnerabilities in business applications.

Onapsis is headquartered in Boston, MA, with offices in Dallas, TX, Heidelberg, Germany, Bucharest, Romania, and Buenos Aires, Argentina, and proudly serves hundreds of the world’s leading brands, including close to 30% of the Forbes Global 100, six of the top 10 automotive companies, five of the top 10 chemical companies, four of the top 10 technology companies, and three of the top 10 oil and gas companies.

For more information, connect with Onapsis on LinkedIn or visit https://www.onapsis.com.

#LI-AC1

#LI-Hybrid

 

Location & Eligibility

Where is the job
Dallas, United States
On-site at the office
Who can apply
US

Listing Details

Posted
October 8, 2026
First seen
October 8, 2026
Last seen
October 8, 2026

Posting Health

Days active
0
Repost count
1
Trust Level
61%
Scored at
October 8, 2026

Signal breakdown

freshnesssource trustcontent trustemployer trust

Onapsis open source projects

Employees
350
Founded
2009
View company profile
Newsletter

Stay ahead of the market

Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.

A
B
C
D
Join 12,000+ marketers

No spam. Unsubscribe at any time.

Site Reliability Engineer II