Site Reliability Engineer II (AI Platform)
Quick Summary
This hybrid role requires working in the office two days per week. With millions of diners, 70,000+ restaurant partners and 25+ years of experience, OpenTable, part of Booking Holdings, Inc.
This hybrid role requires working in the office two days per week.
With millions of diners, 70,000+ restaurant partners and 25+ years of experience, OpenTable, part of Booking Holdings, Inc. (NASDAQ: BKNG), is an industry leader with a passion for helping restaurants thrive. Our world-class technology empowers restaurants to focus on what matters most – their team, their guests, and their bottom line – while enabling diners to discover and book the perfect restaurant for every occasion.
Every employee at OpenTable has a tangible impact on what we do and how we do it. You’ll also be part of a global team and its portfolio of metasearch brands. Hospitality is all about taking care of others, and it defines our culture.
About the Role
~1 min readAs a Site Reliability Engineer II on the Serving Platforms team within Infrastructure Engineering, you will design, automate, and manage the core container stack and infrastructure powering our global business applications. Operating in a high-scale, self-hosted environment, you will serve as a subject matter expert for Kubernetes, Linux systems, and cloud-native automation, directly driving the reliability, security, and efficiency of our platform. In this role, you will collaborate with cross-functional engineering teams worldwide, lead greenfield infrastructure projects, resolve complex incidents, and build self-service capabilities that empower application developers across the organization.
Responsibilities
~1 min read- →Maintain, tune, and ensure high availability for the low-level Linux operating system and Kubernetes control plane across our self-hosted bare-metal infrastructure.
- →Architect, build, and maintain scalable container management, configuration management, and automation tools across global environments.
- →Investigate, resolve, and conduct root-cause analysis for complex infrastructure disruptions and performance bottlenecks at the system call level.
- →Participate in high-impact platform engineering projects and collaborate with globally distributed engineering teams to drive infrastructure standardization.
- →Participate in the team's on-call rotation to support critical production systems and ensure operational resilience.
- →Develop and maintain self-service tools, automation pipelines, and robust infrastructure monitoring to eliminate manual operational overhead.
Requirements
~1 min read- 5+ years of hands-on Linux experience (e.g., Ubuntu, CentOS) with expertise in kernel tuning (sysctl), process management (cgroups/namespaces), system calls, and performance optimization.
- 3+ years of experience using configuration management systems such as Puppet, Chef, Ansible, or SaltStack in production environments.
- Proven experience building, operating, and troubleshooting bare-metal Kubernetes clusters from the ground up, including control plane, etcd, and CNI plugin management.
- Proficiency with continuous system automation and scripting in languages such as Go, Python, Ruby, Perl, or Bash.
- Demonstrated experience responding to live service disruptions, leading root-cause analysis, and operating messaging systems (e.g., Kafka or RabbitMQ) in production.
- Experience operating, scaling, and monitoring AI/ML or LLM-powered services and workloads in high-concurrency production environments.
- Hands-on expertise with public cloud providers (AWS, GCE, or Azure) and containerized CI/CD pipelines (e.g., GitHub, Jenkins, CircleCI, Docker).
- Experience with distributed key-value stores (e.g., Consul, etcd, Zookeeper, Redis) and enterprise observability/alerting tools (e.g., Prometheus, Sensu).
- Familiarity with server virtualization infrastructure (e.g., Proxmox, VMware, Xen, OpenStack) and low-level networking concepts (IPtables/NFTables, routing, load balancing).
- Experience developing and maintaining OS-level software packaging (RPM/DEB) and participating in globally distributed software engineering teams.
This posting is for an existing vacancy.
What We Offer
~2 min readAt OpenTable, we pride ourselves on fostering a global and dynamic work environment. As a team member with us, you will benefit from a schedule tailored to accommodate a global workforce operating across multiple time zones. While the majority of your responsibilities may align with conventional business hours, there will be instances where you are expected to manage communications - via calls, Slack messages, or emails - outside of regular working hours to effectively collaborate with international colleagues, respond to restaurant partners, and/or address urgent matters. OpenTable will always abide by and consider local laws and regulations.
We’re committed to creating a workplace where everyone feels they belong and can thrive. We know the best ideas come when we bring different voices to the table, so we're building a team as dynamic as the diners and restaurants we serve—and fostering a culture where everyone feels welcome to be themselves.
If you need accommodations during the application or interview process, or on the job, we’re here to support you. Please reach out to your recruiter to request any accommodations.
#LI-Hybrid
Location & Eligibility
Listing Details
- Posted
- September 29, 2026
- First seen
- September 30, 2026
- Last seen
- September 30, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 67%
- Scored at
- September 30, 2026
Signal breakdown
Browse Similar Jobs
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
No spam. Unsubscribe at any time.