IT Site Reliability Engineer — API Management Platforms

United StatesUnited States·Dallasmid
EngineeringDevops Engineer
2 views0 saves0 applied

Quick Summary

Key Responsibilities

Serve as the go-to subject matter expert (SME) for TI's Apigee Edge private cloud platform, owning the full platform lifecycle from installation and configuration through ongoing operations, upgrades,

Requirements Summary

Monitor platform resource utilization (CPU, memory, disk, JVM heap); conduct capacity planning and performance tuning across all Apigee Edge components and underlying infrastructure.

Technical Tools
EngineeringDevops Engineer

About the Role

~1 min read

The Enterprise Platforms team within our Data & Agentic Platform Solutions organization is responsible for deploying, operating, and continuously improving the platforms that power TI's digital integration, automation and DevOps capabilities. As an IT Site Reliability Engineer within the Enterprise Platforms team, you will serve as the primary technical platform owner for TI's Apigee Edge private cloud environment — the backbone of TI's API management and integration strategy — while also providing platform SRE support across a broader portfolio of automation and DevOps tooling used throughout the enterprise.

 

This role sits at the intersection of infrastructure administration, DevOps engineering, and platform reliability. You will be expected to bring deep technical expertise in API management operations while also developing familiarity with adjacent platforms such as CI/CD pipelines, artifact repositories, source control systems, automation orchestration tools, and integration middleware. You will work closely with application development teams, security, network, and infrastructure teams to ensure all supported platforms are stable, performant, secure, and scalable.

 

 

Responsibilities

~1 min read
  • Primary Technical Owner: Serve as the go-to subject matter expert (SME) for TI's Apigee Edge private cloud platform, owning the full platform lifecycle from installation and configuration through ongoing operations, upgrades, and eventual roadmap evolution.
  • Platform Administration: Install, configure, upgrade, patch, and maintain Apigee Edge private cloud components (Management Server, Router, Message Processor, Cassandra, ZooKeeper, Qpid, Postgres) across dev/test and production environments.
  • Monitoring & Intervention: Own and improve system monitoring solution to ensure early detection of issues and implementation of self-healing solutions where appropriate.
  • Incident Management: Serve as the primary escalation point for all Apigee-related incidents; lead root-cause analysis (RCA) and drive post-incident reviews (PIRs) to prevent recurrence.
  • API Proxy Lifecycle Support: Partner with development teams on API proxy deployment pipelines, policy troubleshooting, and runtime performance optimization within the Apigee environment.
  • Security & Compliance: Apply and enforce security hardening standards on the Apigee platform; manage TLS/SSL certificates, OAuth configurations, and role-based access controls (RBAC); ensure alignment with TI IT security policies and audit requirements.
  • Capacity Planning & Performance Tuning: Monitor platform resource utilization (CPU, memory, disk, JVM heap); conduct capacity planning and performance tuning across all Apigee Edge components and underlying infrastructure.
  • Backup & Disaster Recovery: Own and maintain backup, restore, and disaster recovery procedures for all Apigee Edge components and associated data stores (Cassandra, PostgreSQL, ZooKeeper).
  • Vendor Engagement: Manage the technical relationship with Google/Apigee support for escalated issues, product defects, version roadmaps, and advisory services.
  • Platform Roadmap Input: Provide technical guidance and recommendations on Apigee platform strategy, including future-state considerations such as migration paths to Apigee Hybrid hosted solution.

 

  • Multi-Platform SRE: Provide site reliability engineering support for other automation and DevOps platforms within the ITS DevOps Platforms portfolio — which may include CI/CD tools, artifact management, source control, integration middleware, and automation orchestration platforms.
  • Operational Consistency: Apply consistent SRE principles across all supported platforms — defining SLOs/SLIs, building observability dashboards, maintaining runbooks, and driving reliability improvements.
  • Cross-Platform Incident Support: Respond to and triage incidents across the broader platform portfolio; coordinate with platform-specific SMEs and infrastructure teams to restore services rapidly.
  • Automation & Tooling Development: Develop and maintain automation scripts and Infrastructure-as-Code (IaC) tooling (Python, Bash, Ansible, Terraform, or similar) to streamline operations, deployments, and configuration management across all supported platforms.
  • Documentation & Knowledge Management: Create and maintain thorough technical documentation including architecture diagrams, operational runbooks, change records, and knowledge base articles for all platforms under support.
  • Cross-Team Collaboration: Partner with networking, server infrastructure, application development, and security teams to support integrations, troubleshoot complex issues, and deliver platform enhancements across the automation portfolio.

Requirements

~2 min read
  • Bachelor's degree in Computer Science, Information Technology, Information Systems, or a related field (or equivalent practical experience)
  • 3+ years of hands-on experience with API management on Apigee or similar API Management platform

 

  • Strong proficiency with Linux/Unix system administration (RHEL, Rocky, or similar)
  • Proficiency in at least one scripting/automation language: Python, Bash, or PowerShell
  • Familiarity with REST API concepts, API lifecycle management, and API security patterns (OAuth 2.0, API keys, JWT)
  • Experience with monitoring and observability tools, including one or more of the following: Elastic (ELK Stack — Elasticsearch, Logstash, Kibana), Splunk, Prometheus, Grafana, or equivalent
  • Understanding of networking concepts: TCP/IP, DNS, load balancing, TLS/SSL, firewalls, and proxies
  • Demonstrated experience supporting or administering at least one other DevOps or automation platform beyond Apigee (CI/CD, artifact management, source control, or similar)
  • Ability to provide on-call support for critical platform incidents outside normal work schedule
  • Apigee Certified API Engineer or equivalent certification
  • Experience with Apigee Edge components: Management Server, Router, Message Processor, Cassandra, ZooKeeper, Qpid, and PostgreSQL
  • Experience with infrastructure-as-code tools: Terraform, Ansible, Chef, or Puppet
  • Familiarity with CI/CD pipeline tools: Jenkins, GitLab CI, GitHub Actions, or similar
  • Experience managing artifact repositories such as JFrog Artifactory or Sonatype Nexus
  • Experience with containerization and orchestration: Docker, Kubernetes
  • Knowledge of cloud platforms (AWS, GCP, or Azure) and hybrid cloud architectures
  • Exposure to API mediation patterns, microservices architecture, and event-driven integration
  • Experience working in ITIL-based environments (Change Management, Incident Management, Problem Management)
  • Prior experience in a semiconductor, manufacturing, or large enterprise IT environment

Location & Eligibility

Where is the job
Dallas, United States
On-site at the office
Who can apply
US

Listing Details

Posted
August 13, 2026
First seen
August 13, 2026
Last seen
August 13, 2026

Posting Health

Days active
0
Repost count
0
Trust Level
51%
Scored at
August 13, 2026

Signal breakdown

freshnesssource trustcontent trustemployer trust

3 other jobs at 118-WW TMG MFG OPS

View all →

Explore open roles at 118-WW TMG MFG OPS.

Newsletter

Stay ahead of the market

Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.

A
B
C
D
Join 12,000+ marketers

No spam. Unsubscribe at any time.

118-WW TMG MFG OPSIT Site Reliability Engineer — API Management Platforms