Ensono
Ensono2d ago
New

Senior Mainframe Systems Programmer - Site Reliability Engineering

IndiaIndia·Punesenior
OtherProgrammer
0 views0 saves0 applied

Quick Summary

Key Responsibilities

- Design and deploy Infrastructure-as-Code (IaC) solutions using Ansible, Zowe CLI, and z/OSMF workflows to automate system provisioning, configuration management, and recovery processes.

Requirements Summary

IBM z/OS System Programming, Broadcom Mainframe SRE, or Hashicorp Terraform. - Familiarity with Zowe Desktop for modern IDE-driven development or Dynatrace APM for CICS/Db2 monitoring.

Technical Tools
OtherProgrammer

 

 

Ensono is a hybrid IT services provider that sees a world in transformation. We understand the world of business has become more complex, layered, and interdependent because we’re enabling the IT infrastructure for that transformation. Our clients are some of the most innovative and forward-thinking companies in the world and we help keep their businesses thriving. With nearly 50 years of experience, we optimize and modernize IT infrastructure by amplifying the power of mainframes and mid-range servers with the agility of cloud. Check us out at www.ensono.com

We are seeking a visionary Mainframe Site Reliability Engineer (SRE) to redefine the reliability, automation, and efficiency of our mission-critical z/OS systems. This role combines deep mainframe expertise with cutting-edge SRE practices, focusing on innovations in observability, AI-driven operations, and DevOps integration to transform legacy workflows into modern, self-healing systems. You will drive initiatives to eliminate manual toil, optimize performance, and ensure the platform’s resilience aligns with business-critical service level objectives (SLOs).

Responsibilities

~1 min read

- Design and deploy Infrastructure-as-Code (IaC) solutions using Ansible, Zowe CLI, and z/OSMF workflows to automate system provisioning, configuration management, and recovery processes.

- Develop self-healing workflows for critical subsystems (CICS, Db2, IMS) to auto-resolve incidents like JVM failures or transaction bottlenecks.

- Convert legacy operational scripts (REXX, NCL) into modern, version-controlled pipelines integrated with Git and CI/CD tools like Jenkins.

- Implement predictive analytics tools (e.g., IBM Watson AIOps, Splunk ITSI) to detect anomalies in system metrics, logs, and message queues.

- Build dashboards using Grafana or Prometheus to visualize the Four Golden Signals (latency, traffic, errors, saturation) across mainframe workloads.

- Centralize alert management to reduce noise and prioritize actionable alerts using AI-driven correlation.

  1. DevOps Integration & Modernization

- Streamline software delivery pipelines for COBOL/PL/I applications using IBM Dependency-Based Build (DBB) and UrbanCode Deploy (UCD).

- Integrate mainframe SDLC processes with enterprise Git repositories (GitHub, GitLab) to enable collaborative development and audit trails.

- Enable automated testing and phased rollouts for z/OS middleware updates.

- Performance & Capacity Engineering:

- Optimize CPU/MIPS utilization through runtime tuning (e.g., CICS Threadsafe, AT-TLS offloading) to reduce software licensing costs.

- Forecast capacity demands using historical SMF/RMF data and propose dynamic hardware scaling strategies.

- Conduct load testing for batch and OLTP workloads to validate system limits and error budgets.

  1. Incident Management & Reliability

- Lead blameless postmortems for critical incidents, focusing on root cause analysis (RCA) and preventive actions (e.g., monitoring gaps, automation fixes).

- Reduce MTTR by implementing automated incident response playbooks (e.g., auto-restart failed subsystems, reroute traffic).

- Maintain 24/7 operational readiness through on-call rotations and cross-training in z/OS, CICS, Db2, and storage management.

  1. Platform Hardening & Knowledge Sharing

- Enforce security best practices (RACF, TLS) and vulnerability remediation for z/OS and middleware.

- Develop reusable workbooks and runbooks to document system configurations, troubleshooting steps, and automation workflows.

- Mentor teams on SRE principles, fostering a T-shaped skill model (deep mainframe + DevOps/Agile practices).

 

  1. Batch Optimization & Resource Management

  - Design dynamic resource allocation strategies (e.g., WLM policies, enclaves) to prioritize critical batch jobs and minimize contention for CPU, memory, and I/O resources.  

  - Implement parallel processing (e.g., multi-task JCL, SYSAFF routing) to reduce runtime and avoid bottlenecks in long-running batch cycles.  

  - Streamline job dependencies using graph-based scheduling tools (e.g., IWS, CA7, Control-M) to eliminate idle wait times between interdependent jobs.  

  1. Proactive Batch Health Monitoring :  

  - Develop automated checks for batch job SLAs, including real-time alerts for delays, resource starvation, or dataset contention.  

  - Integrate predictive analytics (e.g., historical SMF data analysis) to forecast and mitigate delays caused by seasonal peaks or data volume spikes.  

---

- xx+ years in z/OS system programming, performance tuning, or infrastructure support.

- Proficiency in JCL, REXX, Python, and mainframe automation tools (IBM Z System Automation, Broadcom OPS/MVS).

- Hands-on experience with Zowe, Ansible, Git, and CI/CD pipelines.

- Mastery of SRE tenets: SLOs/SLIs, error budgets, and Infrastructure-as-Code (IaC).

- Innovation Focus:

- Proven track record in implementing AI/ML-driven monitoring or auto-remediation for mainframe environments.

- Experience modernizing legacy workflows (e.g., replacing CA Endevor with Git-based SDLC).

- Ability to lead cross-functional teams during high-severity incidents.

- Strong communication to align technical execution with business objectives.

- Bachelor’s degree in Computer Science, Engineering, or related field.

---

Requirements

~1 min read

- Experience with AI-Driven Automation platforms (e.g. AMELIA AIOps) to standardize and migrate legacy workflows, integrate with event management systems (e.g., BigPanda), and orchestrate ITIL processes (Incident, changes) via ServiceNow

- Certifications: IBM z/OS System Programming, Broadcom Mainframe SRE, or Hashicorp Terraform.

- Familiarity with Zowe Desktop for modern IDE-driven development or Dynatrace APM for CICS/Db2 monitoring.

- Knowledge of mainframe open-source ecosystems (Zowe, Feilong) or hybrid-cloud integrations.

Shift Timing- 1:30 PM to 10:30 PM

 

We are an equal opportunity employer. All qualified applicants will be considered for employment without regard to caste, colour, creed, religion, gender, gender identity, sexual orientation, age, disability, HIV status, or any other status protected by law. Candidates with disabilities who require accommodations during the recruitment process are encouraged to contact our Talent Acquisition team to place a request.

Location & Eligibility

Where is the job
Pune, India
On-site at the office
Who can apply
IN

Listing Details

Posted
September 19, 2026
First seen
September 19, 2026
Last seen
September 21, 2026

Posting Health

Days active
0
Repost count
0
Trust Level
67%
Scored at
September 19, 2026

Signal breakdown

freshnesssource trustcontent trustemployer trust
Ensono
Ensono
greenhouse

Ensono is a premier managed service provider delivering complete Hybrid IT solutions, committed to optimizing clients' digital transformation journeys.

Employees
350
Founded
2016
View company profile
Newsletter

Stay ahead of the market

Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.

A
B
C
D
Join 12,000+ marketers

No spam. Unsubscribe at any time.

EnsonoSenior Mainframe Systems Programmer - Site Reliability Engineering