New
New
Site Reliability Engineer (Onsite, Lahore, PKR Salary)
EngineeringDevops Engineer
0 views0 saves0 applied
Quick Summary
Key Responsibilities
Own the reliability, availability, and performance of production Linux GPU clusters, covering the operating system, drivers, GPUs, high-speed networking, and storage. Lead deep,
Requirements Summary
5+ years of experience in systems, infrastructure, or SRE engineering, operating production systems at scale. Deep Linux troubleshooting skills across the OS, networking, storage, and performance,
Technical Tools
EngineeringDevops Engineer
Requirements
~1 min read- 5+ years of experience in systems, infrastructure, or SRE engineering, operating production systems at scale.
- Deep Linux troubleshooting skills across the OS, networking, storage, and performance, with hands-on experience working as root on production systems. Experience with Ubuntu is highly relevant, as it is used almost exclusively.
- Hands-on experience operating GPU servers in production, including troubleshooting driver, device, and hardware-level issues, rather than only the workloads running on top of them.
- Practical network troubleshooting experience, including diagnosing physical-layer faults.
- Strong automation mindset with programming skills in Python or a comparable language.
- Experience with configuration management, node provisioning, and infrastructure-as-code (IaC) using Ansible, Terraform, or similar tools.
- Experience building observability and alerting solutions using Grafana and Prometheus.
- Bachelor's degree in Computer Science or equivalent experience.
- Experience operating GPU clusters or AI infrastructure at production scale.
- Production experience with Kubernetes or Slurm; experience with both is a bonus.
- Background in HPC or research computing.
- Familiarity with the NVIDIA GPU stack, InfiniBand/RDMA, and NCCL.
- Experience with CLI-based AI coding agents such as Claude Code, rather than browser-based assistants alone.
- Contributions to open-source projects within the cloud-native, HPC, or AI infrastructure ecosystem.
Responsibilities
~1 min read- →Own the reliability, availability, and performance of production Linux GPU clusters, covering the operating system, drivers, GPUs, high-speed networking, and storage.
- →Lead deep, end-to-end troubleshooting of complex distributed systems, GPU nodes, networking, and storage issues.
- →Troubleshoot and resolve GPU rail and NCCL performance issues across multi-GPU and multi-node collective communication paths.
- →Diagnose network faults end to end, including configuration, routing, and physical-layer issues such as cabling, transceivers, and link errors.
- →Configure and maintain workload managers that schedule customer jobs, including Kubernetes, Slurm, or both, along with the identity, storage, and networking services they depend on.
- →Build automation and tooling to eliminate operational toil, using a modern language such as Python or Go to design solutions, review implementations, and redirect approaches when needed.
- →Use AI-assisted engineering tools such as Claude to accelerate automation, runbook development, and incident analysis.
- →Automate provisioning, image deployment, configuration, and remediation using Ansible and infrastructure-as-code.
- →Design and operate observability using Grafana, Prometheus, and Loki, while tuning alerts for meaningful signals and building self-healing capabilities that reduce the need for human intervention.
- →Lead incident response, on-call activities, blameless postmortems, and reliability improvements that maintain customer SLAs.
- →Partner with Platform and Systems Engineering teams on capacity planning, rollouts, and continuous improvement.
8 PM - 4 AM
Location & Eligibility
Where is the job
Lahore, Pakistan
On-site at the office
Who can apply
PK
Listing Details
- First seen
- September 1, 2026
- Last seen
- September 1, 2026
Posting Health
- Days active
- 0
- Repost count
- 0
- Trust Level
- 51%
- Scored at
- September 1, 2026
Signal breakdown
freshnesssource trustcontent trustemployer trust
External application · ~5 min on hr-pod-hiring-talent-globally's site
Please let hr-pod-hiring-talent-globally know you found this job on Jobera.
3 other jobs at hr-pod-hiring-talent-globally
View all →Explore open roles at hr-pod-hiring-talent-globally.
Similar Devops Engineer jobs
View all →Senior Client Platform Engineer
$175k–$240k/yr
Site Reliability Engineer
K
K2SpacecorporationSenior Platform Engineer (DevX)
$165k–$200k/yr
Senior Site Reliability Engineer (Capacity) - Platform Infrastructure
N
New Era TechnologyRemoteSenior QRadar Platform Engineer
USD 80-125
Remote
C
CHAOS IndustriesSenior DevOps Engineer
$140k–$240k/yr
Browse Similar Jobs
Security2.2kDevOps & Infrastructure2.1kFullstack Developer2.1kEngineering Manager2.1kSoftware Architect1.8kQa Engineer1.7kBackend Developer1.5kMechanical Engineer1.4kSecurity Engineer1.4kFrontend Developer1.1kElectrical Engineer1.1kMobile Developer1kData Engineering945Backend Engineering936Project Engineer915Design Engineer829Product Engineer506Automation Engineer504Embedded Engineer500Frontend Engineering495
Newsletter
Stay ahead of the market
Get the latest job openings, salary trends, and hiring insights delivered to your inbox every week.
A
B
C
D
No spam. Unsubscribe at any time.