Ensure availability, performance, scalability, and reliability of production systems. Monitor systems with observability tools, define SLIs/SLOs, lead incident management and RCAs, automate operations via IaC, and support CI/CD pipeline stability and improvements.
- Job Summary
- The Site Reliability Engineer (SRE) is responsible for ensuring the availability, performance, scalability, and reliability of enterprise platforms and applications. The role focuses on monitoring, automation, incident management, and continuous improvement, working closely with engineering and DevOps teams to build resilient and highly available systems.
- 2. Key Responsibilities
- Ensure high availability and reliability of production systems and services
Monitor system health using observability tools (logs, metrics, traces)
Define and track SLIs, SLOs, and SLAs to measure system performance
Lead/support incident management, root cause analysis (RCA), and post-incident reviews
Automate operational tasks and implement Infrastructure as Code (IaC) practices
Support and improve CI/CD pipelines for stable and efficient releases
3. Skills & Competencies - Technical Skills
- Cloud Platforms: Azure
Monitoring & Observability: Prometheus, Grafana, Splunk, ELK, Datadog
Containers & Orchestration: Docker, Kubernetes
CI/CD Tools: Jenkins, GitHub Actions, GitLab CI, Azure DevOps
Infrastructure as Code: Terraform, Ansible, CloudFormation
- 2. Key Responsibilities
- Ensure high availability and reliability of production systems and services
Monitor system health using observability tools (logs, metrics, traces)
Define and track SLIs, SLOs, and SLAs to measure system performance
Lead/support incident management, root cause analysis (RCA), and post-incident reviews
Automate operational tasks and implement Infrastructure as Code (IaC) practices
Support and improve CI/CD pipelines for stable and efficient releases
3. Skills & Competencies
- Technical Skills
- Cloud Platforms: AWS / Azure / GCP
Monitoring & Observability: Prometheus, Grafana, Splunk, ELK, Datadog
Containers & Orchestration: Docker, Kubernetes
CI/CD Tools: Jenkins, GitHub Actions, GitLab CI, Azure DevOps
Infrastructure as Code: Terraform, Ansible, CloudFormation
Part of the $4.8 billion RPG Group, we’re a community of 10,000+ innovators across 30+ global locations, including Milpitas, Seattle, Princeton, Cape Town, London, Zurich, Singapore, and Mexico City. Explore Life at Zensar and join us to Grow. Own. Achieve. Learn. to be the best version of yourself.
We believe the best work happens when individuality is celebrated, growth is encouraged, and well-being is prioritized. We are an equal employment opportunity (EEO) and affirmative action employer, committed to creating an inclusive workplace. All qualified applicants will be considered without regard to race, creed, color, ancestry, religion, sex, national origin, citizenship, age, sexual orientation, gender identity, disability, marital status, family medical leave status, or protected veteran status.
Similar Jobs
Fintech • Payments • Financial Services
Leads application support for critical investment platforms, including P1/P2 incident resolution, root-cause analysis, monitoring, uptime improvement, release validation, and change management. Mentors L1/L2 support teams, establishes SOPs and runbooks, automates support processes using scripting, implements SRE practices, and coordinates communication among business stakeholders, vendors, leadership, development, and QA teams.
Top Skills:
.NetAPIsAppdynamicsAWSAzure DevopsCi/CdDevOpsElkGrafanaIisItilJavaLinuxMicrosoft Sql ServerOraclePowershellPythonServicenowShellSQLSreUnixWindows
Insurance
Leads the design and improvement of DevOps platforms, CI/CD pipelines, infrastructure automation, cloud environments, container orchestration, observability, security controls, and SRE practices. Partners with engineering teams to improve delivery speed, reliability, deployment quality, and developer self-service. Provides technical leadership, mentoring, governance, incident response, resilience engineering, and cloud cost optimization across production and non-production workloads.
Top Skills:
AnsibleAWSAzureAzure DevopsBashCi/CdCloudFormationContainer RegistriesDatadogDevsecopsDockerEfkElkGCPGithub ActionsGitlab CiGrafanaGroovyHelmInfrastructure As CodeJenkinsKubernetesOpentelemetryPowershellPrometheusPythonService MeshSplunkSreTerraform
Edtech • Software
Leads cloud platform architecture, reliability, modernization, automation, cost optimization, and operational excellence across Azure-based SaaS infrastructure with some AWS. Manages and develops cloud and database engineers, establishes observability and incident-response practices, improves CI/CD, supports highly available systems, and partners with engineering, product, and security teams. The role remains hands-on with Infrastructure-as-Code while overseeing hiring, performance management, capacity planning, and technical strategy.
Top Skills:
Arm TemplatesAWSAzure Kubernetes ServiceBicepC#Ci/CdFinopsInfrastructure-As-CodeAzurePowershellPythonSreTerraform
What you need to know about the Mumbai Tech Scene
From haggling for the best price at Chor Bazaar to the bustle of Crawford Market, the energy of Mumbai's traditional markets is a key part of the city's charm. And while these markets will always have their place, the city also boasts a thriving e-commerce scene, ranking among the largest in the region. Driven by online sales in everything from snacks to licensed sports merchandise to children's apparel, the local industry is worth billions, with companies actively recruiting to meet the demands of continued growth.


