Ensure availability, performance, scalability, and reliability of production systems. Monitor systems with observability tools, define SLIs/SLOs, lead incident management and RCAs, automate operations via IaC, and support CI/CD pipeline stability and improvements.
- Job Summary
- The Site Reliability Engineer (SRE) is responsible for ensuring the availability, performance, scalability, and reliability of enterprise platforms and applications. The role focuses on monitoring, automation, incident management, and continuous improvement, working closely with engineering and DevOps teams to build resilient and highly available systems.
- 2. Key Responsibilities
- Ensure high availability and reliability of production systems and services
Monitor system health using observability tools (logs, metrics, traces)
Define and track SLIs, SLOs, and SLAs to measure system performance
Lead/support incident management, root cause analysis (RCA), and post-incident reviews
Automate operational tasks and implement Infrastructure as Code (IaC) practices
Support and improve CI/CD pipelines for stable and efficient releases
3. Skills & Competencies - Technical Skills
- Cloud Platforms: Azure
Monitoring & Observability: Prometheus, Grafana, Splunk, ELK, Datadog
Containers & Orchestration: Docker, Kubernetes
CI/CD Tools: Jenkins, GitHub Actions, GitLab CI, Azure DevOps
Infrastructure as Code: Terraform, Ansible, CloudFormation
- 2. Key Responsibilities
- Ensure high availability and reliability of production systems and services
Monitor system health using observability tools (logs, metrics, traces)
Define and track SLIs, SLOs, and SLAs to measure system performance
Lead/support incident management, root cause analysis (RCA), and post-incident reviews
Automate operational tasks and implement Infrastructure as Code (IaC) practices
Support and improve CI/CD pipelines for stable and efficient releases
3. Skills & Competencies
- Technical Skills
- Cloud Platforms: AWS / Azure / GCP
Monitoring & Observability: Prometheus, Grafana, Splunk, ELK, Datadog
Containers & Orchestration: Docker, Kubernetes
CI/CD Tools: Jenkins, GitHub Actions, GitLab CI, Azure DevOps
Infrastructure as Code: Terraform, Ansible, CloudFormation
Part of the $4.8 billion RPG Group, we’re a community of 10,000+ innovators across 30+ global locations, including Milpitas, Seattle, Princeton, Cape Town, London, Zurich, Singapore, and Mexico City. Explore Life at Zensar and join us to Grow. Own. Achieve. Learn. to be the best version of yourself.
We believe the best work happens when individuality is celebrated, growth is encouraged, and well-being is prioritized. We are an equal employment opportunity (EEO) and affirmative action employer, committed to creating an inclusive workplace. All qualified applicants will be considered without regard to race, creed, color, ancestry, religion, sex, national origin, citizenship, age, sexual orientation, gender identity, disability, marital status, family medical leave status, or protected veteran status.
Similar Jobs
Fintech • Information Technology
Develop and maintain telemetry and automation tooling to monitor and manage global platform health. Participate in on-call rotations, diagnose and resolve production incidents, implement automated incident response and reliability improvements, and proactively mitigate system stability risks.
Top Skills:
Aws CloudwatchAws DynamodbAws Ec2Aws EksAws ElbAws LambdaAws RdsAws SqsGoLinuxPythonTerraform
Information Technology
Design, operate, and automate AWS-based production and development infrastructure for eCommerce/enterprise platforms. Implement CI/CD, infrastructure-as-code, observability, security, disaster recovery, and large-scale automation. Support microservices, databases, caching, virtualization, and application servers while troubleshooting, performance tuning, and onboarding new tools.
Top Skills:
AkamaiAndroidApacheAppdynamicsAptitude/DpkgAwkAWSBashCassandraCdnChefDatadogDynatraceEc2Elastic CloudElkGeodnsGlobal Traffic ManagementGraphiteGroovyHadoopHaproxyHbaseHelmHypervisorIptablesJavaJavaScriptJbossJenkinsJettyJSONKeycloakKubernetesLdapLinuxMemcachedMicroservicesMongoDBMySQLNagiosNessusNetappNew RelicNfsNginxNmapNtpObjective-COktaOpen DirectoryOraclePerlPHPPuppetPythonRackspace CloudRedisRestRubyService MeshSoftlayerSplunkSsl/TlsTerraformTomcatVarnishVdiVirtualizationWeblogicXMLYum/Rpm
Information Technology
Design, implement, and maintain scalable, secure GCP infrastructure and Kubernetes deployments. Build and optimize CI/CD pipelines (Jenkins), automate IaC (Terraform/Deployment Manager), monitor systems (Prometheus, Grafana, Cloud Monitoring, ELK), troubleshoot production issues, apply SRE practices (SLIs/SLOs, incident management), and harden Linux-based environments.
Top Skills:
ArgocdBashCloud MonitoringCloud StorageCompute EngineDeployment ManagerDockerElasticsearchElkGCPGitGkeGoGrafanaIamJbossJenkinsKibanaKubernetesLinuxLogstashPrometheusPythonSpinnakerStackdriverTerraformVaultVpcWildfly
What you need to know about the Mumbai Tech Scene
From haggling for the best price at Chor Bazaar to the bustle of Crawford Market, the energy of Mumbai's traditional markets is a key part of the city's charm. And while these markets will always have their place, the city also boasts a thriving e-commerce scene, ranking among the largest in the region. Driven by online sales in everything from snacks to licensed sports merchandise to children's apparel, the local industry is worth billions, with companies actively recruiting to meet the demands of continued growth.

