This role involves managing and automating infrastructure, deploying applications, ensuring system availability on cloud platforms, and collaborating with cross-functional teams.
Position: Site Reliability Engineer
Job Summary
As a Site Reliability Engineer, you will play a critical role in ensuring the availability and performance of our customer-facing platform. You will work closely with DevOps, DBA, and Development teams to provision and maintain infrastructure, deploy and monitor our applications, and automate workflows. Your contributions will have a direct impact on customer satisfaction and overall experience.
Responsibilities and Deliverables
- Manage, monitor, and maintain highly available systems (Windows and Linux)
- Analyze metrics and trends to ensure rapid scalability.
- Address routine service requests while identifying ways to automate and simplify.
- Create infrastructure as code using Terraform, ARM Templates, Cloud Formation.
- Maintain data backups and disaster recovery plans.
- Design and deploy CI/CD pipelines using GitHub Actions, Octopus, Ansible, Jenkins, Azure DevOps.
- Adhere to security best practices through all stages of the software development lifecycle
- Follow and champion ITIL best practices and standards.
- Become a resource for emerging and existing cloud technologies with a focus on AWS.
Organizational Alignment
- Reports to the Senior SRE Manager
- This role involves close collaboration with DevOps, DBA, and security teams.
Technical Proficiencies
- Hands-on experience with AWS is a must-have.
- Proficiency analyzing application, IIS, system, security logs and CloudTrail events
- Practical experience with CI/CD tools such as GitHub Actions, Jenkins, Octopus
- Experience with observability tools such as New Relic, Application Insights, AppDynamics, or DataDog.
- Experience maintaining and administering Windows, Linux, and Kubernetes.
- Experience in automation using scripting languages such as Bash, PowerShell, or Python.
- Configuration management experience using Ansible, Terraform, Azure Automation Run book or similar.
- Experience with SQL Server database maintenance and administration is preferred.
- Good Understanding of networking (VNET, subnet, private link, VNET peering).
- Familiarity with cloud concepts including certificates, Oauth, AzureAD, ASE, ASP, AKS, Azure Apps, Load Balancers, Application Gateway, Firewall, Load Balancer, API Management, SQL Server, Databases on Azure
Experience
- 5+ years of experience in SRE or System Administration role
- Demonstrated ability building and supporting high availability Windows/Linux servers, with emphasis on the WISA stack (Windows/IIS/SQL Server/ASP.net)
- 3+ years of experience with CI/CD tools
- 3+ years of experience working with cloud technologies including AWS, Azure.
- 1+ years of experience working with container technology including Docker and Kubernetes.
- Comfortable using Scrum, Kanban, or Lean methodologies.
Education
- Bachelor’s Degree or College Diploma in Computer Science, Information Systems, or equivalent experience.
Similar Jobs
Digital Media • eCommerce • Gaming • Mobile • News + Entertainment
Lead reliability, scalability, observability, automation, infrastructure, disaster recovery, and security initiatives for Crunchyroll’s cloud-native data platforms. Establish SRE practices including SLIs, SLOs, error budgets, incident management, and postmortems. Operate Kubernetes and GCP environments, implement Infrastructure as Code, optimize capacity and performance, and drive vulnerability remediation, penetration-testing support, and cloud platform security.
Top Skills:
Ci/CdDatadogGCPGoGrafanaIdentity And Access ManagementInfrastructure As CodeJavaKubernetesLinuxOpentelemetryOwasp Top 10PrometheusPythonShellTerraform
Fintech • Real Estate • Software
Leads the architecture, reliability, security, and operational excellence of cloud infrastructure systems. Owns SLOs, monitoring, incident response, on-call practices, and preventative reliability improvements. Drives medium-to-large SRE projects, collaborates across engineering and product, manages project risks, and mentors engineers. Requires deep experience with AWS, Linux, infrastructure as code, Kubernetes, PostgreSQL, observability, CI/CD, cloud security, and production AI/LLM integrations.
Top Skills:
Api GatewaysAWSAws CdkAws Well-Architected FrameworkCi/CdCloudFormationDistributed TracingEmbeddingsGitopsGoIamKubernetesLinuxLlmsLoggingMetricsMicroservicesObservabilityPostgresPythonRagSecrets ManagementService MeshTerraform
Artificial Intelligence • Information Technology • Machine Learning • Professional Services • Software • Analytics • Consulting
Respond to P1/P2 production incidents in a 24/7 rotation, triage infrastructure failures using Kubernetes, RabbitMQ, and PostgreSQL logs, stabilize systems, implement infrastructure fixes, escalate application issues, communicate incident status, and support system hardening and scaling improvements at senior levels.
Top Skills:
AWSAzureGCPKubernetesPostgresRabbitMQ
What you need to know about the Mumbai Tech Scene
From haggling for the best price at Chor Bazaar to the bustle of Crawford Market, the energy of Mumbai's traditional markets is a key part of the city's charm. And while these markets will always have their place, the city also boasts a thriving e-commerce scene, ranking among the largest in the region. Driven by online sales in everything from snacks to licensed sports merchandise to children's apparel, the local industry is worth billions, with companies actively recruiting to meet the demands of continued growth.



