Senior DevOps Engineer responsible for improving platform resilience and performance by collaborating with architects and engineering teams. Implement infrastructure improvements, ensure high availability, strengthen CI/CD and DR plans, assist incident response and post-mortems, document runbooks/playbooks, and share knowledge across the team.
We are looking for an experienced and collaborative Senior DevOps Engineer to join our DevOps team and help us improve the resilience and performance of the Arbor platform, enabling the business to rapidly scale. The remit and focus of the role is to use metrics and tooling to collaborate with the solution architect and engineering teams to continuously fix and/or improve the architecture, infrastructure and the way we work. It’s a broad and exciting role, so we’re looking for someone up for a challenge - if you’re a highly technical and diligent Engineer, this is the role for you.
Key Responsibilities- Work with the Head of Platform and Head of SRE to identify improvements within the platform infrastructure and implement plans to address
- Work with the Platform teams to improve the maturity of all components within the system, including ensuring High Availability, and adequate testing and DR plans
- Contribute to improving our CI/CD pipelines and providing patterns of deployment
- Assist in incident response and resolution, and subsequent post-mortems and retrospectives
- Participate in tech-talks and team based learning to ensure knowledge is spread
- Document obsessively, relying on Playbooks/Runbooks and systems documentation to aid knowledge transfer
Requirements
- 6-12 years of experience
- Extensive experience of DevOps Engineering and operating a large scale platform
- Extensive experience of distributed cloud systems, and specifically Amazon Web Services
- Extensive experience of Infrastructure as Code tooling, such as Terraform, Ansible, Cloudformation etc
- Understanding of relational database technologies and their cloud versions (e.g. AWS Aurora)
- Experience with messaging and distributed asynchronous workloads
- Experience with nginx or similar technologies
- Experience with DataDog, Prometheus or similar tools
- A positive and proactive attitude to problem solving
- A team player, willing to muck in and help others when needed, driven personality who asks questions and actively participates in discussions
- Good written and spoken English so you can present your ideas
Desired skills
- Past experience with enterprise solutions running at scale
- Familiarity with kanban and agile development processes
- Experience with Docker and containerisation
- Familiarity with software best practices such as Refactoring, Clean Code, Domain-Driven Design, Test-Driven Development, etc.
Benefits
The chance to work alongside a team of hard-working, passionate people in a role where you’ll see the impact of your work everyday. We also offer:
- Hybrid work environment
- Group Term Life Insurance paid out at 3x Annual CTC (Arbor India)
- 32 days holiday (plus Arbor Holidays). This is made up of 25 days annual leave plus 7 extra companywide days given over Easter, Summer & Christmas
- Work time: 9.30 am to 6 pm (8.5 hours only)
- Compensation - 100% fixed salary disbursement and no variable component
Similar Jobs
Edtech
The Senior DevOps Engineer will improve the Arbor platform’s resilience, performance, scalability, and infrastructure maturity. Responsibilities include enhancing high availability, testing, disaster recovery, CI/CD pipelines, deployment patterns, incident response, post-mortems, documentation, and knowledge sharing. The role requires extensive experience with AWS, large-scale distributed cloud platforms, Infrastructure as Code, databases, messaging systems, reverse proxies, monitoring tools, and containerization.
Top Skills:
Amazon Web Services (Aws)AnsibleAws AuroraAws CloudformationCi/CdDatadogDevOpsDistributed SystemsDockerInfrastructure As CodeMessaging SystemsNginxPrometheusRelational DatabasesTerraform
Analytics
Design, automate, and operate highly available AWS infrastructure and CI/CD pipelines. Provide 24x7 production support, incident management, monitoring (Splunk/New Relic/CloudWatch), networking, disaster recovery, and collaborate with development and data teams to improve reliability and automation.
Top Skills:
Amazon CloudwatchAmazon DynamodbAmazon EcsAmazon RdsAmazon S3Application Load BalancerAws CloudformationAws Secrets ManagerAws VpcBambooBashBitbucketDnsDockerGitHttp/HttpsIamKubernetes (Eks)Nat GatewayNew RelicPythonRoute 53Shell ScriptingSplunkSsl/TlsTcp/IpTerraformVpn
Agency • Information Technology
Design, implement, and maintain automated systems for deploying, scaling, and managing infrastructure and applications. Build CI/CD pipelines, manage IaC and orchestration (Kubernetes/EKS), monitor system health, implement security best practices, and collaborate with development and IT teams. Requires scripting/coding and cloud expertise to optimize reliability and performance.
Top Skills:
AWSBuild ToolsDockerEksGoGroovyInfrastructure-As-CodeJavaJenkinsKubernetesShell Scripts
What you need to know about the Mumbai Tech Scene
From haggling for the best price at Chor Bazaar to the bustle of Crawford Market, the energy of Mumbai's traditional markets is a key part of the city's charm. And while these markets will always have their place, the city also boasts a thriving e-commerce scene, ranking among the largest in the region. Driven by online sales in everything from snacks to licensed sports merchandise to children's apparel, the local industry is worth billions, with companies actively recruiting to meet the demands of continued growth.


