Job Summary
EXL is looking for a skilled Managed Service Engineer with strong experience in DevOps practices, AWS cloud services, infrastructure monitoring, automation, incident management, and production support. The candidate will be responsible for maintaining cloud environments, supporting CI/CD pipelines, managing incidents, ensuring system reliability, and driving operational excellence for managed service engagements.
The role requires hands-on technical expertise, strong troubleshooting skills, and the ability to work with cross-functional teams to ensure high availability, performance, and security of cloud-based platforms
ResponsibilitiesManaged Services & Production Support
- Provide day-to-day support for cloud infrastructure, applications, and data platforms.
- Monitor system health, availability, performance, and security alerts.
- Handle incidents, service requests, change requests, and problem management.
- Perform root cause analysis for recurring issues and drive permanent fixes.
- Ensure adherence to SLA, OLA, and operational support processes.
- Maintain runbooks, SOPs, knowledge articles, and support documentation.
DevOps Engineering
- Manage and support CI/CD pipelines for application and infrastructure deployments.
- Automate repetitive operational tasks using scripts and DevOps tools.
- Support build, release, and deployment activities across environments.
- Work with development teams to troubleshoot deployment and environment issues.
- Implement infrastructure automation using Infrastructure as Code practices.
AWS Cloud Operations
- Manage and support AWS services such as:
- EC2
- S3
- IAM
- VPC
- Lambda
- RDS
- CloudWatch
- CloudTrail
- ELB / ALB
- Auto Scaling
- ECS / EKS, if applicable
- Monitor AWS resource utilization and recommend optimization opportunities.
- Support backup, recovery, patching, access management, and cloud security activities.
- Assist in cloud cost optimization and operational governance.
Monitoring, Incident & Problem Management
- Configure and manage monitoring dashboards, alerts, and logs.
- Respond to alerts and incidents based on defined severity levels.
- Coordinate with application, infrastructure, security, and business teams during outages.
- Perform proactive health checks and preventive maintenance.
- Support post-incident reviews and implement corrective actions.
Security & Compliance
- Follow cloud security best practices for access, network, and data protection.
- Support IAM role management, key rotation, vulnerability remediation, and audit requests.
- Ensure compliance with organizational security and change management policies.
Assist in maintaining operational readiness for audits and compliance reviews
QualificationsRequired Technical Skills
- Strong hands-on experience with AWS cloud services.
- Good understanding of DevOps tools and CI/CD pipelines.
- Experience with tools such as:
- Jenkins / GitHub Actions / GitLab CI / Azure DevOps
- Git / Bitbucket
- Terraform / CloudFormation
- Docker / Kubernetes
- Ansible, if applicable
- Knowledge of Linux/Unix administration and shell scripting.
- Experience with monitoring and logging tools such as:
- AWS CloudWatch
- Splunk
- Datadog
- Grafana / Prometheus
- ELK Stack
- Good understanding of networking concepts such as VPC, subnets, routing, security groups, load balancers, and DNS.
- Familiarity with ITSM tools such as ServiceNow, Jira, or similar platforms.
- Basic understanding of security, backup, disaster recovery, and cloud governance.
Required Experience
- 3 to 8 years of experience in IT operations, cloud support, DevOps, or managed services.
- Minimum 2+ years of hands-on AWS experience.
- Experience working in production support or managed service environments.
- Experience handling incidents, service requests, and change management.
- Exposure to 24x7 support models, shift operations, or on-call support is preferred


