SID Global Solutions Logo

SID Global Solutions

SRE

Posted One Month Ago
In-Office
Mumbai, Maharashtra, IND
Junior
In-Office
Mumbai, Maharashtra, IND
Junior
Monitor production infrastructure and observability stacks (Datadog, Dynatrace, Prometheus, Grafana), triage alerts, escalate incidents, manage API traffic (Apigee), support Nginx and Kubernetes, use GCP monitoring, execute runbooks, and operate in a 24/7 rotating shift model to ensure availability and performance of banking services.
The summary above was generated by AI

Job Title: Site Reliability Engineer (SRE) / L1 Monitoring Engineer

Job Summary

We are seeking a proactive and technically driven SRE / L1 Monitoring Engineer with 1 to 3 years of experience to join our core digital infrastructure operations team. In this role, you will serve as the first line of defense ensuring the high availability, security, and performance of critical financial services and digital banking applications. You will be responsible for real-time system monitoring, tracking alerts across modern observability stacks, performing initial triage on infrastructure bottlenecks, and managing API traffic performance. This is an excellent opportunity for an early-career engineer looking to scale their skills in a high-volume, secure cloud infrastructure environment. [1]


Key Responsibilities

L1 Infrastructure Monitoring & Alerts

  • Real-time Surveillance: Actively monitor production environments, enterprise dashboards, and telemetry feeds using toolsets like Datadog, Dynatrace, and Grafana to spot anomalies before they impact end-users. [1, 2, 3]
  • Alert Triage: Acknowledge, validate, and categorize incoming infrastructure, database, and application alerts generated by Prometheus and application performance monitoring (APM) agents using predefined Standard Operating Procedures (SOPs). [1, 2, 3, 4, 5]
  • Incident Escalation: Document incident details clearly in the ticketing system and swiftly escalate unresolved P1/P2 issues to L2 engineers or specialized DevOps teams with complete log snippets and context.

Application Delivery & API Traffic Management

  • Nginx Operations: Monitor web server logs, verify reverse proxy configurations, and troubleshoot basic traffic routing or SSL/TLS certificate errors. [1, 2, 3]
  • API Gateways: Use Apigee to monitor API proxy performance, track error rates (5xx/4xx codes), track latency spikes, and check developer portal connectivity. [1, 2, 3, 4]
  • Kubernetes Support: Monitor cluster health, inspect pod statuses, view application logs using kubectl, and track resource usage (CPU/Memory limits). [1, 2]

Cloud Operations & Reliability

  • GCP Monitoring: Utilize Google Cloud logging, native monitoring tools, and integrated observability dashboards to check the health of virtual machines, storage, and networking layers. [1, 2, 3, 4]
  • Health Checks: Perform routine daily morning sanity checks and post-deployment validation steps for critical banking services.
  • Runbook Execution: Execute automated or manual scripts to restart failed services, clear disk space, or cycle pods safely in staging and production environments.

Required Qualifications & Technical Skills

  • Experience: 1 to 3 years of hands-on experience in an L1 Support, Infrastructure Monitoring, or Junior SRE role.
  • Observability Tools: Hands-on experience navigating and tracking alerts within Datadog, Dynatrace, Prometheus, and Grafana.
  • Web Servers: Practical understanding of Nginx (reverse proxy, load balancing, log analysis).
  • Containerization: Foundational knowledge of Kubernetes (K8s) (understanding pods, deployments, services, and basic troubleshooting commands like kubectl logs and kubectl get pods).
  • Cloud Platform: Familiarity with Google Cloud Platform (GCP) core services and cloud monitoring concepts.
  • API Management: Exposure to Apigee or equivalent API gateways for monitoring traffic flow and checking endpoint health.
  • Operating Systems: Strong command-line comfort in Linux/Unix environments for navigating directories and tailing logs.
  • Shift Flexibility: Readiness to work in a 24/7 rotating shift model (including night shifts and weekends) to maintain uninterrupted banking infrastructure support

 



Similar Jobs

Yesterday
In-Office
Mumbai, Maharashtra, IND
Mid level
Mid level
Healthtech • HR Tech • Professional Services
Designs, implements, and operates secure CI/CD pipelines and platform automation. The role builds infrastructure-as-code and DevSecOps controls, supports application onboarding and deployments, and standardizes build and release processes. Responsibilities include coding, testing, integration, observability, compliance, vulnerability tracking, change management, service management, and enterprise DevOps enablement. The engineer also supports cloud and container platforms and provides hands-on assistance to application teams in a regulated banking environment.
Top Skills: AksAnsibleAqua SecurityArtifactoryAstronomer AirflowAWSAws CloudformationAzureAzure Resource ManagerBashBitbucketCheckmarxCi/CdConfiguration As CodeDevOpsDevsecopsDockerEksGitGitGithub ActionsGitopsGkeGoogle Cloud PlatformHarness CdInfrastructure As CodeJenkinsJfrogKubernetesOpenshiftPlatform EngineeringPolicy As CodePowershellPythonRancherRemedyServicenowSonarqubeSreTerraformXlr
2 Days Ago
In-Office or Remote
India
Senior level
Senior level
Insurance
Leads the design and improvement of DevOps platforms, CI/CD pipelines, infrastructure automation, cloud environments, container orchestration, observability, security controls, and SRE practices. Partners with engineering teams to improve delivery speed, reliability, deployment quality, and developer self-service. Provides technical leadership, mentoring, governance, incident response, resilience engineering, and cloud cost optimization across production and non-production workloads.
Top Skills: AnsibleAWSAzureAzure DevopsBashCi/CdCloudFormationContainer RegistriesDatadogDevsecopsDockerEfkElkGCPGithub ActionsGitlab CiGrafanaGroovyHelmInfrastructure As CodeJenkinsKubernetesOpentelemetryPowershellPrometheusPythonService MeshSplunkSreTerraform
6 Days Ago
In-Office
Senior level
Senior level
Logistics • Transportation
Leads Technology Operational Excellence strategy, roadmap, governance, and execution across recovery assurance, resilience, technology risk, audit, standards assurance, and operational improvement. Drives enterprise recovery exercises, risk remediation, automated controls, AI-enabled transformation, reliability improvements, and engineering standards adoption across cloud, SaaS, on-premises, and hybrid environments. Builds and develops multidisciplinary teams while influencing senior technology, platform, cybersecurity, audit, and business stakeholders.
Top Skills: AIAutomationCloudObservabilityPolicy-As-CodeReliability EngineeringSaaSSre

What you need to know about the Mumbai Tech Scene

From haggling for the best price at Chor Bazaar to the bustle of Crawford Market, the energy of Mumbai's traditional markets is a key part of the city's charm. And while these markets will always have their place, the city also boasts a thriving e-commerce scene, ranking among the largest in the region. Driven by online sales in everything from snacks to licensed sports merchandise to children's apparel, the local industry is worth billions, with companies actively recruiting to meet the demands of continued growth.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account