Orion Innovation

Senior Site Reliability Engineer

Reposted 15 Days Ago

Be an Early Applicant

Remote

Hiring Remotely in India

Senior level

Remote

Hiring Remotely in India

Senior level

Drive operational maturity for Kubernetes workloads on GKE, improve CI/CD pipelines, troubleshoot production issues, and standardize infrastructure resources.

The summary above was generated by AI

Orion Innovation is a premier, award-winning, global business and technology services firm. Orion delivers game-changing business transformation and product development rooted in digital strategy, experience design, and engineering, with a unique combination of agility, scale, and maturity. We work with a wide range of clients across many industries including financial services, professional services, telecommunications and media, consumer products, automotive, industrial automation, professional sports and entertainment, life sciences, ecommerce, and education.

Job Overview:

Drive reliability and operational maturity for Kubernetes workloads on GKE through safe rollout patterns, high-signal observability, resilient IaC, and effective incident response. Collaborate with developers to harden CI/CD pipelines and address infrastructure concerns within application code.

Key responsibilities:

Design and maintain resilient deployment patterns (blue-green, canary, GitOps syncs) across services.
Instrument and optimize logs, metrics, traces, and alerts to reduce noise and improve signal.
Review backend code (e.g., Django, Node.js, Go, Java) with a focus on infra touchpoints like database usage, timeouts, error handling, and memory consumption.
Tune and troubleshoot GKE workloads, HPA configs, network policies, and node pool strategies.
Improve or author Terraform modules for infrastructure resources (e.g., VPC, CloudSQL, Secrets, Pub/Sub).
Diagnose production issues from logs, traces, dashboards, and lead or support incident response.
Reduce config drift across environments and standardize secrets, naming, and resource tagging.
Collaborate with developers to harden delivery pipelines, standardize rollout readiness, and clean up infra smells in code.

Key skills:

Have 4–6+ years of experience in backend or infra-focused engineering roles (e.g., SRE, platform, DevOps, or fullstack).
Can confidently write or review production-grade code and infra-as-code (Terraform, Helm, GitHub Actions, etc.).
Have deep hands-on experience with Kubernetes in production, ideally on GKE, including workload autoscaling and ingress strategies.
Understand cloud concepts like IAM, VPCs, secret storage, workload identity, and CloudSQL performance characteristics.
Think in systems: you understand cascading failure, timeout boundaries, dependency health, and blast radius.
Regularly contribute to incident mitigation or long-term fixes (not just closing alerts).
Can influence through well-written PRs, documentation, and thoughtful design reviews.

Good to have:

Exposure to GitOps tooling such as ArgoCD or FluxCD.
Experience developing or integrating Kubernetes operators.
Familiarity with service-level indicators (SLIs), service-level objectives (SLOs), and structured alerting.

Tools and Expectations:

Datadog - Monitor infrastructure health, capture service-level metrics, reduce alert fatigue through high signal thresholds.
PagerDuty - Own incident management pipeline. Route alerts by severity and align with business SLAs.
GKE / Kubernetes - Improve cluster stability and workload isolation. Define auto-scaling configurations and tune for efficiency.
Helm / GitOps (ArgoCD/Flux) - Validate release consistency across clusters. Monitor sync status and rollout safety.
Terraform Cloud - Support DR planning and detect infrastructure changes through state comparisons.
CloudSQL / Cloudflare - Diagnose DB and networking issues. Monitor latency, enforce access patterns, and validate WAF usage.
Secret Management - Audit access to secrets, apply short-lived credentials, and define alerts for abnormal usage.

Orion is an equal opportunity employer, and all qualified applicants will receive consideration for employment without regard to race, color, creed, religion, sex, sexual orientation, gender identity or expression, pregnancy, age, national origin, citizenship status, disability status, genetic information, protected veteran status, or any other characteristic protected by law.

Candidate Privacy Policy

Orion Systems Integrators, LLC and its subsidiaries and its affiliates (collectively, “Orion,” “we” or “us”) are committed to protecting your privacy. This Candidate Privacy Policy (orioninc.com) (“Notice”) explains:

What information we collect during our application and recruitment process and why we collect it;
How we handle that information; and
How to access and update that information.

Your use of Orion services is governed by any applicable terms in this notice and our general Privacy Policy.

Top Skills

Cloudsql

Datadog

Github Actions

Gke

Helm

Kubernetes

Pagerduty

Terraform

Similar Jobs

Teikametrics

Senior Site Reliability Engineer

11 Days Ago

Remote

India

Mid level

Artificial Intelligence • eCommerce • Software

Teikametrics seeks a Senior Site Reliability Engineer to manage cloud infrastructure, improve DevOps practices, and support application deployment using technologies like Docker and Kubernetes.

Top Skills: Argo WorkflowsAWSAws RdsBashCi/CdCircleCIDatadogDockerGitJavaJavaScriptKubernetesOpensearchPostgresPythonSentryTerraform

Teikametrics

Senior Site Reliability Engineer

11 Days Ago

Remote

India

Mid level

Artificial Intelligence • eCommerce • Software

The Senior Site Reliability Engineer will manage cloud infrastructure, develop internal DevOps tools, and enhance automation for application deployment and management.

Top Skills: Argo WorkflowsAWSAws RdsBashCircleCIDatabricksDatadogDockerJavaJavaScriptKafkaKubernetesOpensearchPostgresPythonSentryTerraform

Nexthink

Senior Site Reliability Engineer

2 Days Ago

Remote or Hybrid

Senior level

Artificial Intelligence • Big Data • Cloud • Information Technology • Machine Learning • Software

The Senior Site Reliability Engineer will manage cloud-native systems, improve infrastructure, monitor systems, and automate while ensuring high availability and performance of digital experiences at Nexthink.

Top Skills: AWSBashDatadogDockerGithub ActionsGitlab CiGoJenkinsKubernetesLinuxPythonTerraform

What you need to know about the Mumbai Tech Scene

From haggling for the best price at Chor Bazaar to the bustle of Crawford Market, the energy of Mumbai's traditional markets is a key part of the city's charm. And while these markets will always have their place, the city also boasts a thriving e-commerce scene, ranking among the largest in the region. Driven by online sales in everything from snacks to licensed sports merchandise to children's apparel, the local industry is worth billions, with companies actively recruiting to meet the demands of continued growth.