Akamai Technologies Logo

Akamai Technologies

Senior Site Reliability Engineer

Reposted An Hour Ago
Be an Early Applicant
In-Office or Remote
Hiring Remotely in India
Senior level
In-Office or Remote
Hiring Remotely in India
Senior level
Designs, develops, and operates infrastructure, observability platforms, and internal tooling for Akamai’s Compute products. Responsibilities include proactive troubleshooting, automation, systems programming, Kubernetes and production-system operations, reliability and scalability improvements, cross-team collaboration, and guidance on service performance. The role requires Linux administration, networking knowledge, CI/CD, infrastructure as code, configuration management, cloud storage exposure, and scripting expertise.
The summary above was generated by AI

Do you like collaborating across teams to solve complex problems?

Do you enjoy solving large scale distributed content delivery challenges?

Join our highly skilled Compute Site Reliability team

Our team designs, develops, and manages applications and infrastructure that support Akamai's Compute products and services. We specialize in creating solutions that help improve observability and enforce SLAs across all internal teams.

Partner with the best

As a Site Reliability Engineer Senior, you will collaborate across operations teams and application development teams. Together, you will be creating tooling and software that monitors and improves the reliability of our systems. You'll work with a diverse range of technologies as we release new applications and modernize existing tooling

As a Site Reliability Engineer Senior, you will be responsible for:

  • Solving complex problems in a timely and accurate manner through proactive troubleshooting, automation and systems programming
  • Deploying and maintaining our observability platform and internal tooling
  • Partnering across teams to ensure the reliability, scalability and usability of our products and services
  • Providing guidance to engineers and developers to increase confidence that their services are performing as expected
  • Collaborating with our support, operations, and engineering teams to investigate and troubleshoot complex problems

Do what you love

To be successful in this role you will:

  • Have a Bachelor's degree in Computer Science, Engineering
  • Have 6 years of experience in Site Reliability Engineering or a related engineering role, with Linux system administration expertise.
  • Have an understanding of networking fundamentals, including TCP/IP, DNS, routing/switching, and storage concepts.
  • Have hands-on experience with containerized environments and Kubernetes, including operating and troubleshooting production systems.
  • Have knowledge of CI/CD and DevOps practices, with hands-on experience using tools such as Jenkins, Git, Prometheus, and Grafana.
  • Have experience with Infrastructure as Code and configuration management using Terraform, Ansible, SaltStack, or similar tools; exposure to cloud storage systems
  • Have automation/scripting skills using Python, Bash, Go, Rust, or similar languages.

About us

At Akamai, we make life better for billions of people, trillions of times a day.
Whether you're streaming live events, scrolling social media, watching your favorite series, or managing your savings, we're the engine behind the scenes. We provide the world's most distributed platform from Cloud to Edge to help the giants of the digital world work faster and stay more secure, making the internet a better experience for everyone.
Our focus is simple:
Cloud and Edge: Running apps closer to users for instant performance.
Security: Neutralizing threats before they ever reach your data.
Content Delivery: Scaling the world's biggest moments without a glitch.
AI: Enabling our customers to build, secure, and scale AI apps on the world's most distributed cloud platform.
At Akamai, we don't just support the internet; we power and protect it, because behind every great digital experience is a massive hidden challenge. And we're the ones who solve it. When millions of people hit play or pay, Akamai ensures it just works.

Benefits at Akamai: We support your health, well-being, finances, and life beyond work. See our benefits.

FlexBase adapts to your job's needs

Akamai's FlexBase program is yet another way we show our commitment to providing employees with an exceptional workplace experience. It's not about telling employees where to work; it's about supporting employees to do their best work.
We trust our incredible employees to work in ways that suit them best: at home, in an office, or a combination of both.

Connect with us on social and see what life at Akamai is like!

Akamai Technologies Mumbai, Maharashtra, IND Office

Mumbai, India

Similar Jobs

3 Days Ago
Remote or Hybrid
Senior level
Senior level
Digital Media • eCommerce • Gaming • Mobile • News + Entertainment
Lead reliability, scalability, observability, automation, infrastructure, disaster recovery, and security initiatives for Crunchyroll’s cloud-native data platforms. Establish SRE practices including SLIs, SLOs, error budgets, incident management, and postmortems. Operate Kubernetes and GCP environments, implement Infrastructure as Code, optimize capacity and performance, and drive vulnerability remediation, penetration-testing support, and cloud platform security.
Top Skills: Ci/CdDatadogGCPGoGrafanaIdentity And Access ManagementInfrastructure As CodeJavaKubernetesLinuxOpentelemetryOwasp Top 10PrometheusPythonShellTerraform
2 Days Ago
Remote
India
Senior level
Senior level
Information Technology • Productivity • Software • Manufacturing
Own reliability for major AWS production domains by defining SLOs, building observability and automation, managing capacity and self-healing, and leading complex incident response. Design Terraform modules and progressive delivery pipelines, operate ECS, EKS, Lambda, and PostgreSQL workloads, and reduce operational toil. Establish security and compliance controls, apply governed AI to operations, mentor SRE engineers, and standardize reliability practices across global teams.
Top Skills: AWSCi/CdCloudwatchDnsDockerEcs FargateEksGenerative AiGithub ActionsIamIso 27001KubernetesLambdaLlmsOpenobserveOpentelemetryPagerdutyPythonRds PostgresqlSoc 2TerraformVpc
17 Days Ago
In-Office or Remote
India
Senior level
Senior level
Cloud • Security • Software • Cybersecurity
Oversee, scale, and optimize high-density AI hardware infrastructure across regional data centers. Build Python automation and infrastructure-as-code tooling, integrate incident workflows, develop telemetry pipelines and monitoring dashboards, and improve reliability across private cloud, bare-metal, and virtualized environments. Lead on-call incident response, runbooks, post-mortems, service rollouts, vendor coordination, and field technician activities while driving uptime, performance, and operational readiness.
Top Skills: Ai-Based Anomaly DetectionApi IntegrationsBare-Metal InfrastructureBgpGrafanaInfrastructure As CodeIpv4Ipv6LlmsLokiOpentelemetryPagerdutyPrivate CloudPrometheusPythonSlackTelemetry PipelinesVirtualization

What you need to know about the Mumbai Tech Scene

From haggling for the best price at Chor Bazaar to the bustle of Crawford Market, the energy of Mumbai's traditional markets is a key part of the city's charm. And while these markets will always have their place, the city also boasts a thriving e-commerce scene, ranking among the largest in the region. Driven by online sales in everything from snacks to licensed sports merchandise to children's apparel, the local industry is worth billions, with companies actively recruiting to meet the demands of continued growth.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account