NVIDIA Logo

NVIDIA

Senior Platform Engineer, Network Infrastructure

Reposted 7 Days Ago
Be an Early Applicant
In-Office or Remote
Hiring Remotely in India
Senior level
In-Office or Remote
Hiring Remotely in India
Senior level
Design, build, and operate a Kubernetes-based platform for global network infrastructure. Own cluster lifecycle, provisioning, upgrades, GitOps delivery, observability, capacity, and recovery. Develop automation, provide production support and on-call incident response for network services, and drive issues from detection through verified resolution.
The summary above was generated by AI
Cloud Foundations Reliability (CFR) is part of NVIDIA’s Global Network Infrastructure (GNI) organization. We deploy, integrate, and operate the Kubernetes-based platform and shared services used to provision, monitor, and operate NVIDIA’s global network across data centers, colocation facilities, and cloud environments. The team owns the architecture and lifecycle of this platform, including cluster provisioning and upgrades, GitOps delivery, observability, capacity, and service enablement. We build software and automation to standardize how network platforms and services are deployed, scaled, and managed across environments.

We are looking for a hands-on senior engineer to own the lifecycle and automation of the Kubernetes platform supporting GNI network systems. You will also provide production support for network services running on the platform, partnering with their engineering owners when issues or changes cross the platform boundary. You will take complex problems from design through production and remain accountable for the outcome. You will bring deep Kubernetes expertise and help establish consistent engineering practices across the US and Bangalore teams. This is a senior individual contributor role with end-to-end ownership and production responsibility.

What You’ll Be Doing:

  • Design, build, and operate the Kubernetes platform that powers GNI network automation, telemetry, and operations across data center, colocation, and cloud environments.

  • Own the lifecycle management for GNI Kubernetes environments, including cluster onboarding, upgrades, capacity, availability, and recovery.

  • Develop production-quality software and automation for cluster provisioning, validation, upgrades, remediation, and safe multi-cluster delivery through GitOps.

  • Provide production support for network services hosted on the platform, working with Network Automation and service teams that retain ownership of application architecture, code, and features.

  • Diagnose complex Kubernetes platform and hosted-service failures involving control-plane health, cluster networking, storage, scheduling, workload placement, and multi-cluster dependencies. Drive issues from initial signal through verified resolution.

  • Define production-readiness and observability standards for the platform and hosted network services, including health signals, capacity, alerts, runbooks, and recovery.

  • Participate in CFR’s production on-call rotation, including scheduled after-hours and weekend coverage. Lead incident response and recovery, then drive corrective actions to completion.

What We Need to See:

  • Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent experience.

  • 8+ years of experience building or operating production Kubernetes platforms, network infrastructure, or distributed systems.

  • Deep experience with Kubernetes at scale, including cluster lifecycle, upgrades, networking, storage, and recovery.

  • Proficiency in at least one general-purpose programming language, such as Go or Python.

  • Experience with GitOps, infrastructure as code, CI/CD, and automated production delivery.

  • Experience deploying and supporting network automation or telemetry services on Kubernetes.

  • Experience with production on-call, incident response, root-cause analysis, and driving corrective actions to completion.

Ways to Stand Out From the Crowd:

  • Strong knowledge of IP routing, data center fabrics, and cloud networking is a great plus.

  • Experience designing and operating large, multi-region Kubernetes fleets, including fleet-wide upgrades and recovery.

  • Hands-on experience with Cluster API (CAPI) and Metal3 for bare-metal provisioning, cluster lifecycle, machine remediation, and upgrades.

  • Experience building Kubernetes controllers or operators in Go using custom resources and reconciliation patterns. Experience designing or operating network automation and telemetry services on Kubernetes at global scale.

  • Contributions to Cluster API, Metal3, or other open-source Kubernetes infrastructure projects.

NVIDIA’s deep learning platforms have made major impact to various fields is broadly used across leading academic institutions, start-ups, and industry, including the world’s largest Internet companies. We need passionate, hard-working and creative people to help us take on more of these unique opportunities in deep learning cloud solutions. NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hard-working people in the world working for us. Are you creative and autonomous? Do you love a challenge? If so, we want to hear from you.

NVIDIA Mumbai, Maharashtra, IND Office

No. 127, Andheri Kurla Road, Village, Chakala, Andheri East, Mumbai, Maharashtra, India, 400093

Similar Jobs

6 Hours Ago
Remote or Hybrid
Senior level
Senior level
Artificial Intelligence • Big Data • Cloud • Information Technology • Software • Big Data Analytics • Automation
Lead end-to-end QA for Salesforce Sales Cloud and CPQ across Quote-to-Cash. Design and execute functional, regression, API, and end-to-end tests; build and maintain ACCELQ automation suites; validate data flows with external systems; support UAT; collaborate with developers, admins, and product owners; troubleshoot defects and ensure release quality.
Top Skills: AccelqApexBrunoCi/Cd PipelinesClaudeDell BoomiFlowsInformatica CloudLightning Web Components (Lwc)Microsoft CopilotMulesoftNetSuiteOauthPlaywrightPostmanProcess BuilderProvarRest ApisRevenue CloudSalesforce BillingSalesforce CpqSalesforce Sales CloudSalesforce Scale CenterSeleniumSlackbot AiSoap ApisSOQLSoslSsoVaricent
8 Hours Ago
Remote
India
Expert/Leader
Expert/Leader
Artificial Intelligence • Big Data • Cloud • Information Technology • Software • Cybersecurity • Data Privacy
Lead Rubrik's engagement with Global Systems Integrators across India to grow strategic partnerships, drive service creation and enablement, execute account and partner business plans, coordinate cross-functional teams, and accelerate partner adoption and bookings. Manage negotiations, build OEM/ISV ecosystems, and travel extensively to engage partners.
Top Skills: AWSCloud PlatformsGCPMicrosoft (Azure)NutanixPure StorageRubrikSaaSSalesforce
9 Hours Ago
Remote or Hybrid
India
Senior level
Senior level
Artificial Intelligence • Cloud • Sales • Security • Software • Cybersecurity • Data Privacy
Design, build, and deploy event-driven automations to eliminate operational runbook toil. Operate and support production databases and streaming platforms, participate in on-call rotation, implement least-privilege IAM and auditability, and lead automation roadmap across data infrastructure.
Top Skills: AlertmanagerApache AirflowAws CloudtrailAws CloudwatchAws EventbridgeAws IamAws LambdaClaudeCursorDynamoDBElasticsearchGithub ActionsGitopsJavaScriptJenkinsKafkaKubernetesMySQLPostgresPrometheusPythonRedisTerraformTypescript

What you need to know about the Mumbai Tech Scene

From haggling for the best price at Chor Bazaar to the bustle of Crawford Market, the energy of Mumbai's traditional markets is a key part of the city's charm. And while these markets will always have their place, the city also boasts a thriving e-commerce scene, ranking among the largest in the region. Driven by online sales in everything from snacks to licensed sports merchandise to children's apparel, the local industry is worth billions, with companies actively recruiting to meet the demands of continued growth.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account