Quantiphi Logo

Quantiphi

Senior MLOps Engineer

Posted 7 Days Ago
Be an Early Applicant
3 Locations
Mid level
3 Locations
Mid level
The Senior MLOps Engineer will design and manage distributed systems using Kubernetes and Slurm, focusing on Multi-GPU, Multi-Node Deep Learning job scheduling. Responsibilities include scripting for system automation, monitoring performance, troubleshooting issues, and collaborating on project requirements.
The summary above was generated by AI

While technology is the heart of our business, a global and diverse culture is the heart of our success. We love our people and we take pride in catering them to a culture built on transparency, diversity, integrity, learning and growth.
If working in an environment that encourages you to innovate and excel, not just in professional but personal life, interests you- you would enjoy your career with Quantiphi!

Role: Senior Platform Engineer (MLOps)

Experience Level: 3 to 6 Years 

Location: Mumbai/Bangalore (Hybrid)

 

Overview:

We are seeking an experienced Platform Engineer with expertise in MLOps/LLMOps and handling distributed systems, particularly Kubernetes and Slurm, along with a strong background in managing Multi-GPU, Multi-Node Deep Learning job scheduling. Proficiency in Linux (Ubuntu) systems, the ability to create intricate shell scripts, good proficiency in working with configuration management tools and sufficient understanding of deep learning workflow.

Roles and Responsibilities:

  • Design, deploy, and maintain distributed systems using Kubernetes and Slurm for optimal resource utilization and workload management.

  • Lead the configuration and optimization of Multi-GPU, Multi-Node Deep Learning job scheduling, ensuring efficient computation and data processing.

  • Collaborate with cross-functional teams to understand project requirements and translate them into technical solutions.

  • Develop and maintain complex shell scripts for various system automation tasks, enhancing efficiency and reducing manual intervention.

  • Monitor system performance, identify bottlenecks, and implement necessary adjustments to ensure high availability and reliability.

  • Troubleshoot and resolve technical issues related to the distributed system, job scheduling, and deep learning processes.

  • Stay updated with industry trends and emerging technologies in distributed systems, deep learning, and automation.

Skill Set Needed: 

  • Hands-on experience in MLOps - Azure ML (preferred), MLFlow, Kubeflow, AutoML etc.

  • Good to have at least one ML framework understanding - PyTorch / TensorFlow.

  • Good with Python. Experience in shell/linux scripting.

  • Good understanding of logical networks. 

  • Proven experience in designing, deploying, and managing distributed systems, with a focus on Kubernetes and Slurm.

  • Sufficient understanding of AI Model Training and Deployment and Strong background in Multi-GPU, Multi-Node Deep Learning job scheduling and resource management.

  • Proficiency in Linux systems, particularly Ubuntu, and the ability to navigate and troubleshoot related issues.

  • Extensive experience creating complex shell scripts for automation and system orchestration.

  • Familiarity with continuous integration and deployment (CI/CD) processes.

  • Excellent problem-solving skills and the ability to diagnose and resolve technical issues promptly.

  • Strong communication and collaboration skills to work effectively within a cross-functional team.

Good to Have:

  • Previously working on NVIDIA Ecosystem or well aware of NVIDIA Ecosystem - Triton Inference Server, CUDA. Experience in working with On-prem NVIDIA GPU servers.

If you like wild growth and working with happy, enthusiastic over-achievers, you'll enjoy your career with us!

Top Skills

Automl
Azure Ml
Kubeflow
Mlflow
Python
PyTorch
TensorFlow

Similar Jobs

2 Days Ago
Pune, Maharashtra, IND
Senior level
Senior level
Healthtech • Biotech • Pharmaceutical
The Senior MLOps Engineer at Roche will research and implement MLOps tools and frameworks, enhance the MLOps maturity in the organization, and provide internal training on MLOps benefits. The role involves operationalizing Data Science projects and applying Agile methodologies in the team.
Top Skills: Python
33 Minutes Ago
Hybrid
Mumbai, Maharashtra, IND
Senior level
Senior level
Financial Services
As a Lead Software Engineer, you will design and implement scalable data pipelines using Python and PySpark on AWS, manage a team of data engineers, collaborate with stakeholders, ensure data quality, and continuously improve technical solutions.
Top Skills: PysparkPython
3 Hours Ago
Hybrid
Mumbai, Maharashtra, IND
Junior
Junior
Financial Services
The Software Engineer II is responsible for enhancing, designing, and delivering software components while executing standard software solutions, developing code, and troubleshooting technical issues. The role involves working within an agile team and applying knowledge of tools in the Software Development Life Cycle to improve automation and stability.
Top Skills: Java

What you need to know about the Mumbai Tech Scene

From haggling for the best price at Chor Bazaar to the bustle of Crawford Market, the energy of Mumbai's traditional markets is a key part of the city's charm. And while these markets will always have their place, the city also boasts a thriving e-commerce scene, ranking among the largest in the region. Driven by online sales in everything from snacks to licensed sports merchandise to children's apparel, the local industry is worth billions, with companies actively recruiting to meet the demands of continued growth.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account