Design and maintain scalable pipelines for visual and voice data, including ingestion, processing, validation, testing, observability, and automation. Build cloud, on-premise, and containerized data infrastructure; optimize storage, retrieval, and caching for large media assets; debug complex data issues; and collaborate with data scientists, operators, and product teams to support machine learning workflows.
As a Data Engineer, you’ll architect and maintain the pipelines that power our products and services. You’ll work at the intersection of ML, media processing, and infrastructure; owning the data tooling and automation layer that enables scalable, high-quality training and inference. If you’re a developer who loves solving tough problems and building efficient systems, we want you on our team.
Key Responsibilities
- Design and maintain scalable pipelines for ingesting, processing, and validating datasets with main focus visual and voice data.
- Work with other teams to identify workflow optimisation potential, design and develop automation tools, using AI-driven tools and custom model integrations and scripts.
- Write and maintain tests for pipeline reliability.
- Build and maintain observability tooling in collaboration with other engineers to track data pipeline health and system performance.
- Collaborate with data scientists, operators, and product teams to deliver data solutions.
- Debug and resolve complex data issues to ensure system performance.
- Optimise storage, retrieval, and caching strategies for large media assets across environments.
- Deploy scalable data infrastructure using cloud platforms as well as on-premise and containerization.
- Deepen your knowledge of machine learning workflows to support AI projects.
- Stay current with industry trends and integrate modern tools into our stack.
Must Haves
- 3+ years in data engineering or related backend/infrastructure role.
- Strong programming skills in Python or similar languages.
- Experience with software development lifecycle (SDLC) and CI/CD pipelines.
- Proven experience building and testing data pipelines in production.
- Proficiency in Linux.
- Solid SQL knowledge.
- Experience with Docker or other containerisation technologies.
- Proactive approach to solving complex technical challenges.
- Passion for system optimisation and continuous learning.
- Ability to adapt solutions for multimedia data workflows.
Nice to Have
- Experience with Kubernetes (k8s).
- Knowledge of machine learning or AI concepts.
- Familiarity with ETL tools or big data frameworks.
- Familiarity with cloud platforms (e.g., AWS, GCP, Azure).
About You
- Innovative
- Like challenges
- Adaptable
- Calm under pressure
- Strong communication abilities
Similar Jobs
Financial Services
Build and maintain Databricks-based data integration and ETL pipelines using Python, PySpark, DLT, Delta Lake, and Unity Catalog. Design governed semantic models and self-service analytics layers, support data governance and access controls, and deliver executive dashboards in Tableau or Sigma. Partner with business stakeholders to translate requirements into actionable insights while maintaining production-quality code, documentation, lineage, and compliance in a regulated environment.
Top Skills:
AlteryxAWSDatabricksDatabricks GenieDatabricks SqlDelta LakeDelta Live TablesGitGithub CopilotJulesProphecyPysparkPythonServicenowSigmaSnowflakeSQLTableauUnity Catalog
Artificial Intelligence • Healthtech • Professional Services • Analytics • Consulting
Leads end-to-end cloud technology consulting projects, delivering scalable data and technology solutions for clients. Responsibilities include project leadership, technology strategy, delivery management, cloud architecture, data management, analytics, process automation, application development, team mentoring, business case development, and global client collaboration. The role also requires troubleshooting database, operating system, and application interactions and traveling as needed for client engagements.
Top Skills:
AgileAmazon RedshiftAWSAzureDatabricksGoogle Cloud PlatformPower BISalesforceSnowflake
Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Lead the architecture and development of scalable batch, streaming, and event-driven data platforms supporting analytics, machine learning, generative AI, and agentic AI. Build governed lakehouse architectures, reusable data products, ingestion and orchestration pipelines, and AI lifecycle capabilities. Drive security, governance, observability, reliability, performance, and cost optimization while translating R&D concepts into production systems. Provide hands-on technical leadership, architecture reviews, mentorship, and cross-functional collaboration.
Top Skills:
Apache AirflowSparkAWSAws Step FunctionsAzure Ai FoundryAzure Data FactoryAzure Machine LearningCi/CdDatabricksDatabricks WorkflowsDockerInfrastructure As CodeKafkaKubernetesMicrosoft FabricMicrosoft PurviewMlflowPysparkPythonSQLTerraformUnity Catalog
What you need to know about the Mumbai Tech Scene
From haggling for the best price at Chor Bazaar to the bustle of Crawford Market, the energy of Mumbai's traditional markets is a key part of the city's charm. And while these markets will always have their place, the city also boasts a thriving e-commerce scene, ranking among the largest in the region. Driven by online sales in everything from snacks to licensed sports merchandise to children's apparel, the local industry is worth billions, with companies actively recruiting to meet the demands of continued growth.



