Amtech Software Logo

Amtech Software

Senior Site Reliability Engineer

Posted 9 Hours Ago
Be an Early Applicant
Remote
Hiring Remotely in India
Senior level
Remote
Hiring Remotely in India
Senior level
Own reliability for major AWS production domains by defining SLOs, building observability and automation, managing capacity and self-healing, and leading complex incident response. Design Terraform modules and progressive delivery pipelines, operate ECS, EKS, Lambda, and PostgreSQL workloads, and reduce operational toil. Establish security and compliance controls, apply governed AI to operations, mentor SRE engineers, and standardize reliability practices across global teams.
The summary above was generated by AI

About Vista Equity Partners

Vista Equity Partners is a leading global investment firm focused exclusively on enterprise software, data, and technology-enabled businesses. With over $100B in assets under management and a portfolio of 90+ software product companies worldwide, Vista accelerates growth through operational excellence, shared expertise, and long-term partnership. In India, Vista’s presence continues to expand with 45+ portfolio companies employing more than 17,000 professionals across technology, product, customer success, and operations — reinforcing India’s strategic role as a hub of innovation and talent within the Vista ecosystem.

Through its Agentic AI Factory, Vista is embedding Generative AI across its global portfolio — enabling companies to integrate intelligent, responsible AI into products, operations, and decision-making. This initiative is strengthened through portfolio-wide learning programs, leadership workshops, and AI hackathons that foster innovation, build fluency, and accelerate practical AI adoption across teams.

About Amtech

Amtech is a leading provider of enterprise software solutions for the packaging, printing, and manufacturing industries. Our integrated systems streamline order management, production planning, scheduling, inventory, and business analytics — empowering customers to drive efficiency, reduce costs, and improve operational performance. With a strong commitment to innovation and customer success, Amtech delivers reliable technology backed by deep industry expertise.

With Vista’s investment and strategic guidance, we combine the agility of a growing technology organization with the scale, stability, and career mobility of a global software ecosystem.

Our Employee Value Proposition
At Amtech, our people are our greatest differentiator. We create an environment where you can:

Purpose
Shape the future of manufacturing and supply chain operations by delivering mission-critical enterprise software used by industry-leading organizations.

Growth
Access continuous learning, leadership development, and cross-portfolio opportunities through Vista’s global network — accelerating both technical and managerial career paths.

Culture
Work in a collaborative, transparent, and people-first environment where values, accountability, and integrity guide every decision.

Innovation
Engage with cutting-edge technologies, including AI-driven automation, and contribute to modernizing financial systems and operational processes across the business.
Role Description
Amtech is scaling its Platform Engineering organization as our products move to a fully AWS-hosted, multi-tenant SaaS model. Our SRE practice keeps Encore, LabelTraxx, and supporting platforms reliable, observable, and secure across a multi-account AWS estate. As a Senior Platform Engineer on the SRE track, you will own reliability for major production domains end to end: defining SLOs, building the observability and automation that defend them, and leading incident command for the most complex events. You will shape Amtech's global reliability standards, mentor the SRE team, and partner with Cloud-track engineers to make new platforms operable from day one.

KEY RESPONSIBILITIES

Reliability & Performance

  • Own SLIs, SLOs, and error budgets for major production domains, and drive engineering priorities from them.
  • Design the monitoring, alerting, and reliability frameworks the global team standardizes on (OpenTelemetry, OpenObserve, CloudWatch, PagerDuty).
  • Engineer capacity management, autoscaling, and self-healing so systems recover without human intervention.
  • Lead root cause analysis for the highest-severity incidents and verify permanent fixes land.

Automation & Platform Engineering

  • Design reusable Terraform modules, deployment patterns, and progressive delivery (blue/green, canary, automated rollback) in GitHub Actions.
  • Operate and optimize ECS Fargate, EKS, Lambda, and RDS PostgreSQL workloads at production scale.
  • Set the toil-reduction agenda: quantify operational load and eliminate it through engineering.
  • Build reliability into the Encore-on-AWS and LabelTraxx platforms as they scale customer counts.

Incident Response & On-Call

  • Serve as senior incident commander for cross-service, customer-impacting incidents.
  • Own the on-call program's health for your domains: escalation policies, alert quality, and rotation sustainability in PagerDuty.
  • Drive game days and failure testing to validate runbooks and recovery paths.

Security & Compliance

  • Engineer security into reliability tooling: IAM boundaries, secrets, and network controls.
  • Ensure operations satisfy SOC 2 and ISO 27001 obligations with automated audit evidence.

AI Competency

  • Apply AI in operations and incident response with risk tiering: AI assistance for triage, analysis, and hypothesis generation; human gates for production changes.
  • Design validation gates and guardrails for AI-assisted operational changes, including rollback and audit trails.
  • Measure whether AI tooling improves reliability work against baselines (MTTR, rework, defect escape), and enforce data classification policy in all AI-assisted work.

Technical Leadership

  • Mentor SRE I-III engineers through design reviews, paired incident response, and career coaching input.
  • Set and document global reliability standards adopted across U.S. and India teams.
  • Represent reliability in architecture reviews and migration planning.

QUALIFICATIONS

  • 6+ years of hands-on SRE, DevOps, or platform engineering experience, including ownership of customer-facing production services.
  • Deep AWS operational expertise (ECS/EKS, Lambda, RDS, IAM, VPC, multi-account organizations).
  • Strong Terraform and CI/CD (GitHub Actions or equivalent) skills, including progressive delivery patterns.
  • Proven software engineering ability in Python (or similar) for automation and tooling at team scale.
  • Demonstrated incident command experience for high-severity, multi-service incidents.
  • Experience designing and operating SLO-driven observability with OpenTelemetry or equivalent.
  • Strong grasp of networking, DNS, and cloud security architecture.
  • Demonstrated risk-tiered AI usage in operations: AI-assisted triage with human decision gates, and measured impact on reliability outcomes.
  • Bachelor's degree in Computer Science, Engineering, or a related discipline, or equivalent demonstrated skills.

PREFERRED QUALIFICATIONS

  • AWS Certified DevOps Engineer Professional or CKA.
  • Experience with multi-tenant SaaS or account-per-customer architectures and ERP-class workloads.
  • PagerDuty administration at scale (service ownership models, escalation design).
  • Experience operating AI/LLM workloads or building agentic automation under governance controls.
  • Prior tech-lead experience in a distributed global team.

Why Join Amtech

At Amtech, you will drive meaningful financial impact in a growing enterprise software organization while benefiting from Vista’s world-class ecosystem. You’ll collaborate with talented peers, leverage cross-portfolio learning programs, and help shape the future of Amtech’s financial operations and systems. Build your career with Amtech — backed by the strength, scale, and innovation culture

Similar Jobs

Yesterday
Remote or Hybrid
Senior level
Senior level
Digital Media • eCommerce • Gaming • Mobile • News + Entertainment
Lead reliability, scalability, observability, automation, infrastructure, disaster recovery, and security initiatives for Crunchyroll’s cloud-native data platforms. Establish SRE practices including SLIs, SLOs, error budgets, incident management, and postmortems. Operate Kubernetes and GCP environments, implement Infrastructure as Code, optimize capacity and performance, and drive vulnerability remediation, penetration-testing support, and cloud platform security.
Top Skills: Ci/CdDatadogGCPGoGrafanaIdentity And Access ManagementInfrastructure As CodeJavaKubernetesLinuxOpentelemetryOwasp Top 10PrometheusPythonShellTerraform
Yesterday
Remote or Hybrid
India
Senior level
Senior level
Automotive
Own and optimize Elastic-based observability for a mission-critical 3DX PLM platform. Design monitoring, alerting, health checks, dashboards, and custom Vega visualizations; analyze performance and capacity; manage Elastic clusters, ILM, Fleet, APM, access controls, and query optimization. Automate operational tasks using Python, Bash, or Go, collaborate with infrastructure and engineering teams, and maintain documentation, runbooks, and incident knowledge.
Top Skills: AnsibleAzureBashDassault Systemes 3DxDynatraceElastic StackElasticsearch Query Language (Es|Ql)GCPGoKibanaKibana Query Language (Kql)PythonSiemens TeamcenterTerraformVegaVega-Lite
6 Days Ago
In-Office or Remote
India
Senior level
Senior level
Cloud • Security • Software • Cybersecurity
The Senior Site Reliability Engineer improves the reliability, scalability, availability, and performance of distributed content delivery systems. Responsibilities include defining SLOs and SLIs, monitoring platforms, debugging incidents, implementing corrective actions, automating operational processes, participating in design reviews, and guiding scalable infrastructure design. The role collaborates with Product and Engineering teams and applies software engineering, systems administration, cloud, DevOps, and SRE practices.
Top Skills: AdbmsBashCloud ComputingDatadogDevOpsGrafanaJavaScriptOracle SqlPrometheusPythonUnix/Linux

What you need to know about the Mumbai Tech Scene

From haggling for the best price at Chor Bazaar to the bustle of Crawford Market, the energy of Mumbai's traditional markets is a key part of the city's charm. And while these markets will always have their place, the city also boasts a thriving e-commerce scene, ranking among the largest in the region. Driven by online sales in everything from snacks to licensed sports merchandise to children's apparel, the local industry is worth billions, with companies actively recruiting to meet the demands of continued growth.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account