Jobs and Careers
CO

Sr. Manager, Site Reliability Engineering

Coupa Software, Inc.
Germanyfull_timeVerifiedPosted 30 Apr 2024

About the role

Coupa makes companies operate smarter and grow faster. Our leading AI-driven platform connects and optimizes sourcing, purchasing, supply chains, and financial management. More than 3,000 global organizations large and small trust Coupa to transform operating margins, increase efficiencies and growth, optimize cash, and reduce risk.
Coupa makes companies operate smarter and grow faster. Our leading AI-driven platform connects and optimizes sourcing, purchasing, supply chains, and financial management. More than 3,000 global organizations large and small trust Coupa to transform operating margins, increase efficiencies and growth, optimize cash, and reduce risk. We are looking for a highly talented and innovative Senior Manager to lead our Site Reliability Engineering (SRE) team. You will lead strategy and execution of a technical roadmap that will increase the velocity of delivering products and unlock new engineering capabilities. The ideal candidate has deep technical expertise to improve application performance, capacity benchmarking, improve availability and reliability, design and evolve cloud infrastructure and architecture
Note : This is a Hybrid role which will involve travel to our Karlsruhe office twice per week.

Responsibilities:

  • Lead a team of SRE professionals responsible for maintaining and executing technology programs to improve overall systems and services stability, availability, performance, and service monitoring across all of our platforms
  • Partner with internal stakeholders to develop multi-year roadmaps influencing the direction and evolution of the operating environment and support protocols
  • Have strong technical expertise and leadership, you are able to lead from the trenches and have proven knowledge in your field
  • Identify and raise appropriate project risks, in addition to presenting detailed and implementable solutions or alternatives
  • Establish and maintain Key Performance Indicators for the overall health of the service and build tools to exercise and evaluate KPIs
  • Lead cross-functionally with other teams to surface common pain points, architect solutions, establish conventions and evangelize application development and operations best practices
  • Build relationships with peers and stakeholders to foster cross-functional/team/department collaboration
  • Limiting time spent on operational work, blameless post-mortems and proactive identification of potential outages factor into iterative improvement
  • Manage, scale, and grow a team of SRE professionals by setting performance goals and measuring deliverables to help recruit, source, interview and hire across the entire engineering team
  • Create robust technical solutions for complex business challenges and improve team delivery through Agile practices; help define and communicate technical standards and best practices
  • Manage on-call rotations across continents, using a follow-the-sun mode

Requirements:

  • Experience in leading and supporting a team of DevOps or software engineers, distributed & remote
  • Hands-on DevOps experience using Go, Python, Java, .Net or equivalent
  • Experience building & running modern full-stack cloud applications using public cloud technologies (AWS, Azure, GCP) reliably and at scale
  • Management experience of Linux machines, Windows, Web servers, Application servers, Databases
  • Experience with containers and container orchestration tools (Docker, Kubernetes, EKS, AKS, ECS)
  • Experience with Kafka, MySQL, PostgreSQL, Elasticsearch, and/or Redis
  • Unquenchable thirst for knowing everything within your platform and learning new technologies
  • Experience managing team on large-scale projects with technical deep-dives into code, networking, operating systems, and cloud
  • Knowledge of defining and monitoring system quality measures, including SLIs (Service-level Indicators), SLOs (Service-level Objectives), and Service-level Agreements (SLAs)
  • Built tooling to improve the reliability of systems, automated remediation of issues, or improve scalability
  • Hands-on experience collecting performance data, analyzing, troubleshooting, and tuning
  • Has effective communication skills (written and verbal) to properly articulate complicated problems to all levels of the organization & customers
  • Bachelor’s degree in degree in Computer Science, Computer Engineering or equivalent

Preferred:

  • Experience with modern infrastructure management systems (Chef, Ansible, Terraform)
  • Expertise in building Platform-as-a-Service (PaaS), Infrastructure-as-a-service (IaaS), or Internal Developer Portal (IDP) solutions
#LI-ML1
At Coupa, we’re building a great company that is laser-focused on three core values: ensuring customer success with an obsessive and u

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Coupa Software, Inc.

View company profile →