Jobs and Careers
PE

Sr. Manager, System Reliability Engineering - TrainingPeaks

Peaksware
Louisville, United Statesfull_timeVerifiedPosted 3 Nov 2023
💰 $200,815/yr($120,489/yr$200,815/yr)

About the role

Are you ready to work on a product impacting millions of people? At TrainingPeaks our user base of athletes and coaches is growing rapidly. To meet their demands TrainingPeaks needs innovators, collaborators, and excellent engineers like you. Together we’re building the world’s best training platform. Join TrainingPeaks today.

You may know us as TrainingPeaks, MakeMusic, TrainHeroic and Alfred Music. All these brands are under the Peaksware umbrella. TrainingPeaks develops software for coaches and athletes to track, analyze and plan endurance training. TrainHeroic develops software solutions for the strength and conditioning needs of coaches and athletes. MakeMusic develops software to transform how music is composed, taught, learned and performed. Alfred Music creates and publishes educational music to help teachers, students, professionals and hobbyists experience the joy of making music.

We would love to have you join our ever-growing team! All applicants will receive equal consideration for employment regardless of gender, race, national origin, age, sexual orientation, gender identity, physical disability, religion, or length of time spent unemployed.

General Summary

As Sr. Manager, System Reliability Engineering, you are a proven leader in transforming business operations in a cloud-based environment. You and the team you lead will be crucial in ensuring the reliability, availability, and performance of our software applications and infrastructure. Your technical experience in System Reliability Engineering will be instrumental in setting and maintaining high standards for system availability, performance, and incident response. In this role, you will think beyond the scope of your team and look for opportunities to improve the organization at large, collaborating across multiple teams and leading the execution of those cross-team initiatives.

You are a continuous learner with a hunger for knowledge. You approach challenges as opportunities to improve. You value team members’ input from all levels, and you actively seek ways to support your colleagues.

You will sit directly with the System Reliability Engineering team, collaborate closely with Product and Engineering teams, and report to the Vice President, Operations. 

Core Functions

  • Team Leadership: Lead, mentor, and manage a team of Site Reliability Engineers, fostering a culture of collaboration, innovation, and excellence.
  • Collaboration: Foster strong relationships with software development, operations, and product teams to ensure a unified approach to system reliability.
  • Strategy and Planning: Develop and implement SRE strategies, goals, and objectives to achieve high system reliability and availability. Collaborate with cross-functional teams to align SRE objectives with business goals.
  • Incident Management & On-Call Rotation: Oversee incident response and resolution processes, ensuring a swift and effective response to system issues. Continuously improve incident management practices. Develop and manage an on-call rotation schedule for SREs and senior Developers, providing 24/7 support for critical systems.
  • Monitoring and Alerting: Establish and maintain robust monitoring and alerting systems to proactively identify and address system performance issues. Drive automation of monitoring and alerting processes.
  • Security and Compliance: Ensure security and compliance of the cloud-based applications by implementing and maintaining robust security measures, access controls, and monitoring systems to protect sensitive data and infrastructure from threats and vulnerabilities.
  • Capacity Planning: Work closely with  the SRE and development teams to plan and manage system capacity, ensuring systems can handle current and future workloads.
  • Reliability Engineering: Drive the implementation of best practices for reliability engineering, including chaos engineering, fault tolerance, and system resilience.
  • Performance Optimization: Collaborate with engineering teams to optimize application performance and improve system efficiency.

The work characteristics described here are representative of those an employee encounters while performing the essential functions of this job. Reasonable accommodations may be made to enable individuals with disabilities to perform the essential functions.

Requirements

  • Experience managing a collaborative team of System Reliability Engineers.
  • Proven software engineering experience.
  • Prior experience working in a DevOps or SRE environment and continuou

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Peaksware

View company profile →