Jobs and Careers
CA
Site Reliability Engineering Manager (SRM)
Casper LabsNew York City, United StatesRemotefull_timeVerifiedPosted 6 Feb 2024
About the role
The role of a Site Reliability Manager is especially critical due to the complex nature of large-scale software systems and the high demand for reliability and performance. Some of the specific purposes and responsibilities are ensure system reliability, lead / develop and mentor SRE team members, collaborate with development teams, define and monitor SLOs and SLIs and take a hands-on approach to incident management, performance monitoring and optimization, automation and tooling and security and compliance.
Responsibilities:
- Team Leadership: Lead and manage a team of Site Reliability Engineers (SREs) in maintaining the operational aspects of our software infrastructure. Foster a culture of collaboration, reliability, and continuous improvement within the SRE teams.
- Collaboration with Development Teams: Work closely with software development teams to integrate reliability practices into the software development life cycle. Ensure a seamless collaboration between SRE and development teams (DevOps).
- Incident Management: Oversee incident management activities, including coordinating responses to incidents, leading post-incident reviews, and implementing preventive measures.
- Automation and Tooling: Implement automation and tooling to streamline operational processes, including deployment, monitoring, and recovery.
- Performance Monitoring and Optimization: Monitor and optimize the performance of enterprise software systems, collaborating with development teams to address performance bottlenecks.
- Capacity Planning: Conduct capacity planning to ensure that our infrastructure can handle current and future workloads.
- Security and Compliance: Collaborate with security teams to ensure the security and compliance of our software systems, implementing security best practices and conducting regular audits.
- Communication and Reporting: Effectively communicate with executive leadership and stakeholders, providing updates on system reliability, improvement initiatives, and potential risks.
- Define and Monitor SLOs and SLIs: Set and monitor Service Level Objectives (SLOs) and Service Level Indicators (SLIs) to measure and maintain the reliability and performance of our software systems.
Requirements
- Bachelor's degree in Computer Science, Information Technology, or a related field.
- Proven experience in a leadership role overseeing Site Reliability Engineering teams in an enterprise software development environment.
- Strong background in software development and operations, with a focus on reliability, scalability, and performance.
- Experience with incident management, automation, and performance monitoring tools.
- Excellent communication and interpersonal skills, with the ability to collaborate effectively with cross-functional teams.
- Knowledge of security best practices and experience ensuring compliance with industry standards.
- Experience in one or more of the following: C, C++, Rust, Java, Python, Go, Perl, Ruby or shell scripting.
- Experience with Unix/Linux operating systems internals and administration (e.g., filesystems, inodes, system calls) or networking (e.g., TCP/IP, routing, network topologies and hardware, SDN).
- Experience with monitoring and aggregating systems, such as Prometheus, Graphite, etc, and visualizing systems such as Grafana.
- Experience with CI/CD systems such as Travis, Harness, Drone, CircleCI, etc.
- Experience with version control systems such as Git, Perforce, etc.
- Experience with log aggregator systems such as ELK, Splunk, etc.
- Experience with configuration management systems such as Puppet, Ansible, Chef, Salt, etc.
- Expertise in designing, analyzing and troubleshooting large-scale distributed systems.
- Experience with a cloud based infrastructure platform (i.e. AWS).
- Ability to debug and optimize code and automate routine tasks.
- Experience with managing a fully remote workforce with time zone variation.
Benefits
- Fully remote, work from home environment
- Flexible working hours
- Paid Time-Off
- Periodic in-person offsites globally (travel permitting)
- Long-term incentive programs
- Continued education support
- Advancement opportunity
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s