Sr Mgr, Site Reliability Engineering
Palo Alto NetworksAbout the role
Company Description
Your Career
Palo Alto Networks runs a large infrastructure and is one of the largest GCP customers. As a Senior Manager of Site Reliability Engineering, you will lead a team supporting the services running on this infrastructure. This includes automation, architecture, performance, observability, troubleshooting, security, and reliability.
Our stack includes Terraform, Kubernetes, GitLab CI/CD, GitOps, Prometheus, Grafana, Loki, Docker, GCP, AWS, Vault, Kafka, MySQL, Python, Bash, and Go.
Job Description
Your Impact
- Lead the SRE team in operational excellence
- Define standards and processes around monitoring, incident response, operations
- Define long term vision for the future of managing large scale systems
- Drive improvements in availability, observability, costs, security
- Mentor and train individual contributors of all levels from junior to principal
- Work across teams to define requirements, architect solutions, and drive projects from implementation to completion
- Design, build and operate reliable, cost effective, and secure cloud infrastructure
- Work with application development leaders to plan infrastructure, deployment, scaling, and application changes to support operational improvements
- Define tools and processes to improve developer efficiency
- Work with platform engineering team to build tools and services which can be used across multiple different teams
- Participate in the on-call rotation
- Lead root cause analysis of critical business and production issues
- Lead through dynamic high growth environment
Qualifications
Qualifications and Experience
- 6+ years as a engineer in Infrastructure, Operations, DevOps, or System Engineering
- 3+ years building high availability, scalable cloud native applications on AWS or GCP
- 3+ years directly managing SREs or SRE teams
- BS or MS in Computer Science, a related field, or equivalent professional experience
- Expertise in managing large scale production environments
- Expertise in high scale workload management and systems debugging
- Ability to motivate and inspire a strong SRE team
- Broad knowledge across different areas of the computing stack, from networking to computing, to databases.
- Solid understanding of Kubernetes operations
- Solid understanding of managing databases, both SQL and NoSQL.
- Experience with provisioning and configuration management tools like Terraform, Puppet, Chef, Ansible
- Linux administration, internals, and network troubleshooting
- Proficiency with programming languages like Python, Java, Golang, and shell scripting to automate tasks
- Experience Continuous Integration tools like Gitlab
- Experience building Continuous Deployment with tools like Spinnaker or ArgoCD
- Excellent written and verbal communication, able to collaborate and rally support
Additional Information
Remote
- This position is not eligible for remote work. The candidate will be expected to be in the Santa Clara HQ office 3-5 days per week.
The Team
To stay ahead of the curve, it’s critical to know where the curve is, and how to anticipate the changes we’re facing. For the fastest growing cybersecurity company, the curve is the evolution of cyberattacks, and the products and services that proactively address them. Our engineering team is at the core of our products – connected directly to the mission of preventing cyberattacks. They are constantly innovating – challenging the way we, and the industry, think about cybersecurity. These engineers aren’t shy about creating products to solve problems no one has tackled before. They define the industry, instead of waiting for directions. We need individuals who feel comfortable in ambiguity, excited by the prospect of challenge, and empowered by the unknown risks facing our everyday lives that are only enabled by a secure digital environment.
Our engineering team is provided with an unrivaled opportunity to build the products and practices that will support our company growth over the next decade, defining the cybersecurity industry as we know it. If you see the potential of how incredible people products can transform a business, this is the team for you. If you don’t wait for directions, instead, identifying new features and opportunities we have to just get better, this is your new career.
Our Commitment
We’re trailblazers that dream big, take risks, and challenge cybersecurity’s status quo. It’s simple: we can’t accomplish our mission without diverse teams innovating, together.
We are committed to providing reasonable accommodations for all qua
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s