Senior Manager Site Reliability Engineering (Cortex Cloud Infrastructure)
Palo Alto NetworksAbout the role
Company Description
Our Mission
At Palo Alto Networks® everything starts and ends with our mission:
Being the cybersecurity partner of choice, protecting our digital way of life.
We have the vision of a world where each day is safer and more secure than the one before. These aren’t easy goals to accomplish – but we’re not here for easy. We’re here for better. We are a company built on the foundation of challenging and disrupting the way things are done, and we’re looking for innovators who are as committed to shaping the future of cybersecurity as we are.
We’re changing the nature of work. Palo Alto Networks is evolving to meet the needs of our employees now and in the future through FLEXWORK, our approach to how we work. From benefits to learning, location to leadership, we’ve rethought and recreated every aspect of the employee experience at Palo Alto Networks. And because it FLEXes around each individual employee based on their individual choices, employees are empowered to push boundaries and help us all evolve, together.
Job Description
Your Career
Palo Alto Networks runs a large infrastructure and is one of the largest GCP customers. As a Senior Manager of Site Reliability Engineering, you will lead a team supporting the services running on this infrastructure. This includes automation, architecture, performance, observability, troubleshooting, security, and reliability.
Our stack includes Terraform, Kubernetes, GitLab CI/CD, GitOps, Prometheus, Grafana, Loki, Docker, GCP, AWS, Vault, Kafka, MySQL, Python, Bash, and Go.
Your Impact
- Lead the SRE team in operational excellence
- Define standards and processes around monitoring, incident response, operations
- Define long term vision for the future of managing large scale systems
- Drive improvements in availability, observability, costs, security
- Mentor and train individual contributors of all levels from junior to principal
- Work across teams to define requirements, architect solutions, and drive projects from implementation to completion
- Design, build and operate reliable, cost effective, and secure cloud infrastructure
- Work with application development leaders to plan infrastructure, deployment, scaling, and application changes to support operational improvements
- Define tools and processes to improve developer efficiency
- Work with platform engineering team to build tools and services which can be used across multiple different teams
- Participate in the on-call rotation
- Lead root cause analysis of critical business and production issues
- Lead through dynamic high growth environment
Qualifications
Your Experience
- 6+ years as a engineer in Infrastructure, Operations, DevOps, or System Engineering
- 3+ years building high availability, scalable cloud native applications on AWS or GCP
- 3+ years directly managing SREs or SRE teams
- BS or MS in Computer Science, a related field, or equivalent professional experience or equivalent military experience required
- Expertise in managing large scale production environments
- Expertise in high scale workload management and systems debugging
- Ability to motivate and inspire a strong SRE team
- Broad knowledge across different areas of the computing stack, from networking to computing, to databases
- Solid understanding of Kubernetes operations
- Solid understanding of managing databases, both SQL and NoSQL
- Experience with provisioning and configuration management tools like Terraform, Puppet, Chef, Ansible
- Linux administration, internals, and network troubleshooting
- Proficiency with programming languages like Python, Java, Golang, and shell scripting to automate tasks
- Experience Continuous Integration tools like Gitlab
- Experience building Continuous Deployment with tools like Spinnaker or ArgoCD
- Excellent written and verbal communication, able to collaborate and rally support
- This position is not eligible for remote work - The candidate will be expected to be in the Santa Clara HQ office 3-5 days per week
Additional Information
The Team
To stay ahead of the curve, it’s critical to know where the curve is, and how to anticipate the changes we’re facing. For the fastest growing cybersecurity company, the curve is the evolution of cyberattacks, and the products and services that proactively address them. Our engineering team is at the core of our products – connected directly to the mission of preventing cyberattacks. They are constantly innovating – challenging the way we, and the industry, think about cybersecurity. These engineers aren’t shy about creating products to solve problems no one ha
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s