Principal Site Reliability Engineer (Cloud Management Platform)
Palo Alto NetworksAbout the role
Company Description
Our Mission
At Palo Alto Networks® everything starts and ends with our mission:
Being the cybersecurity partner of choice, protecting our digital way of life.
We have the vision of a world where each day is safer and more secure than the one before. These aren’t easy goals to accomplish – but we’re not here for easy. We’re here for better. We are a company built on the foundation of challenging and disrupting the way things are done, and we’re looking for innovators who are as committed to shaping the future of cybersecurity as we are.
We’re changing the nature of work. Palo Alto Networks is evolving to meet the needs of our employees now and in the future through FLEXWORK, our approach to how we work. From benefits to learning, location to leadership, we’ve rethought and recreated every aspect of the employee experience at Palo Alto Networks. And because it FLEXes around each individual employee based on their individual choices, employees are empowered to push boundaries and help us all evolve, together.
Job Description
Your Career
Palo Alto Networks runs a large infrastructure and is one of the largest GCP customers. As a Site Reliability Engineer, you will be part of a team supporting the services running on this infrastructure. This includes automation, architecture, performance, observability, troubleshooting, security, and reliability.
Our Infrastructure Platform stack includes Terraform, Kubernetes, GitLab CI/CD, GitOps, Prometheus, Grafana, Loki, Docker, GCP, AWS, Vault, Kafka, MySQL, Python, Bash, and Go.
Your Impact
- Contribute to the success of SRE and DevOps
- Develop expertise in new technologies
- Work with developers, researchers, data scientists, and security experts
- Design, build and operate reliable, secure Cloud infrastructure
- Ensure that applications are production-ready, scalable, and reliable
- Develop tools and automation frameworks
- Automate robust deployment of robust services
- Orchestrate end-to-end monitoring and alerting
- Participate with SRE and Dev teams in the on-call rotation
- Lead root cause analysis of critical business and production issues
Qualifications
Your Experience
- 6+ years as a engineer in Infrastructure, Operations, DevOps, or System Engineering
- 3+ years building high availability, scalable cloud native applications on AWS or GCP
- BS or MS in Computer Science, a related field, or equivalent professional experience or equivalent military experience required
- Expertise in configuration management with a framework such as Ansible, Terraform, Helm
- Experience in Site Reliability Engineering, Production Engineering, or DevOps
- Expertise in public or private cloud
- Solid experience in Kubernetes and containers
- Linux administration, internals, and network troubleshooting
- Proficiency with programming languages like Python, Java, Golang, and shell scripting to automate tasks
- Familiarity with CI/CD pipelines, GitLab and GitHub preferred
- Ability to diagnose and troubleshoot complex distributed systems handling high volume transactions
- Excellent written and verbal communication, able to collaborate and rally support
- Self-disciplined, self-managed, self-motivated and strong sense of ownership, urgency, and drive
- Passion for infrastructure and monitoring as code
- Ready to understand and dissect new technology stacks quickly
Additional Information
The Team
Drawing on the near real-time data collected through PAN-OS device telemetry, our industry-leading next generation insights product (AIOps for NFGW) gives large cybersecurity operators a force multiplier that provides visibility into the health of their next-generation-firewall (NGFW) devices. It enables early detection of issues at various levels of the stack via Deep Learning based advanced time-series forecasting and anomaly detection using novel deep learning techniques. Our goal is to be able to predict and prevent service-impacting issues in critical security infrastructure that operates 24/7/365 with zero false positives and zero false negatives.
Our Commitment
We’re trailblazers that dream big, take risks, and challenge cybersecurity’s status quo. It’s simple: we can’t accomplish our mission without diverse teams innovating, together.
We are committed to providing reasonable accommodations for all qualified individuals with a disability. If you require assistance or accommodation due to a disability or special need, please contact us at accommodations@paloaltonetworks.com.
Palo Alto Networks is an equal opportunity employer. We celebrate diversity in our workp
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s