Jobs and Careers
PA

Sr Principal Site Reliability Engineer (Advanced Threat Protection)

Palo Alto Networks
Santa Clara, United Statesfull_timeVerifiedPosted 15 Mar 2024
💰 $275,000/yr($170,000/yr$275,000/yr)

About the role

Company Description

Our Mission

At Palo Alto Networks® everything starts and ends with our mission:

Being the cybersecurity partner of choice, protecting our digital way of life.

Our vision is a world where each day is safer and more secure than the one before. We are a company built on the foundation of challenging and disrupting the way things are done, and we’re looking for innovators who are as committed to shaping the future of cybersecurity as we are.

Our Approach to Work

We lead with flexibility and choice in all of our people programs. We have disrupted the traditional view that all employees have the same needs and wants. We offer personalization and offer our employees the opportunity to choose what works best for them as often as possible - from your wellbeing support to your growth and development, and beyond!

At Palo Alto Networks, we believe in the power of collaboration and value in-person interactions. This is why our employees generally work from the office three days per week, leaving two days for choice and flexibility to work where you feel most effective. This setup fosters casual conversations, problem-solving, and trusted relationships. While details may evolve, our goal is to create an environment where innovation thrives, with office-based teams coming together three days a week to collaborate and thrive, together!

Job Description

Your Career

Palo Alto Networks has been rapidly moving towards the future where cloud-based applications are increasingly common. As a Site Reliability Engineer, you will develop the frameworks and pathways to help move our internal applications to microservices. You will be a critical link between engineering and the Infrastructure Platform, building Infrastructure as Code and working in partnership with the App developers to deploy the applications in GCP, AWS and data centers across the globe.

As a member of the SRE team, you will work on producing mission-critical platforms, tools, and processes that will ensure the highest levels of availability and reliability of all our applications. We need creative and innovative problem solvers who can partner with our Application development teams to make their services more usable. Our SRE team is furnished with a standout opportunity to build tools, frameworks, and cloud platforms that will support our company’s growth over the next decade. If you are a self-starter and jump on new ideas to make the platform more stable, secure and feature-rich, this is your new career.

Your Impact

  • Write automation code for provisioning and operating infrastructure at massive scale
  • Design, build and operate Cloud infrastructure to enable reliable and rapid deployment of microservices with effective monitoring and resilient operations
  • Work with development teams to make sure the applications are production ready, scalable and reliable from the grounds up
  • Identify and drive opportunities to improve automation for code deployment, management, and visibility of application services
  • Develop tools and framework to automate operational tasks, deployment of machines, services, applications
  • Establish end-to-end monitoring and alerting on all critical components of the application
  • Participate in the on-call rotation supporting the platform and or the production application
  • Directs root cause analysis of critical business and production issues
  • Develop and mentor other SREs on standard methodology from Infra orchestration and troubleshooting application service in production
  • Represent SRE in design reviews and work cross-functionally with Engineering teams on operational readiness

Qualifications

Your Experience

  • Expertise in configuration management with a framework such as Terraform, Ansible, and Helm
  • Strong Linux administration, internals, and network troubleshooting
  • Experience in DevOps, Site Reliability, or infrastructure engineering 
  • Expertise in Google cloud computing (GCP) and its related services
  • Proficiency with a programming language like Python and shell scripting to automate tasks
  • Strong experience with CI/CD pipeline, GitHub, Jenkins, Artifactory 
  • Ability to diagnose and troubleshoot complex distributed systems handling high volume transactions
  • Strong fundamentals in HTTP including HTTP headers and web servers 
  • BS or MS in Computer Science, a related field, or equivalent professional experience or equivalent military experience required
  • Excellent problem solving, critical thinking, communication, and teamwork skills
  • Excellent written and verbal communication, able to collaborate and rally support
  • Self-disciplined, self-managed, self-motivated and strong sense of owner

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Palo Alto Networks

View company profile →