Staff Site Reliability Engineer
ZscalerAbout the role
Company Description
Zscaler (NASDAQ: ZS) accelerates digital transformation so that customers can be more agile, efficient, resilient, and secure. The Zscaler Zero Trust Exchange is the company’s cloud-native platform that protects thousands of customers from cyberattacks and data loss by securely connecting users, devices, and applications in any location.
With more than 10 years of experience developing, operating, and scaling the cloud, Zscaler serves thousands of enterprise customers around the world, including 450 of the Forbes Global 2000 organizations. In addition to protecting customers from damaging threats, such as ransomware and data exfiltration, it helps them slash costs, reduce complexity, and improve the user experience by eliminating stacks of latency-creating gateway appliances.
Zscaler was founded in 2007 with a mission to make the cloud a safe place to do business and a more enjoyable experience for enterprise users. Zscaler’s purpose-built security platform puts a company’s defenses and controls where the connections occur—the internet—so that every connection is fast and secure, no matter how or where users connect or where their applications and workloads reside.
Zscaler enables the world’s leading organizations to securely transform their networks and applications for a mobile and cloud-first world. Applications have moved from the data center to the cloud and users are connecting to their workloads from everywhere, but security has remained anchored to the data center. Zscaler is redefining security by moving it out of the data center and into the cloud.
The Zscaler Zero Trust Exchange uses software-defined business policies, not appliances, to securely connect the right user to the right application, regardless of device, location, or network. Zscaler operates 4 pillars of Trust Exchange. Zscaler Internet Access™ which scans every byte of traffic to ensure that nothing bad comes in and nothing good leaks out. Zscaler Private Access™ offers authorized users secure and fast access to internal applications hosted in the data center or public clouds—without a VPN. Zscaler Cloud Workload Protection, to identify and remediate risks associated with customer’s cloud infrastructure. Zscaler Workload Segmentation provides micro-segmentation of processes across multiple systems with Machine Learning based policy.
Zscaler services are 100% cloud delivered and offer the simplicity, enhanced security, and improved user experience that traditional appliances or hybrid solutions are unable to match. Used in more than 185 countries, the Zscaler multi-tenant, distributed security cloud protects thousands of customers from cyberattacks and data loss, enabling customers to embrace the agility, speed, and cost containment of the cloud—securely.
Job Description
- Deploy, support and troubleshoot multiple large-scale distributed software applications and networks
- Lead and Coordinate the response to critical incidents to support reliability targets
- Be an escalation contact for critical service incidents
- Communicate effectively with engineering teams and executives to provide status updates and drive urgency
- Determine the situation, options, and risks and make timely decisions to facilitate incident resolution
- Drive Post-Incident Analysis to ensure a thorough understanding of failures and identify critical follow-ups
- Create and deploy scalable monitoring and control tools and processes to support a highly available global infrastructure.
- Design, code and deploy automation to eliminate operational toil and improve reliability.
- Champion best practices for reliability within Engineering Department
- Participate and/or lead projects from intake to closure
Qualifications
- Strong communication and interpersonal skills
- Decisiveness and ability to make sound judgments under pressure
- Problem-solving and analytical skills
- Strong hands-on Linux system administration experience (7+ years).
- Experience in operations or site reliability engineering supporting large scale production environments with high uptime requirements
- Hands-on experience with ansible, coding language such as Python, Go, etc (7+ years)
- Experience in configuration and maintenance of applications such as web servers, load balancers, relational databases, storage systems and network devices
- Experience with large-scale commercial networking and security applications is a plus.
Bachelor's degree in Computer Science, a related technical field involving computer systems engineering, or equivalent practical experience.
Additional Information
What You Can Expect From Us
- An environment where you will be working on cutting edge technologies and architectures
- A fun, passionate and collaborative
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s