Sr. Manager, Site Reliability Engineering
IllumioAbout the role
Onwards Together!
Illumio is the leader in ransomware and breach containment, redefining how organizations contain cyberattacks and enable operational resilience. Powered by the Illumio AI Security Graph, our breach containment platform identifies and contains threats across hybrid multi-cloud environments – stopping the spread of attacks before they become disasters.
Recognized as a Leader in the Forrester Wave™ for Microsegmentation, Illumio enables Zero Trust, strengthening cyber resilience for the infrastructure, systems, and organizations that keep the world running.
Location: 5 on-site days a week in Sunnyvale, CA Headquarters.
Our Team's Vision:
Our Engineering team is shaping the future of cybersecurity. We thrive on visionary leadership, autonomy, and ownership, fostering a culture of innovation that propels us forward in the ever-evolving cybersecurity landscape.
As a leader in Zero Trust Segmentation, we are redefining security for a world facing unprecedented cyber threats. You’ll work with a highly scalable SaaS service built using cloud-native technologies while simultaneously shipping the solution on-premises.
Our guiding philosophy in Engineering is to get things right through practicing disciplined engineering, focusing, not cutting corners, and of course having fun while we are at it. We believe in enabling ownership at all levels of the organization and empowering teams. If you thrive in this culture, come join us!
Your Impact:
In this role, you will lead a team of talented engineers to help build a world-class SaaS security platform so we can continue to provide quality security solutions for our customers.
Every day you will lead a small team to ensure our SaaS security platform is available and performing, finding problems before our customers do, building tools to improve speed, confidence, and visibility, while embedding security into every step of the software and infrastructure life cycle.
To thrive in this role you must have at least 5 years of people leadership experience; be fluent in AWS/Azure cloud platforms and have programming language experience while hands-on building infrastructure tooling and automation at least 50% of the time.
Manage Illumio’s SRE team to deliver SaaS security products to companies including the Fortune 100.
Work closely with Development, QA, Customer success, and Technical Support to ensure the health of our products and that all SLAs are being met for our customers.
Manage infrastructure to scale globally, utilizing automation tools to maximize operational efficiency on public clouds.
Lead the team responsible for supporting the infrastructure that powers the Illumio SaaS products.
Own and improve the SaaS delivery efficiency with end-to-end responsibility for application lifecycle, availability, performance, and SLAs.
Work with the team to improve CI/CD automation for deploying applications using Infrastructure as Code to minimize downtime and ensuring adherence to any contractual commitments.
Work with senior management in developing a long-term product reliability and technology road map, using strategies to align with business objectives and large scale.
Develop and maintain automation which can be consumed by multiple teams to deploy SaaS clusters.
Create a high-performing Cloud Operations and SRE team through career development, mentorship, and training
Ensure adherence of infrastructure and processes to FedRAMP, SOC2 and other requirements, and work with PM to adjust the infrastructure to meet any federal changes.
Continue to evolve product architecture and DevOps processes to ensure reliable CI/CD pipelines and continuous delivery.
Your Toolkit:
Bachelor's or Master’s degree in Computer Engineering, Computer Science, or related field, or equivalent relevant experience
5+ years of Unix or Linux system administration experience with Chef/Ansible, Ruby and/or Python
7+ years of hands-on technical experience managing/developing CI/CD solution using Concourse, Gitlab, or equivalent
Proven track record of improving uptime (at least 99.9%) and SLAs.
Experience establishing support SLOs.
Working experience with Cloud Technology such as application gateway, HAProxy, Nginx, microservices, databases(Postgresql), Redis
Experience building observability for 24/7 monitoring with Grafana, Prometheus, Splunk, Datadog, or equivalent and ability to improve uptime and meet SLAs.
Experience building/running revenue generating enterprise applications in at least one of the three big public cloud providers: AWS(Pref
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s