Staff Site Reliability Engineer
OktaAbout the role
Get to know Okta
Okta is The World’s Identity Company. We free everyone to safely use any technology—anywhere, on any device or app. Our Workforce and Customer Identity Clouds enable secure yet flexible access, authentication, and automation that transforms how people move through the digital world, putting Identity at the heart of business security and growth.
At Okta, we celebrate a variety of perspectives and experiences. We are not looking for someone who checks every single box - we’re looking for lifelong learners and people who can make us better with their unique experiences.
Join our team! We’re building a world where Identity belongs to you.
The Business Technology Team
This role joins the Business Technology organization and plays a critical part in realizing our vision to accelerate the delivery of business outcomes across Okta by driving clarity, collaboration, and accountability in everything we do. We enable the broader Business Technology organization in the mission to “Accelerate Okta’s Scale and Growth”.
The Site Reliability Engineer Opportunity
We are looking for an experienced Staff BT Site Reliability Engineer to join our Business Technology team to build, improve, and maintain our cloud platform services. The Site Reliability Engineering team builds foundational back-end infrastructure services and tooling for Okta’s corporate teams. We enable teams to build infrastructure at scale and automate their software reliably and predictably. SREs are team players and innovators who build and operate technology using best practices and an agile mindset.
We are looking for a smart, innovative, and passionate engineer for this role, someone who is interested in designing and implementing complex cloud-based infrastructure. This is a lean and agile team, and the ideal candidate welcomes the challenge of building in a dynamic and ever changing environment. They enjoy seeing their designs run at scale with automation, testing, and an excellent operational mindset. If you exemplify the ethics of, "If you have to do something more than once, automate it," we want to hear from you!
What you’ll be doing
- Build and run development tools, pipelines, and infrastructure with a security-first mindset
- Actively participate in Agile ceremonies, write stories, and support team members through demos, knowledge sharing, and architecture sessions
- Promote and apply best practices for building secure, scalable, and reliable cloud infrastructure
- Develop and maintain technical documentation, network diagrams, runbooks, and procedures
- Designing, building, running, and monitoring Okta's IT infrastructure and cloud services
- Driving initiatives to evolve our current cloud platforms to increase efficiency and keep it in line with current standards and best practices
- Recommend, develop, implement, and manage appropriate policy, standards, processes, and procedural updates, especially for highly secure and regulated infrastructure
- Working with software engineers to ensure that development follows established processes and works as intended
- Backup and recovery strategy with the ability to contribute to the business continuity planning
- Provide excellent customer service to our internal users and be an advocate for SRE services and DevOps practices
What you’ll bring to the role
- Proficient in developing applications running on AWS or other cloud infrastructure resources, including compute, storage, networking, and virtualization. Especially critical is knowledge of AWS authentication, governance, and org management suite, including, but not limited to, AWS Orgs, AWS IAM, AWS Identity Center, and Stacksets
- Proficient with automating systems and infrastructure via Terraform
- Proficient with Git and building deployment pipeline using commercial tools, especially Gitlab
- Demonstrated ability to develop complex applications for cloud infrastructure at scale and deliver projects on schedule and within budget
- Experience with monitoring tools, especially Splunk, Cloudwatch, and Grafana
- Experience with reliability engineering concepts and security best practices on public cloud platforms
- Experience with developing tooling and automation in Bash, Python, Go, etc.
- Knowledgeable with Linux system administration skills
- Knowledge of VDI concepts and implementing various VDI systems at scale
- Good communication skills, with the ability to influence others and communicate complex technical concepts to different audiences
- Exposure to FedRAMP, SOC2, FIPS, or other compliance programs is preferred
- US Person Status (e.g., a U.S. Citizen, National, Lawful Permanent Resident, Re
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s