Principal Site Reliability Engineer
SecurityScorecardAbout the role
About SecurityScorecard:
SecurityScorecard is the global leader in cybersecurity ratings, with over 12 million companies continuously rated, operating in 64 countries. Founded in 2013 by security and risk experts Dr. Alex Yampolskiy and Sam Kassoumeh and funded by world-class investors, SecurityScorecard’s patented rating technology is used by over 25,000 organizations for self-monitoring, third-party risk management, board reporting, and cyber insurance underwriting; making all organizations more resilient by allowing them to easily find and fix cybersecurity risks across their digital footprint.
Headquartered in New York City, our culture has been recognized by Inc Magazine as a "Best Workplace,” by Crain’s NY as a "Best Places to Work in NYC," and as one of the 10 hottest SaaS startups in New York for two years in a row. Most recently, SecurityScorecard was named to Fast Company’s annual list of the World’s Most Innovative Companies for 2023 and to the Achievers 50 Most Engaged Workplaces in 2023 award recognizing “forward-thinking employers for their unwavering commitment to employee engagement.” SecurityScorecard is proud to be funded by world-class investors including Silver Lake Waterman, Moody’s, Sequoia Capital, GV and Riverwood Capital.
Role Overview
As a Principal Site Reliability Engineer, you will play a strategic and technical leadership role in shaping the reliability, scalability, and velocity of our engineering platform. Your primary focus will be advancing our Kubernetes-based infrastructure and CI/CD systems to support high-scale, high-availability services. You will partner with engineering leaders across the organization to define and drive platform-wide initiatives that enable fast, safe, and repeatable deployments, and foster a culture of reliability and operational excellence.
Key Responsibilities
- Lead the design and evolution of Kubernetes-based infrastructure to support multi-tenant, high-scale applications with strong isolation, resilience, and security.
- Architect and optimize CI/CD pipelines to support fast and reliable build, test, and deploy cycles across a polyglot environment.
- Establish and evangelize best practices for GitOps, canary deployments, rollback strategies, and progressive delivery.
- Define and implement scalable Infrastructure as Code (IaC) patterns using tools such as Terraform, Helm, and Crossplane.
- Drive the adoption of automated testing throughout the delivery lifecycle—unit, integration, load, and chaos testing—to ensure high confidence in production changes.
- Guide teams in designing for observability, SLOs, and alerting, ensuring actionable signals and minimizing alert fatigue.
- Partner with security, compliance, and development teams to ensure infrastructure and delivery systems meet modern security and governance standards.
- Lead incident response retrospectives and foster a blameless culture of continuous improvement.
- Mentor and influence senior engineers across multiple teams, helping to up-level platform reliability capabilities organization-wide.
Qualifications
- 8+ years of experience in Site Reliability Engineering, DevOps, or Infrastructure roles, with 2+ years in a technical leadership or principal capacity.
- Deep expertise with Kubernetes internals (controllers, networking, autoscaling, operators, etc.) and production-grade clusters on cloud providers (EKS, GKE, or AKS).
- Proven experience designing and scaling CI/CD systems using tools such as GitHub Actions, Argo CD, Tekton, Spinnaker, or similar.
- Strong proficiency in Terraform and modern IaC practices.
- Advanced knowledge of automated testing strategies, including performance, load, and failure testing.
- Proficient in one or more programming/scripting languages (Python, Go, Bash, etc.).
- Deep experience with monitoring and observability stacks such as Prometheus, Grafana, OpenTelemetry, and Datadog.
- Strong communicator with the ability to align technical initiatives to business objectives and influence across engineering teams.
Nice-to-Have
- Experience implementing multi-cluster or multi-region Kubernetes strategies.
- Exposure to chaos engineering and building resilient distributed systems.
- Familiarity with compliance frameworks (SOC 2,
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s