Jobs and Careers
QU
Principal Engineer - Platform Architecture & Core Reliability
QuizletSan Francisco, United Statesfull_timeVerifiedPosted 13 Oct 2025
💰 $320,000/yr($260,000/yr – $320,000/yr)
About the role
About Quizlet:
At Quizlet, our mission is to help every learner achieve their outcomes in the most effective and delightful way. Our $1B+ learning platform serves tens of millions of students every month, including two-thirds of U.S. high schoolers and half of U.S. college students, powering over 2 billion learning interactions monthly.
We blend cognitive science with machine learning to personalize and enhance the learning experience for students, professionals, and lifelong learners alike. We’re energized by the potential to power more learners through multiple approaches and various tools.
Let’s Build the Future of LearningJoin us to design and deliver AI-powered learning tools that scale across the world and unlock human potential.
About the Role:
We're hiring a Principal Engineer to lead critical architectural decisions that establish industry-leading standards for reliability and operational excellence. This is a high-leverage, hands-on role focused on optimizing performance, driving engineering velocity, and leading systemic architectural change. The role reports to the Senior Director of Technical Infrastructure.
We’re happy to share that this is an onsite position in our San Francisco office. To help foster team collaboration, we require that employees be in the office a minimum of three days per week: Monday, Wednesday, and Thursday and as needed by your manager or the company. We believe that this working environment facilitates increased work efficiency, team partnership, and supports growth as an employee and organization.
At Quizlet, our mission is to help every learner achieve their outcomes in the most effective and delightful way. Our $1B+ learning platform serves tens of millions of students every month, including two-thirds of U.S. high schoolers and half of U.S. college students, powering over 2 billion learning interactions monthly.
We blend cognitive science with machine learning to personalize and enhance the learning experience for students, professionals, and lifelong learners alike. We’re energized by the potential to power more learners through multiple approaches and various tools.
Let’s Build the Future of LearningJoin us to design and deliver AI-powered learning tools that scale across the world and unlock human potential.
About the Role:
We're hiring a Principal Engineer to lead critical architectural decisions that establish industry-leading standards for reliability and operational excellence. This is a high-leverage, hands-on role focused on optimizing performance, driving engineering velocity, and leading systemic architectural change. The role reports to the Senior Director of Technical Infrastructure.
We’re happy to share that this is an onsite position in our San Francisco office. To help foster team collaboration, we require that employees be in the office a minimum of three days per week: Monday, Wednesday, and Thursday and as needed by your manager or the company. We believe that this working environment facilitates increased work efficiency, team partnership, and supports growth as an employee and organization.
In this role, you will:
- Platform Reliability (SLO Focus): Lead the strategy and implementation necessary to achieve and maintain our 99.95% availability target. This includes evolving our multi-region deployment strategy and optimizing for resilience under sustained, high-volume traffic
- Data Backbone Scaling: Define the architectural approach for scaling our core data systems, optimizing performance across Cloud Spanner, PlanetScale MySQL, and BigQuery
- Compute & Service Mesh: Drive performance and efficiency improvements across our managed compute environment, specifically optimizing Kubernetes (GKE) clusters and managing the performance and operational complexity of Istio
- Developer Velocity & CI/CD: Architect high-leverage internal platforms, designing the pipelines across tools like GitHub Actions, CircleCI, and ArgoCD, and driving organizational influence to standardize high-velocity, safe deployment practices
- Incident & Learning Culture: Drive reliability change across the engineering organization by leveraging deep-dive analysis of incidents (Jeli) and proactive monitoring (Datadog), improving operational practices, and influencing architectural design decisions
- Performance & Cost Engineering: Act as a technical owner for the cost-per-request metric, identifying and implementing architectural efficiencies (caching, connection pooling, resource utilization) that scale down infrastructure spend while maintaining service speed
What you bring to the table:
- Deep technical mastery and hands-on experience across three or more of the following high-leverage domains with 10+ years of experience in software development, site reliability engineering, or platform engineering
- A proven track record of driving significant architectural outcomes as a Principal or Staff Engineer in a high-scale platform or infrastructure role
- Global Scale & Resilience: A history of scaling consumer-facing systems that reliably handle tens of thousands of requests per second (RPS) and successfully achieving high-availability targets across multi-region cloud environments
- Data Systems: Deep expertise in architecting and optimizing complex data backbones involving transactional, globally consistent (e.g., Cloud Spanner / PlanetScale), and analytical systems (BigQuery).
- Container Orchestration & Networking: Strong operational knowledge of Kubernetes (GKE) orchestration layered with a Service Mesh (Istio), including traffic shaping, security, and performance tuning.
- Developer Velocity Tooling: Proven ability to design and implement automated CI/CD pipelines leveraging tools like GitHub Actions, CircleCI, and ArgoCD, with a track record of successfully deploying AI developer tools and workflows
- Observability & Incident Management: Mastery in implementing comprehensive monitoring using Datadog for defining SLOs and performing deep-dive investigations, and a strong background in leading structured incident response using tools like Jeli
- Cloud Architecture: Extensive experience in designing and optimizing large-scale cloud-native architecture on GCP (or equivalent cloud providers), incl
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s