Principal Site Reliability Engineer
UpstartAbout the role
About Upstart
Upstart is the leading AI lending marketplace partnering with banks and credit unions to expand access to affordable credit. By leveraging Upstart's AI marketplace, Upstart-powered banks and credit unions can have higher approval rates and lower loss rates across races, ages, and genders, while simultaneously delivering the exceptional digital-first lending experience their customers demand. More than 80% of borrowers are approved instantly, with zero documentation to upload.
Upstart is a digital-first company, which means that most Upstarters live and work anywhere in the United States. However, we also have offices in San Mateo, California; Columbus, Ohio; and Austin, Texas.
Most Upstarters join us because they connect with our mission of enabling access to effortless credit based on true risk. If you are energized by the impact you can make at Upstart, we’d love to hear from you!
The Team
Upstart’s Site Reliability Engineering (SRE) team owns the reliability, resiliency, and observability of Upstart’s production systems. We build automation, tooling, and frameworks to ensure our infrastructure is healthy, scalable, and able to support a seamless experience for both engineers and customers. Our scope includes defining Upstart’s technology operations risk strategy, implementing disaster recovery planning, and setting company-wide reliability standards.
As a Principal Site Reliability Engineer at Upstart, you will serve as a thought leader and SRE evangelist - driving adoption of best practices, mentoring engineers across the organization, and influencing both technical and business decisions. Your impact will extend beyond SRE into cross-functional collaboration with Product Engineering, DevEx, Development Productivity (Quality), DevOps, Data Engineering, and Machine Learning teams to elevate operational excellence across the company.
How you’ll make an impact:
- Lead the definition, advocacy, and adoption of SRE principles across engineering teams
- Partner with leadership to shape long-term reliability, resiliency, and observability strategies
- Champion distributed tracing, real user monitoring (RUM), and key performance metrics such as Largest Contentful Paint (LCP) to improve system visibility and user experience
- Build and scale self-healing systems to minimize manual intervention and reduce downtime
- Drive enterprise-wide improvements to incident response processes, including those related to Machine Learning systems
- Collaborate closely with Development Productivity and Quality teams to improve engineering velocity without sacrificing reliability
- Influence technical and operational roadmaps through data-driven insights and hands-on technical contributions
- Own and deliver cross-functional initiatives from concept through execution, applying program management skills to align stakeholders and achieve results
What we’re looking for:
- Minimum requirements:
- 10+ years combined experience across Software Engineering and Site Reliability Engineering, with a balanced background in both disciplines
- Proven track record as an SRE thought leader and evangelist, driving adoption of reliability best practices across organizations
- Strong communication and mentoring skills to influence engineers across disciplines
- Proficiency in Python, Go, and JavaScript/TypeScript
- Proficiency with Infrastructure as Code (Terraform, CDK, CloudFormation, etc.)
- Experience building internal tooling from scratch in agile development environments
- Expertise with observability, distributed tracing, RUM, LCP, and performance monitoring tools (e.g., Datadog, Prometheus)
- Experience with on-call and incident management, including large-scale or ML-related incidents
- Strong background in automation and building self-healing systems
- Hands-on experience with LLM/GenAI to improve SRE efficiency and processes
- Program management skills, including the ability to propose innovative solutions, influence leadership, improve processes, and drive cross-functional projects to completion
- Preferred qualifications:
- Experience with service mesh
- Full stack development skills
- Experience building or extending observability platforms
- Background in Development Productivity or Quality Platforms
- Experience in high-scale SaaS, microservice-oriented cloud environments
Position Location - This role is available in the following locations: Remote, San Mateo, Columbus, Austin
Time Zone Requirements - This
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s