Principal Staff Software Engineer, Systems Infrastructure
LinkedInAbout the role
Company Description
LinkedIn is the world’s largest professional network, built to create economic opportunity for every member of the global workforce. Our products help people make powerful connections, discover exciting opportunities, build necessary skills, and gain valuable insights every day. We’re also committed to providing transformational opportunities for our own employees by investing in their growth. We aspire to create a culture that’s built on trust, care, inclusion, and fun – where everyone can succeed.
Join us to transform the way the world works.
Job Description
At LinkedIn, our approach to flexible work is centered on trust and optimized for culture, connection, clarity, and the evolving needs of our business. This role may be remote or hybrid. At LinkedIn, hybrid roles are performed both from home and from a LinkedIn office on select days, as determined by the business needs of the team. Remote roles are performed from the designated home work location upon time of hire, and any changes to this home work location requires a review of remote status and approval.
LinkedIn’s Reliability Infrastructure team is responsible for defining and driving the reliability strategy, standards, and practices that keep LinkedIn’s most critical systems stable, resilient, and available at massive scale.
As a Principal Staff Software Engineer, Reliability Infrastructure, you will serve as a senior technical authority for reliability across LinkedIn Engineering. You will help define how critical services are designed, built, operated, and measured, partnering broadly across infrastructure and product engineering teams to improve resiliency, reduce incidents, and raise the reliability bar across the company.
A key focus of this role is driving the adoption and evolution of LinkedIn’s service criticality framework, including reliability expectations for the most business-critical systems. You will help classify services based on criticality and blast radius, define appropriate reliability standards, and influence system architecture to ensure the right levels of availability, redundancy, observability, and failure handling are in place.
As AI-assisted software development, agent-based automation, and autonomous operational systems become more prevalent, this role will also help define how LinkedIn safely builds and operates reliable AI-enabled systems. You will shape standards for evaluating, deploying, monitoring, and governing AI-generated code and agentic workflows, ensuring that automation introduced into critical environments is observable, explainable, auditable, and designed with appropriate safeguards, rollback mechanisms, and human oversight.
This is not a traditional SRE role focused on operating a single service or team. It is a company-wide technical leadership role for someone with deep distributed systems expertise, strong reliability judgment, and the ability to influence architecture and engineering practices across large organizations.
Responsibilities
Define and drive company-wide reliability strategy, standards, and best practices across LinkedIn Engineering
Lead adoption and evolution of service criticality models that set reliability expectations based on business impact and blast radius
Serve as a technical authority for architecture decisions related to reliability, resiliency, availability, and failure handling
Partner with infrastructure and product engineering teams to improve system design, reduce incident risk, and strengthen operational readiness
Identify high-risk systems and drive cross-organizational initiatives to improve reliability of critical services
Establish and evolve reliability standards including SLOs, SLIs, uptime expectations, redundancy, monitoring, alerting, and failover patterns
Influence engineering culture by promoting reliability-focused design, incident review rigor, and postmortem-driven improvements
Provide architectural guidance and mentorship to senior engineers and technical leaders across teams
Balance technical strategy, hands-on engineering judgment, and cross-functional influence to drive measurable improvements in site stability
Help shape how LinkedIn builds and operates resilient systems as the platform continues to scale
Drive the strategy for applying LLMs to alert triage, root cause analysis, and incident summarization at scale, ensuring systems are explainable, auditable, and safe to operate autonomously in Ring0/Ring1 environments.
Qualifications
Basic Qualifications
BA/BS degree in Computer Science or related technical field, or equivalent practical experience
10+ years of experience in software engineering, infrastructure engin
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s