Sr. Staff Software Engineer, Systems Infrastructure
LinkedInAbout the role
Company Description
LinkedIn is the world's largest professional network, built to create economic opportunity for every member of the global workforce. Our products help people make powerful connections, discover exciting opportunities, build necessary skills, and gain valuable insights every day. We're also committed to providing transformational opportunities for our own employees by investing in their growth. We aspire to create a culture that's built on trust, care, inclusion, and fun – where everyone can succeed.
Join us to transform the way the world works.
Job Description
This role will be based in Mountain View, CA, or New York City, CA.
At LinkedIn, our approach to flexible work is centered on trust and optimized for culture, connection, clarity, and the evolving needs of our business. The work location of this role is hybrid, meaning it will be performed both from home and from a LinkedIn office on select days, as determined by the business needs of the team.
LinkedIn’s HALO team is building the Evaluation Operating System (EOS), a foundational platform that defines how all AI agents and GenAI products at LinkedIn are measured, evaluated, and continuously improved in production. This is a brand-new, industry-defining problem space with no established playbook, focused on evaluating multi-step, non-deterministic, and personalized AI systems where traditional metrics and testing approaches fall short.
EOS acts as the central intelligence layer for AI quality, combining large-scale data pipelines, evaluator models (e.g., LLM-as-judge, reward models), and real-time production monitoring to understand how AI systems behave, where they fail, and how to improve them. The platform includes capabilities like synthetic data generation, adversarial testing, golden dataset management, and live “agent arena” experimentation frameworks (champion/challenger testing) to measure performance across multiple dimensions of quality.
As a Senior Staff Engineer, you will own the end-to-end technical vision, architecture, and execution of this platform. This includes designing the data infrastructure for capturing and labeling interactions, building systems to train and deploy evaluation models, and creating real-time monitoring and feedback loops that detect regressions, model drift, and quality degradation in production. You’ll work closely with AI product teams, ML engineers, and infrastructure partners to embed evaluation deeply into the development lifecycle, making it possible for teams across LinkedIn to ship high-quality AI systems with confidence.
This role sits at the intersection of distributed systems, data platforms, and machine learning, and is ideal for engineers who want to define how AI quality is measured at scale. The impact is company-wide: the systems you build will directly determine the quality ceiling, safety, and trustworthiness of every AI-powered experience at LinkedIn.
Responsibilities
Own the technical vision, architecture, and execution of the Evaluation Operating System (EOS), solving complex, open-ended challenges at the intersection of distributed systems, data infrastructure, and machine learning.
Design and build large-scale evaluation infrastructure that enables LinkedIn teams to measure, understand, and continuously improve the quality, reliability, safety, and performance of AI agents and GenAI products.
Architect scalable data pipelines and platforms for capturing, processing, labeling, and managing large volumes of AI interactions, evaluation data, golden datasets, and synthetic data.
Build and evolve evaluation systems powered by LLM-as-judge, reward models, and other automated evaluators to assess AI systems across multiple dimensions of quality and performance.
Develop experimentation and testing frameworks, including adversarial testing, champion/challenger experiments, and agent arena capabilities, to identify weaknesses and drive continuous improvement of AI systems.
Establish real-time observability, monitoring, and feedback loops that detect regressions, model drift, quality degradation, and unexpected behavior in production AI systems.
Partner closely with AI product teams, ML engineers, and infrastructure organizations to integrate evaluation deeply into the AI development lifecycle and establish consistent evaluation standards across LinkedIn.
Lead multiple high-impact, cross-functional initiatives, influencing technical strategy and architectural decisions across AI Platforms and the broader engineering organization.
Mentor and develop engineers, raise the technical bar, and help shape the engineering culture and practices of a growing AI platform organization.
Qualifications
Basic Qualifications:
Experience building or evaluati
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s