Senior AI Linguist
LinkedInAbout the role
Company Description
LinkedIn is the worlds largest professional network, built to create economic opportunity for every member of the global workforce. Our products help people make powerful connections, discover exciting opportunities, build necessary skills, and gain valuable insights every day. Were also committed to providing transformational opportunities for our own employees by investing in their growth. We aspire to create a culture thats built on trust, care, inclusion, and fun where everyone can succeed.
Job Description
This role will be based in Mountain View, CA.
At LinkedIn, our approach to flexible work is centered on trust and optimized for culture, connection, clarity, and the evolving needs of our business. The work location of this role is hybrid, meaning it will be performed both from home and from a LinkedIn office on select days, as determined by the business needs of the team.
HALO (Human Judgment, Annotation, Localization, and Operations) is a horizontal team within Core AI that partners across the company to enable high-quality human judgment for AI development. We partner closely with cross-functional stakeholders and internal teams to define quality goals, design evaluation and data pipelines, and scale repeatable measurement systems. Our work spans multiple initiatives at once, supported by shared standards, platforms, and best practices that help teams move faster without compromising quality.
Ā
Key Responsibilities
Partner cross-functionally with Engineering, Product, Data Science, domain SMEs, Trust/Legal, TPM, and vendor operations to align on quality goals, tradeoffs, and delivery plans
Define and maintain evaluation frameworks, rubrics, rating scales, defect taxonomies, and agent-specific guideline addenda for ambiguous and multi-step agent behaviors across evolving product use cases and i18n markets
Design and run scalable evaluation systems, including metrics, scorecards, regression sets, monitoring plans, scenario suites, evaluation strategies, and success criteria for agent quality assessment
Build and operate high-quality annotation and evaluation pipelines, including task design, in-context evaluation surveys, evaluation strategies, QA gates, adjudication, and workflow maintenance across internal and vendor platforms
Generate, annotate, and validate high-quality human, synthetic, and adversarial data, including reasoning-rich judgments where needed, to improve evaluation coverage, identify blind spots, and support LLM-as-a-judge and reward model development
ĀLead calibration, drift detection, and disagreement analysis between human and model judgments, and translate findings into edge cases, retraining opportunities, and quality improvements
Manage vendor and internal workforce quality, including onboarding, task coordination, guideline training, audits, escalations, and cost-quality tradeoff decisions to maintain high annotation standards
Run error analyses and workflow experiments, document results, and drive iterative improvements based on evidence
Define requirements for human judgment and evaluation tooling, and partner on build, test, deployment, and adoption
Establish reusable best practices, enable partner teams on evaluation methodology, criteria interpretation, and judge score usage, and mentor junior team members
Demonstrate learning agility and adaptability in a fast-evolving field by quickly absorbing new tools, methodologies, and domain knowledge, staying current with changes, and continuously applying new learnings to improve evaluation quality and workflows
Apply native-speaker linguistic and cultural expertise in French (France), German (Germany), Spanish (Spain), Portuguese (Brazil), or other i18n market(s) to define and uphold market-appropriate quality standards for AI products with i18n
Qualifications
Required Qualifications
BA/BS in Computational Linguistics, Linguistics, Language Technologies, or related field
2+ years of industry experience owning end-to-end human judgment and quality workflows for AI development
Proven ownership of medium-to-large evaluation or annotation initiatives (method + delivery)
Demonstrated cross-functional collaboration with Engineering/Product/Data partners, including managing tradeoffs, dependencies, and execution risks
Experience building datasets and evaluation workflows for LLMs and agentic systems, including prompt-based labeling and evaluation, hybrid human-in-the-loop review, automated validation or consistency checks, and iterative dataset development to improve model and agent performance
Experience using AI-assisted data workflows to improve annotation, evaluation, o
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights ā in under 60 seconds.
Apply Now āGenerate Application KitFree account required ā sign up in 30s