Jobs and Careers
HP

Principal AI Systems Performance & Benchmarking Engineer

HP
Fort Collins, United Statesfull_timeVerifiedPosted 26 Feb 2026
💰 $230,850/yr($147,050/yr$230,850/yr)

About the role

Principal AI Systems Performance & Benchmarking Engineer

Description -

HP is seeking a Principal‑level AI Systems Performance & Benchmarking Engineer to own end‑to‑end validation, characterization, and proof of AI performance across HP workstation platforms. This role is responsible for defining and executing credible, repeatable, and defensible AI performance evaluations using real‑world AI workloads—including LLM inference, RAG pipelines, fine‑tuning, and GenAI applications—and translating results into actionable insights for engineering, product, and business stakeholders.

This role serves as HP’s technical authority on AI system performance, influencing platform design, performance claims, and industry benchmarking standards.

Core Responsibilities

AI System Performance Validation & Analysis

  • Lead hands‑on validation and characterization of AI performance across HP desktop and mobile workstation platforms.
  • Execute AI workload testing across local inference, training, and hybrid execution models using production‑representative configurations and datasets.
  • Evaluate system behavior across CPU, GPU, NPU/neural engines, memory, storage, and software stacks under sustained AI workloads.
  • Measure and analyze AI performance metrics including:
    • Time to First Token (TTFT)
    • Inter‑token latency
    • Throughput (tokens/sec)
    • Batch processing efficiency
    • KV cache behavior and optimization
    • Memory pressure, data movement, and unified memory utilization
    • Stability under sustained load
  • Identify performance limits, regressions, and bottlenecks without changing system architecture, providing evidence‑based feedback to engineering and product teams.
  • Conduct performance comparisons across on‑device, hybrid, and cloud/data‑center execution models.

AI Benchmarking Strategy & Execution

  • Own the selection, design, execution, and interpretation of AI benchmarks, including:
    • Industry benchmarks (MLPerf, LLM Perf, SPEC‑based tests)
    • Hugging Face benchmarks and GenAI‑specific evaluation frameworks
    • Custom benchmarks when standard methods fail to represent real workflows
  • Execute LLM inference benchmarks using modern serving frameworks such as vLLM, TensorRT‑LLM, text‑generation‑inference, and similar stacks.
  • Evaluate advanced AI performance characteristics including:
    • Multi‑GPU scaling efficiency
    • Tensor, pipeline, and distributed parallelism
    • GPU/CPU utilization and accelerator efficiency
    • Quantization impacts (GPTQ, AWQ, GGUF, NF4)
    • Distributed training and communication efficiency
  • Ensure all results are repeatable, documented, and defensible, using containerized, version‑controlled benchmark environments (Docker/Kubernetes).

Workflow‑First Performance Interpretation

  • Map benchmark results directly to real customer AI workflows, including enterprise LLM deployments, RAG systems, fine‑tuning pipelines, and multi‑modal AI applications.
  • Translate raw performance data into workflow‑level outcomes such as responsiveness, iteration speed, and scalability.
  • Validate and explain performance deltas across platforms, configurations, generations, and execution modes—focusing on why results differ, not just what changed.
  • Develop clear technical narratives that connect system behavior to user experience and business impact.
  • Publish performance findings through technical blogs, conference presentations, and contributions to industry benchmarking efforts.

Cross‑Functional Influence & Leadership

  • Partner closely with systems engineering, firmware, QA, AI software, and product teams to ensure performance findings are understood and actionable.
  • Support go‑to‑market, marketing, and sales enablement through:
    • Performance proof points
    • Customer‑facing benchmark claims
    • Demos, charts, and validation artifacts
  • Serve as a trusted authority on AI performance credibility, comparability, and methodology.
  • Influence HP’s multi‑year AI performance roadmap and internal benchmarking standards.
  • Represent HP in industry benchmarking consortia and standards bodies (e.g., MLPerf working groups).

Technical Leadership & Mentorship

  • Act as a senior technical reviewer for AI performance results, ensuring rigor, consistency, and accuracy.
  • Mentor engineers on AI benchmarking best practices, systems‑level performance analysis, and responsible interpretation of results.
  • Shape internal guidance on how AI performance should be measured, compared, and communicated.

Education & Experience

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

HP

View company profile →