Technical Lead - HPC
Era4About the role
Era4 develops, owns and operates AI infrastructure across the UK, powered by renewable energy. Converting legacy industrial and energy sites into modern data-centre facilities, Era4 is combining brownfield regeneration opportunities with cleaner, efficient, scalable compute capacity for healthcare, research, finance, enterprise, and public-sector organisations.
Role Summary:
This is a hands-on Technical Lead who can support the build, operation and scale Era4’s sovereign AI/HPC infrastructure. We need someone who can operate at the intersection of HPC platform engineering, GPU infrastructure, Linux systems, Kubernetes/Slurm, automation, observability and production incident response.
You will act as a technical authority, helping shape how GPU clusters are deployed, monitored, automated, supported and improved. You will work closely with SRE, platform engineering, infrastructure, vendors and customer-facing teams to ensure our platform is reliable, observable, scalable and ready for production workloads.
Key Responsibilities:
HPC & GPU Platform Leadership:
- Lead infrastructure, supporting GPU, compute, storage and networking platforms.
- Technical authority across platform engineering, SRE, infrastructure and customer-facing teams.
- Mentor engineers and help define technical standards, best practices and operational excellence.
Platform Reliability & Incident Response:
- Lead technical investigations during major incidents and production outages.
- Improve observability, monitoring and alerting across the platform.
- Drive root-cause analysis and implement long-term reliability improvements.
Automation & Platform Engineering:
- Improve platform scalability, deployment processes and operational efficiency.
- Oversee runbook creation, automation and development of inhouse Agent capabilities to support Operations
- Contribute to the design, build, enhancement of future capabilities
Customer & Technical Engagement:
- Support customer onboarding, complex technical escalations and platform adoption.
- Work with vendors, partners and internal teams to resolve infrastructure issues.
- Translate complex technical challenges into clear, actionable communication.
AI/HPC Infrastructure Evolution:
- Contribute to Technical and Operational Roadmaps
- Contribute to the deployment and optimisation of GPU clusters, Kubernetes environments and next-generation AI infrastructure.
Experience:
You do not need to tick every technology box, but you must bring hands-on experience in production infrastructure and clear depth in HPC, GPU, AI infrastructure, research computing, Neocloud, cloud HPC or high-density compute environments.
- Linux engineering background
- Experience in production of leading teams supporting HPC, GPU infrastructure, AI infrastructure, research computing, cloud HPC, Neocloud, platform engineering or SRE.
- Experience operating or supporting production infrastructure across compute, networking, storage and observability.
- Hands-on experience with at least one workload or orchestration layer such as Slurm, Kubernetes, Run:ai, LSF, PBS or equivalent.
- Experience with Open Source monitoring and troubleshooting using tools such as Prometheus, Grafana, OpenTelemetry, Loki or equivalent.
- Proven involvement in major incidents, on-call, escalation, root-cause analysis or production troubleshooting.
- Ability to mentor engineers, influence technical direction and lead by technical credibility.
- Comfortable working with internal teams, customers, suppliers and vendor engineering teams.
Why Join Era4:
You’ll be joining a mission-driven start-up building critical national infrastructure, where operational excellence directly enables growth. This role offers high visibility with leadership, real autonomy, and the chance to shape how a next-generation company operates at scale.
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s