Engineering Leader, Foundational Infra and Reliability
BenchlingAbout the role
Biotechnology is rewriting life as we know it, from the medicines we take, to the crops we grow, the materials we wear, and the household goods that we rely on every day. But moving at the new speed of science requires better technology.
Benchling’s mission is to unlock the power of biotechnology. The world’s most innovative biotech companies use Benchling’s R&D Cloud to power the development of breakthrough products and accelerate time to milestone and market.
Come help us bring modern software to modern science.
ROLE OVERVIEW
You will be the engineering manager for the Foundational Infrastructure and Reliability Engineering team at Benchling.
This team is responsible for driving reliability at Benchling. This means defining, evangelizing, and directly executing processes and programs to improve the reliability, resilience, and recoverability of Benchling’s SaaS offerings, and offering visibility into the success of such initiatives.
The team is heavily reliant on observability tooling to accomplish its goals, and owns the observability tools we use for application metrics, logging, and tracing.
The team provides oversight on all cloud spend at Benchling and puts safeguards in place to proactively provide cost governance.
Finally, the team also provides other Infra teams with tools to implement infrastructure changes in a reliable and safe manner. This involves, for example, owning and managing Terraform Cloud workspaces for all teams at Benchling.
RESPONSIBILITIES
As the engineering leader, you will:
-
- Lead and manage a team of Software Engineers, fostering a culture of collaboration, innovation, and continuous improvement.
- Provide mentorship and guidance to team members, supporting their professional growth and development.
- Oversee the design, implementation, and maintenance of highly reliable and scalable infrastructure components, including servers, networks, databases, and storage.
- Own critical business processes including incident, problem, site availability, and performance management, and associated tools and business outcomes
- Be the cross-functional leader of Technical Operations process, and drive operational excellence and accountability. Drive continuous improvement initiatives, analyzing incidents and problems to identify root causes and implement preventive measures.
- Own observability tools and improve ROI. Develop strategies for monitoring the performance and health of critical systems and applications, ensuring proactive detection and resolution of issues.
- Contribute to the technical strategy and architecture of the infrastructure, aligning it with the company's business objectives and growth plans.
- Collaborate with external vendors and service providers to evaluate and integrate third-party tools and services that enhance reliability.
- Build and maintain comprehensive documentation of systems, processes, and best practices to facilitate knowledge sharing and training.
- Establish and operationalize SRE model for the company
- Implement cost governance for cloud, observability and other spend areas. Collaborate with engineering teams to forecast resource requirements and implement scaling strategies to meet demand effectively.
- Champion a culture of decentralized DevOps of ‘you build it you run it’
- Enhancing Benchling’s platform resilience through innovative practices. Own Disaster Recovery.
- Drive the development of automation tools and processes to streamline deployment, configuration management, and infrastructure scaling.
QUALIFICATIONS
To succeed in this role, you ideally are:
-
- An engineering manager for 3+ years, or a manager of managers for 1+ years
- A hands on software engineer for 5+ years
- An IaaS/PaaS expert with strong technical skills in AWS, Python, GraphQL, Datadog, BuildKite/Jenkins, ArgoCD, and Terraform
- Proficient in driving SRE adoption in a medium or large-sized organization
- Experienced in driving DR strategy
- Exceptional in cross-functional program managemen
The following make you a preferred candidate:
-
-
- A strong understanding of cloud computing and microservices architectures with operations experience at scale
- SRE leadership experience with hands on SW Development skills
- A bachelor’s degree or equivalent in Computer Science, Computer Engineering, or a related field
-
SALARY RANGE
Benchling takes a market-based approach to pay, and pay may vary depending on your location. The candidate's starting pay will be determined based on job-related skills, experience, qualifications, interview performance, a
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s