Jobs and Careers
BI

Sr. Director, Site Reliability Engineering

Biofourmis
United StatesRemotefull_timeVerifiedPosted 17 Mar 2023

About the role

Biofourmis is a rapidly growing, global digital health company filled with committed, passionate professionals who care about augmenting personalized care and empowering people with complex chronic conditions to live better and healthier lives. We are pioneering an entirely new category of medicine by developing clinically validated, software-based therapeutics to provide improved outcomes for patients, smarter engagement & tracking tools for clinicians, and cost-effective solutions for payers. We are collectively devoted to a single-minded idea: powering personally predictive care.

Our dynamic growth has been marked by doubled headcount in the last 12 months via both expansion & acquisition, yielding a global footprint with offices in Boston, Singapore, Bangalore, and Zurich. We are backed by prominent international venture capital investment & have cultivated relationships with worldwide healthcare stakeholders over the last 5 years. Our talented team features numerous PhD’s in Data Science and Biostatistics, over 80 patents, prolific scientific publications, world-class systems, developers & engineers, and leaders in the clinical operations space.

Sr Director – Site Reliability Engineering

The position of Senior Director – SRE  will be responsible for managing global SRE team for all products of Biofourmis globally. This role will play a critical role in establishing and executing efficient Site reliability strategy for Biofourmis ensuring highest SLA for our customers.  In this position the candidate will be responsible for coordinating the tasks and monitoring software products from deployment to day to day monitoring as well as owning and formulating our Disaster Recovery strategy. The right candidate for this position will be an experienced SRE/DevOps leader who has demonstrated success leading SRE teams on global scale with proven track record of maintaining highest SLAs.

Responsibilities

  • The ideal candidate will have a proven track record of building and leading multiple support teams with a focus on problem solving, managing large-scale cloud infrastructure,
  • Lead and manage an SRE cloud team, ensuring high-performance, scalability, and reliability of the cloud infrastructure in AWS.
  • Manage on-call rotations across continents, using a follow-the-sun model
  • Architect, Implement and maintain incident response and disaster recovery plans for the cloud infrastructure
  • Ensure uptime, security, and efficiency of our cloud infrastructure, and establish a culture of continuous improvement through analytics and metrics.
  • Strong People Management skills and experience with managing global teams.
  • Provide leadership and guidance on all technical aspects of the cloud infrastructure, including automation, monitoring, and performance optimization.
  • Develop and maintain positive relationships with vendors and other external partners, and negotiate service contracts and SLAs
  • Own end-to-end availability and performance of key services and execute Disaster recovery activities for all products.
  • Plan and execute scheduled maintenance including cluster upgrades , software updates and infrastructure patching etc.
  • Manage Operations service including 24*7 monitoring of Infrastructure and applications.
  • Lead by example, mentor the team and establish credibility through quality technical execution
  • Promote the use of observability tools to other engineering organization
  • Demonstrated ability to utilize modern monitoring tools (DataDog, Prometheus, ELK, Cloudwatch,Grafana etc.)
  • Partner with the Security Operations  team to address security vulnerabilities and implement security best practices.
  • Partner closely with Customer Engineering team for Troubleshooting infrastructure and application Issues in conjunction with DevOps team members.
  • Provide technology leadership to the team and foster engineering excellence

Requirements

  • Bachelor or Masters in Computer Science or Information Technology
  • 12+ years of  professional experience as SRE Manager (5+ years managing teams).
  • Experience in Site Reliability management and DR strategy in AWS at a global scale
  • Able to operate effectively in a complex organizational structure.
  • Excellent communication skills and influencing skills
  • Strong knowledge in DevOps tools (OpenSource or otherwise) and practices and Agile software development methodology.
  • Independent and self-motivated contributor and passionate about software development.

Skills

  • Experience in monitoring experience in digital therapeutics industry (e.g. Medical Devices, Pharma, high end engineering).
  • Familiar with protections under HIPAA and security practice

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Biofourmis

View company profile →