Jobs and Careers
MI

Team Lead, Site Reliability Engineering - IntelliScript (Remote)

Milliman
Brookfield, United StatesRemotefull_timeVerifiedPosted 26 Dec 2024
💰 $248,000/yr($122,000/yr$248,000/yr)

About the role

What We Do

Milliman IntelliScript is a group of a few hundred experts in fields ranging from actuarial science to information technology to clinical practice. Together, we develop and deploy category-defining, data-driven, software-as-a-service (SaaS) products for a broad spectrum of insurance, health IT and life sciences clients. We are a business unit within Milliman, Inc., a respected consultancy with offices around the world.

Candidates who have their pick of jobs are drawn to IntelliScript’s entrepreneurial and collaborative culture of innovation, excellence, exceptional customer service, balance, and transparency. Every single person has a voice in our company, and we challenge each other to push the outer limits of our full, diverse potential. And, we’ve shown sustained growth that ensures you’ll have room to grow your skillset, responsibilities, and career.

Our team is smart, down-to-earth, and ready to listen to your best ideas. We reward excellence and offer competitive compensation and benefits. Visit our LinkedIn page for a closer look at our company, and learn more about our cultural values here.

Milliman invests in skills training and career development and gives all employees access to a variety of learning and mentoring opportunities. Our growing number of Milliman Employee Resource Groups (ERGs) are employee-led communities that influence policy decisions, develop future leaders, and amplify the voices of their constituents. We encourage our employees to give back to their varied professions, including leadership in professional organizations. Please visit our website to learn more about Milliman’s commitments to our people, diversity and inclusion, social impact, and sustainability.

What this position entails

IntelliScript’s Information Technology has been a key part of our success and is critical to our future. In this position, you will work closely with cross-functional teams to drive operational excellence, automate processes, and continuously improve system reliability. We are looking for someone passionate about identifying and mitigating risks, ensuring timely and effective solutions. The Team Lead, SRE will act as an advocate for reducing complexity and empowering others across the technology organization to drive excellence and innovation. We are looking for a blend of hands-on expertise and leadership experience, acting as a player/coach to help develop the SRE practice at IntelliScript.

What you will be doing

  • Proficiency in Observability, guiding teams on best practices for implementing monitoring, logging, and tracing to improve application reliability and performance based on the application’s Criticality Tier
  • Partner closely with Software Engineering, Cloud Engineering, and Architecture teams to improve standards and Architectural blueprints, ensuring patterns are resilient, performant, and cost effective
  • Collaborate closely with our business stakeholders to understand when our customers experience service degradation, define SLOs, measure adherence to SLAs, and establish Error budgets
  • Help manage risk in our production environment by participating in release reviews, and blocking risky changes that could jeopardize the Error Budget
  • Work closely with Service Desk and Operations teams to proactively identify and prevent disruptions, aiding with incidents, and recommending preventative measures
  • Gain a strong understanding of our environment and establish anomaly detection and resolution automation
  • Establish the SRE practice, setting the North Star vision, and building a team to get us there

What we need

  • 10+ years of relevant experience
  • Prior expertise in Infrastructure/Cloud (AWS), Application Development, Data/Storage, and/or DevOps
  • Strong expertise in Observability (Datadog, CloudWatch, Redgate, App Dynamics, NewRelic, Dynatrace, and/or Prometheus/Grafana)
  • Deep knowledge of on-premises and cloud infrastructure (AWS)
  • Experience in incident response, disaster/disruption recovery, and designing systems for high availability and resilience
  • Experience writing Infrastructure as Code (IaC) with tools like Terraform, Ansible, CDK or Cloudformation
  • Strong leadership qualities and technical skills, with the ability act as a player/coach
  • Strong communication skills to collaborate cross-functionally and make recommendations that align to the business and technology teams
  • Clearly exhibits leadership qualities that foster an inclusive and positiv

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Milliman

View company profile →