Senior Associate - Senior Site Reliability Engineer
New York Life Insurance CoAbout the role
Location Designation: Hybrid - 3 days per week
When you join New York Life, you’re joining a company that values career development, collaboration, innovation, and inclusiveness. We want employees to feel proud about being part of a company that is committed to doing the right thing. You’ll have the opportunity to grow your career while developing personally and professionally through various resources and programs. New York Life is a relationship-based company and appreciates how both virtual and in-person interactions support our culture.
Join Strategic Capabilities and collaborate with a winning team developing and executing game-changing business strategies, informed by cutting-edge data and competitive insights. Fuel business growth through smart acquisitions, innovative partnerships, and impactful technology integration. Leverage AI and strategic data management to bring ideas to life, all while ensuring effective governance and project management.
We are looking for a skilled Platform Engineer alias Site Reliability Engineer to join our Artificial Intelligence and Data (AI&D) group, focusing on managing AWS Infrastructure for our integration platform team. The role involves managing and optimizing AWS Infrastructure and native services using Terraform. The ideal candidate will have a strong background in AWS services, with knowledge on various deployment options and using MuleSoft and microservices architecture hosted on the AWS ecosystem, with a passion for automation using AI/ML frameworks, scalability, and reliability.
This role requires ability to consistently communicate across multiple channels (Teams, emails, meetings, conversations), proactively identifying blockers and escalating as needed, organizational skills to support the team in maximizing its performance and ensuring successful delivery of projects.
Responsibilities
Platform Implementation & Configuration
-
Design, implement, and manage infrastructure on AWS for MuleSoft, microservices on EKS and various other AWS services.
-
Configure monitoring and alerting tools to ensure real-time detection of issues, bottlenecks, and system anomalies.
Maintenance & Upgrades
-
Apply patches, upgrades, and configuration changes to ensure all services remain up-to-date and secure.
-
Maintain AWS infrastructure including EC2, S3, RDS, EKS, Load balancers, and VPCs to support the hosting environment.
-
Conduct regular system backups, disaster recovery drills, and security audits.
Operational Excellence
-
Proactively identify opportunities to automate processes related to deployments, scaling, and incident management using AI/ML frameworks
-
Troubleshoot production issues, debug failures, and resolve platform bottlenecks. Identify opportunities to use AI/ML frameworks and implement them.
Security & Access Management
-
Ensure compliance with security best practices, managing platform access using AWS IAM and ensuring correct role-based permissions.
-
Secure communication between MuleSoft, microservices, and other AWS services through proper configuration of firewalls, security groups, and networking.
Monitoring & Performance Tuning
-
Set up detailed monitoring using tools like CloudWatch, Prometheus, Grafana, ELK, and other relevant tools to track platform performance and health.
-
Optimize performance across all services and ensure system uptime and availability.
Capacity Planning
Monitor system capacity and performance metrics, forecast future needs, and implement strategies for scaling and load balancing.
Collaboration
Work closely with development teams to integrate new features, provide feedback on design and architecture, and ensure smooth deployment processes.
Documentation
Maintain comprehensive documentation for infrastructure, processes, and incident management.
Qualifications
-
5+ years of experience in a Site Reliability Engineer or similar role, with a strong background in AWS services and infrastructure
-
Extensive experience with various AWS services and related tools.
-
Strong background in deploying and managing microservices architectures in AWS using Kubernetes (EKS) and Docker.
-
Experience with CI/CD tools and practices, including pipeline development, automated testing, and infrastructure as code (IaC).
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s