Sr. Site Reliability Engineer
ChubbAbout the role
We are seeking an experienced and proactive Site Reliability Engineer (SRE) with expertise in DevOps to join our robust technology team in the insurance sector. As an SRE, you will play a critical role in ensuring the availability, performance, resiliency, and scalability of Chubb's platforms. You will collaborate with software engineering teams to design, build, and maintain systems, driving continuous improvement in system infrastructure, automation, monitoring, and incident response. We are looking for someone who is highly skilled in designing scalable and resilient business systems, with a strong background in cloud services and configuration management tools.
Responsibilities:
- Design and implement various monitoring strategies to analyze system performance, identify and resolve bottlenecks, and optimize system resource utilization.
- Develop automated scripts and tooling to improve performance, availability, and resiliency of systems.
- Collaborate with development teams to ensure system consistency and enhance deployment methods.
- Create and maintain operational tools for efficient deployment, monitoring, and analysis of systems.
- Drive the design and deployment of scalable and reliable systems, collaborating with architects and engineers to meet SLAs and mitigate risks.
- Perform Root Cause Analysis for production errors and design strategies to mitigate future occurrences.
- Enforce best practices for SRE, DevOps, infrastructure as code, and automated testing.
- Proactively identify and resolve technical and operational issues related to systems.
- Document processes, configurations, and procedures to ensure accurate and up-to-date documentation.
Location:
New Jersey (Jersey City, Whitehouse Station). Remote work 2 days a week will also be considered.
Qualifications:
- Bachelor’s degree in Computer Science, Engineering, or a related field.
- 7+ years of experience as a Site Reliability Engineer or similar role with DevOps responsibilities.
- Strong background and practical experience in designing scalable and resilient business systems.
- Familiarity with application and system reliability design patterns.
- Proficiency in cloud services (preferably Azure) and configuration management tools (GitHub, Jenkins, Ansible, Terraform, Docker, Kubernetes, Jira, Service Now).
- Experience in performance and availability monitoring of applications and servers, infrastructure provisioning, and security certificate management.
- Solid understanding of networking protocols, DNS, VPN, Load Balancing and firewall management.
- Proficiency in programming languages such as Python, PowerShell, .NET, Angular, including experience consuming Python+Flask REST API.
- Proficiency in CI/CD DevSecOps solutions using Jenkins pipelines using Python, Git, Shell, YAML, Ansible, Kubernetes and Docker.
- Knowledge and experience with monitoring, logging, and alerting tools, such as AppDynamics, Azure Monitor, ELK Stack, and App Insights.
- Strong troubleshooting and problem-solving skills, with experience managing complex production infrastructures, including Linux/Unix and/or Windows administration concepts and practices.
- Demonstrate proficiency in PowerShell scripting, Winrm commands, and data manipulation techniques, as well as the ability to develop insightful dashboards and reports.
- Leverage troubleshooting skills to address issues across various components, including server-side, client-side (browser), Windows Server, IIS, SQL Server, network problems, and performance bottlenecks.
- Utilize tools such as SOAP, REST, SQL Client, and Wire Shark to support system integration and analysis.
- Analyze error reports, incidents, and identify recurring patterns using Azure App Insights, AppDynamics, Full Story, and ELK Stack to drive continuous improvement and enhance system reliability.
- Excellent communication and teamwork skills, with the ability to collaborate effectively with cross-functional teams.
- Strong documentation skills, including the ability to develop runbooks, disaster recovery procedures, and KT documents.
- Experience in the insurance industry will be an added advantage.
Join our dynamic team if you are a persistent, ownership-driven professional with a passion for operating scalable and resilient systems. This position offers an opportunity to work in a fast-paced environment and make a substantial impact within the company.
In Jersey City, NJ the pay range for the role is $101,000 to $172,000. The specific offer will depend on an applicant’s skills and other factors. This role may also be eligible to participate in a discretionary annual incentive program. Chubb offers a comprehensive benefits package, more details on which can be found at
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s