Senior Site Reliability Engineer II
RELXAbout the role
Are you a collaborative Azure Sr SRE looking to work for a mission driven global organization?
Do you possess advanced Azure SRE skills and looking to put those skills to use to help drive innovation?
LexisNexis® Risk Solutions provides customers with innovative technologies, information-based analytics, decisioning tools and data management services that help them solve problems, make better decisions, stay compliant, reduce risk and improve operations. Headquartered in metro-Atlanta, Georgia it operates within the Risk market segment of RELX, a global provider of information-based analytics and decision tools for professional and business customers.
About the role; This Azure SRE will administer (Azure) AKS clusters running critical always-on middleware handling thousands of TPS. They will be expected to conduct operations in a manner consistent with a five-9’s availability target.
About the team, this team is entrusted with applying software engineering practices to IT operations tasks to maintain a scalable and reliable production environment. This team also automates recovery to protect critical service levels.
Responsibilities
- Designing, writing, and maintaining Infrastructure as Code (IaC) using Terraform and Helm to provision and manage cloud environments (AWS, Azure, or GCP).
- Creating and maintaining Helm charts for Kubernetes deployments, ensuring consistency across environments.
- Developing modular, reusable Terraform templates for VPCs, subnets, clusters, and observability tooling.
- Conducting code reviews for Terraform and Helm changes to ensure compliance, reusability, and security.
- Individuals are responsible for challenging reliability and toil reduction projects. At this level, SREs have hands-on experience across most SRE practices.
- Contributing to process improvements through experience and knowledge.
- Deploying AKS cluster and cutovers, base image updates, testing IaC changes, and other work focused on daily operations
Requirements
- Current an extensive experience as an Azure SRE.
- Possess a deep understanding of IaC configuration design, software defined networking and infrastructure, and observability platforms such as ELK, Grafana Loki, and/or OpenTelemetry.
- Possess an understanding of how to observe distributed systems and their dependencies, and how to automate recovery to protect service levels.
- Proficiency in scripting and automation (e.g., Python, Bash, Ansible). Knowledge of running software developed in languages such as Java, C++, .NET, and/or Node.js is expected
- This position requires experience with Kubernetes, Terraform, Helm, GitHub (Actions), ArgoCD, and IaC automation as well as solid Linux and cloud networking skills
We are committed to providing a fair and accessible hiring process. If you have a disability or other need that requires accommodation or adjustment, please let us know by completing our Applicant Request Support Form
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s