Jobs and Careers
LU

Sr Lead Site Reliability Engineer

Lumen Technologies
Remote, US, United StatesRemotefull_timeVerifiedPosted 9 Jul 2026
💰 $193,940/yr($132,232/yr$193,940/yr)

About the role

Lumen is the trusted network for the AI‑powered world, connecting people, data, and applications through our expansive fiber network and connected ecosystem. We enable secure, high‑performance connectivity across cloud, edge, and AI workloads for enterprises, governments, and communities.

At Lumen, you’ll work on infrastructure customers rely on today and build for what’s next, where performance, security, and resilience matter.

This is a high accountability environment where bold ideas drive real innovation for our customers, partners, and industry. The work is challenging, expectations are clear, and trust is built into how we operate. If you’re ready to take ownership, deliver meaningful impact, and help shape the future of AI‑ready connectivity, join us today.

The Role

We are seeking a highly skilled and proactive Senior Lead Site Reliability Engineer (SRE) to join our team, focusing on production support and performance optimization across our portal ecosystem. This role is critical to ensuring the reliability, scalability, and efficiency of our systems, with a strong emphasis on AWS infrastructure, observability, automation and AI-assisted engineering practices.  

 

The Senior Lead SRE requires an AI-native mindset, understands the software development lifecycle (from coding to support) and applies modern AI tools to enhance productivity, quality, and operational excellence. This role will shape how Lumen combines the latest technologies, including AI-driven automation, to modernize software delivery and application lifecycle management.

 

This role will collaborate with key stakeholders across the engineering organization — including product owners, developers, and testers — to design, optimize, and automate business and technical processes, while effectively navigating multiple teams within a large and complex organization.

Location

This role is designated as a fully remote position within the United States.

The Main Responsibilities

Production Support & Incident Management

  • Help define and improve the processes and best practices for incident management within the customer facing applications.
  • Implement AI systems and automations to assist during ongoing outages and triage potential ones. You will work with the development teams to ensure that they have all the data normally needed during an outage at their fingertips including preliminary analysis by AI.
  • Design and implement improved SRE processes to handle incidents. From proactive analysis, faster response, automatic remediation, AI guided analysis, partially automatic root cause analysis.

Performance Optimization

  • Monitor system performance and proactively identify bottlenecks or degradation using AI-driven observability and anomaly detection tools.
  • Implement tuning strategies across application layers, databases, and infrastructure. 
  • Drive initiatives to improve latency, throughput, and resource utilization.

Monitoring & Observability

  • Deploy improved alerting for Lumen Connect in depth, focusing on outside in but including early indicators for fulfilment and other areas. Combining traditional monitoring with AI-based anomaly detection and noise reduction.
  • Proactively monitor the errors and performance on Lumen Connect. Implement rules to detect deviations, implement improvements together with the teams.
  • Design and maintain dashboards, alerts, and metrics using tools like Datadog, AppInsights, CloudWatch, or similar.

Automation & Infrastructure as Code

  • Develop and maintain automation scripts and tools for deployment, scaling, and recovery, leveraging AI-assisted code generation and validation tools
  • Use Terraform, or similar IaC tools to manage AWS resources.

Reliability Engineering

  • Perform an in-depth analysis of the overall system and its dependencies, implementing techniques to increase the global availability, reduce the reliance on unstable dependencies and guide ecosystem improvements.
  • Champion SRE principles such as SLIs, SLOs, and error budgets.
  • Advocate for resilient architecture and fault-tolerant design patterns, incorporating AI-assisted design reviews and architecture evaluation.
  • Help define and improve better SRE processes and lead significant improvements in reliability for the Lumen Connect platform.

Collaboration & Communication

  • Work closely with software engineers, DevOps, and product teams t

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Lumen Technologies

View company profile →