Jobs and Careers
DI

Principal, IT Software Engineer 2 - AIOps Lead

DIRECTV
El Segundo, United StatesRemotefull_timeVerifiedPosted 22 May 2025
💰 $255,530/yr($140,790/yr$255,530/yr)

About the role

DIRECTV is seeking an AIOps Lead (Principal, IT Software Engineer 2) who will play a crucial role in driving the adoption and execution of Artificial Intelligence for IT Operations (AIOps) practices across the organization. This individual will be responsible for leading observability standards, AIOps initiatives, automation-first strategies, leveraging AI and machine learning technologies to optimize IT operations, detecting anomalies, improving system performance, and automating incident and problem management processes.

The ideal candidate will have a strong background in IT operations, SRE, a deep understanding of observability platforms and AIOps and tools, DevOps, software development and the ability to lead cross-functional teams to drive innovation in the realm of IT operations automation and monitoring.

Here’s what you’ll do:

Team Leadership and Guidance:

  • Lead projects from a team of 3-4 NPW engineers dedicated to stability and observability improvements and operation efficiency.
  • Technical lead for a team to design and develop end-to-end solutions, managing dependencies and cross-team impacts.
  • Provide hands-on guidance and support to team members (50% hands-on, 50% managerial).
  • Lead a team of AIOps engineers and specialists, ensuring their development, coaching, and alignment with organizational goals.
  • Develop and report on team performance KPIs.
  • Foster a culture of continuous learning, DevOPS excellence through regular technical sessions and internal workshops.
  • Active participant in the development community (Business Unit) to promote best practices through educating their peers.
  • Manage risk and request help from leadership, when necessary, to meet commitments or change directions.

Observability, AIOPS Strategy and Execution:

  • Define and implement an Observability, AIOPS strategy aligned with business objectives and an autonomous IT operations vision.
  • Responsible for planning short term (sprint-to-sprint) and long-term (multiple PI) initiatives and organizing work and designs to meet the long-term target.
  • Implement and optimize AI and machine learning algorithms to detect performance anomalies, predict outages, automate incident response, and improve overall operational efficiency.
  • Implement automated workflows for proactive issue resolution, reducing manual intervention and improving operational agility.
  • Seek opportunities to improve processes and take an automation-first approach.
  • Lead the evaluation, selection, and deployment of AIOps platforms and tools.
  • Design and implement cost-efficient observability and AIOps solutions across cloud and on-premise environments using a mix of commercial, open source, and CNCF solutions.
  • Leverage data analytics and monitoring systems to generate actionable insights that improve system health, application performance, and availability.
  • Develop internal resources and training materials to ease the adoption and implementation of AIOPS tools and practices.

Cross-functional Collaboration:

  • Work closely with IT operations, DevOps, SRE and application development teams to identify pain points and automate processes with AIOps tools and techniques.
  • Present findings, improvements, and key metrics to senior management and stakeholders.

Automation and Process Improvement:

  • Leverage scripting, AI/ML, and automation skills for automation first approach.
  • Embed Observability and AIOps capabilities into reusable platform services by utilizing DevOps, CI/CD, and IaC tools and practices like Terraform, Jenkins, GitHub, ArgoCD, Harness and Ansible.

Technical Implementation and Management:

  • Establish and enforce observability standards, policies, and best practices across the enterprise.
  • Ensure compliance with regulatory and security requirements.
  • Plan and migrate legacy tools and functions to new AIOPS approach.
  • Develop and maintain AIOPS dashboards, extensions, applications, and workflow automation.
  • Integrate AIOPS with tools like Jira, ServiceNow, MS Teams, Slack, xMatters, Confluence/wiki/KB and MoogSoft/BigPanda.
  • Set up and manage observability stacks for cloud monitoring (AWS, Azure), VMs, Kubernetes, and various databases.
  • Optimize naming conventions, management zones, alerting profiles, and tagging to align with business processes.

Performance Monitoring and Reporting:

  • Analyze and report on observability metrics, KPIs, Service Level Indicators (SLI), and Service Level Objectives (SLOs).
  • Develop and recommend baseline monitoring thresholds, SLO, and error budgets to drive continuous improvement in MTR and Availability.

What y

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

DIRECTV

View company profile →