Jobs and Careers
WE

AI Ops Principal Engineer

Wells Fargo
United Statesfull_timeVerifiedPosted 10 Jun 2026
💰 $305,000/yr

About the role

About this role:

Wells Fargo is seeking a Principal Engineer – AIOps to join Platform Strategy & Transformation as part of Commercial & Corporate and Investment Management Technology (CCIBT) group. Learn more about the career areas and business divisions at wellsfargojobs.com.

This role sits at the core of CCIBT’s Zero Touch Production (ZTP) transformation agenda, driving the strategy, architecture, and execution of next‑generation AIOps capabilities across the enterprise. You will define and deliver intelligent, autonomous operations by leveraging AI/ML, observability, automation, and event-driven architectures to minimize manual intervention, improve resilience, and enable self-healing systems.

You will partner closely with senior engineering, platform, SRE, and business leaders to accelerate AIOps adoption, embed intelligence into production ecosystems, and deliver measurable improvements in availability, efficiency, and operational risk reduction.

This is a hands-on senior developer role requiring strong development skills and ability to work with advanced automations using technologies like Robotic Process Automation (RPA), Artificial Intelligence, Low-code technologies like UiPath, Microsoft Power Platforms, Google ADK, LangChain, LangGraph, Alteryx etc.

In this role, you will:

  • Lead the strategy, design, and execution of AIOps platforms and capabilities to enable Zero Touch Production across CCIBT
  • Define and drive enterprise-wide AIOps roadmap, including observability, event correlation, anomaly detection, predictive insights, and automated remediation
  • Architect and implement self-healing systems leveraging AI/ML, event-driven automation, and closed-loop workflows
  • Drive adoption of intelligent incident management, root cause analysis (RCA), noise reduction, and auto-resolution techniques
  • Establish target-state architecture and engineering standards for AIOps platforms, tooling, and integrations
  • Influence enterprise technology strategy by evaluating emerging AIOps trends, tools, and frameworks
  • Partner with SRE, infrastructure, cloud, and application teams to embed AIOps into SDLC, CI/CD, and production operations
  • Lead large-scale engineering initiatives with cross-functional and enterprise impact
  • Provide thought leadership on resilience engineering, reliability, automation, and production excellence
  • Mentor and guide senior engineers and teams on AIOps best practices, architecture, and implementation
  • Collaborate with risk, compliance, and governance teams to ensure secure, compliant, and auditable automation

Required Qualifications:

  • 7+ years of Engineering experience, or equivalent demonstrated through one or a combination of the following: work experience, training, military experience, education
  • 5+ years of experience in AIOps, SRE, production engineering, or large-scale distributed systems operations
  • 4+ years of experience with Python, programming, or scripting languages
  • 2+ years of experience working with Generative AI, large language models (LLM), or foundation models

Desired Qualifications:

  • 2+ Agentic AI and Agent building experience
  • Experience with AI-powered development or GitHub Copilot
  • Proven experience designing and implementing observability, monitoring, and automation platforms at scale
  • Deep expertise in AIOps platforms and tools (e.g., Prometheus, AppDynamics, Splunk, ITRS Geneos, BigPanda, OpenTelemetry ecosystems)
  • Strong experience with AI/ML for IT operations, including anomaly detection, event correlation, forecasting, and intelligent alerting
  • Hands-on experience with automation frameworks (e.g., Ansible, Terraform, or similar) and event-driven architectures
  • Strong understanding of SRE principles, SLIs/SLOs, error budgets, and reliability engineering practices
  • Experience building self-healing systems and closed-loop remediation workflows
  • Proficiency in cloud platforms and cloud-native architectures (Kubernetes, microservices)
  • Knowledge of data pipelines, streaming platforms (Kafka), and telemetry ingestion/processing
  • Familiarity with GenAI/LLM-assisted operations, including incident summarization, knowledge mining, and automated runbook generation
  • Ability to operate across complex organizational structures with strong stakeholder management and communication skills
  • Proven ability t

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Wells Fargo

View company profile →