Jobs and Careers
NB

Site Reliability Engineer

NBCUniversal
Englewood Cliffs, United StatesRemotefull_timeVerifiedPosted 11 Jul 2024
💰 $145,000/yr($110,000/yr$145,000/yr)

About the role

Company Description

We create world-class content, which we distribute across our portfolio of film, television, and streaming, and bring to life through our theme parks and consumer experiences. We own and operate leading entertainment and news brands, including NBC, NBC News, MSNBC, CNBC, NBC Sports, Telemundo, NBC Local Stations, Bravo, USA Network, and Peacock, our premium ad-supported streaming service. We produce and distribute premier filmed entertainment and programming through Universal Filmed Entertainment Group and Universal Studio Group, and have world-renowned theme parks and attractions through Universal Destinations & Experiences. NBCUniversal is a subsidiary of Comcast Corporation.

Here you can be your authentic self. As a company uniquely positioned to educate, entertain and empower through our platforms, Comcast NBCUniversal stands for including everyone. Our Diversity, Equity and Inclusion initiatives, coupled with our Corporate Social Responsibility work, is informed by our employees, audiences, park guests and the communities in which we live. We strive to foster a diverse, equitable and inclusive culture where our employees feel supported, embraced and heard. Together, we’ll continue to create and deliver content that reflects the current and ever-changing face of the world.

Job Description

Are you passionate about digital media, entertainment, and software services? Do you like big challenges and working within a highly motivated team environment?  Keen with respect to Observability and Reliability principles? 

Candidates for the Site Reliability Engineer role in our AIOps group will be responsible for operational management & application support.

As a Site Reliability Engineer for AIOps, you will:  

  • Design, integrate, and provide full-stack lifecycle support for applications supporting our AIOps platform
  • Own CI/CD initiatives, while working with other technical leads to define and maintain compliance with organizational standards
  • Work closely with DevOps teams, customers, and infrastructure partners to identify & understand key system health/performance metrics, and develop monitoring approaches in support of service level objectives
  • Participate in incident cause-analysis and assist in remediation and design efforts to improve reliability/prevent future failure scenarios
  • Advance our AIOps offering through measurement & analysis of complex monitoring and alerting patterns, and provide clear guidance to collaborators on which tool/which pattern/which alert provides the desired Observability outcome 
  • Write code and scripts to automate everything possible
  • Assist in scoping, design, and project work/budget estimations

Qualifications

Technology Expertise and Ownership Requirements:

  • Expertise with implementing Site Reliability Engineering best practices and principles (Resiliency, Observability)
  • Expertise with Agile DevOps methodologies, process & associated software (ServiceNow Agile, Jira)
  • Expertise and troubleshooting skills for large-scale distributed computing systems and software
  • Experience with automation, CI/CD pipeline & software testing (Ansible, Terraform, Puppet, Chef, and Jenkins)
  • Experience with version control management (git, GitHub)
  • Experience with monitoring platforms (OP5/Nagios, Datadog, Splunk)
  • Experience with public cloud service offerings (AWS, Azure, Google)
  • Experience with OS management and features (Windows, Linux distributions)
  • Experience with script language development (Python, Node.js, Perl)
  • Familiarity with network technology concepts (TCP/IP, UDP, IPV4, IPV6, DNS, SSL, Firewalls, F5 LTM)
  • Familiarity with general cybersecurity best practices, close collaboration with Cyber Security team

Collaboration and Technical Communication Requirements

  • Impeccable written and verbal communication and presentation skills for both technical and non-technical audiences
  • Ability to collaborate enthusiastically with DevOps teams, customers, and peers across our organization
  • Ability to communicate directly with software vendors
  • Demonstrate a tolerance for stress and provide a supportive attitude for all colleagues

Desired Characteristics:

  • BS in computer science or related field
  • 2+ years of site reliability experience supporting high volume/large-scale environments
  • Media industry experience (broadcast or production) preferred
  • Cloud platform certification preferred

Additional Requirements:

  • Fully Remote: This position has been d

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

NBCUniversal

View company profile →