Jobs and Careers
GA

Senior Site Reliability Engineer (APM) (Irving)

Gartner
United Statesfull_timeVerifiedPosted 12 Jun 2024
💰 $149,000/yr($100,000/yr$149,000/yr)

About the role

About Gartner IT:

Join a world-class team of skilled engineers who build creative digital solutions to support our colleagues and clients.  We make a broad organizational impact by delivering cutting-edge technology solutions that power Gartner.  Gartner IT values its culture of nonstop innovation, an outcome-driven approach to success, and the notion that great ideas can come from anyone on the team. 

About this role:

The person will primarily be responsible for supporting production or operations when conferences are live and work as part of the conference's SRE team. Between conferences they will be responsible for identifying issues in production, triaging identified issues, partnering with other engineers on the team to identify the root cause. Other responsibilities include managing applications and infrastructure as a code, creating & executing chaos tests, managing alerts & dashboards.

What you’ll do

  • As part of the SRE scrum team, perform full stack triaging of alerts and engage other engineers  to identify root cause of application performance & stability issues.

  • Collaborate with cross functional members of swat team during production incidents and provide critical technical insight as well as thought leadership to identify the root cause.

  • Conduct blameless post mortems to troubleshoot priority incidents.

  • Establish Relationships with stakeholders such as development teams or product owners to define service level objectives (SLOs) for application features/services.

  • Measure performance against SLOs in partnership with stakeholders, and ensure systems continue to meet SLOs over time.

  • Design, develop dashboards and reports to communicate key metrics.

  • Identify opportunities to improve alerting posture and create/update alerts accordingly.

  • Work closely with the Application team to understand application architecture and perform Single point of failure analysis and create scenarios for testing resiliency of the application. Assist in Game Day preparation and execution. 

  • Oversee, design, implement, and manage DevOps capabilities using continuous integration/continuous delivery toolsets and automation.

  •  Identify opportunities to automate manual operational work (i.e., “toil”) using pipelines or by using new software or any other appropriate mechanisms.

  • Perform analytics on previous incidents to understand root causes and use automation to reduce the probability and/or impact of problem recurrence.

  • Create/derive NFR/Workload model and ensure performance & resiliency is considered early in the SDLC. 

  • Execute performance/chaos tests,  analyze using APM and other tools to identify performance & stability issues.

  • Document any findings/analysis/results, communicate and present to stakeholders.

  • Available to work flexible hours and travel as required for preparation and operational support of select events like releases or conferences

  • Participate in on-call schedule, ensuring that issues are addressed promptly and effectively.


 

What you’ll need:

  • Bachelor’s degree or foreign equivalent degree in Computer Science or a related field required

  • 6+ years of information technology experience with 4+ years working on DevOps or SRE team or performance engineering team

  • Experienced in triaging of production issues using APM tools such as Dynatrace or AppDynamics or New Relic and log aggregation tools such as Splunk, ELK, etc. 

  • Skilled in collaborating with Dev/DBA/Architecture teams or other relevant teams and performing root cause analysis with good working knowledge of application, processes, operating system

  • Experience with SRE concepts like SLI/SLOs & error budgets

  • Experience with AWS cloud, specifically services such as EC2, EKS, API GW, Lambda, Route53, SNS, RDS, Elasticcache, OpenSearch, etc. or similar cloud technologies & services

  • Knowledge of Docker containers and related orchestration technologies 

  • Excellent analytical, verbal & written communication skills with data driven analysis

Nice to Have:

  • Experience with CDN such as Cloudflare and various features like bot management features, Advanced WAF, Coding at edge

  • Experience with CI/CD processes and tools ( Jenkins, Argo, Harness, etc.)  

  • Experience with chaos engineering   

  • Experience with Agile and DevOps development methodologies

  • Exposure to automation and scripting skills using jenkins, python, shell, etc.

  • Knowledge of Infrastructure as a code using terraform.

  • Ability t

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Gartner

View company profile →