Principal SRE (Network Performance)
GartnerAbout the role
Hiring near our Barcelona, Spain Center of Excellence.
Hybrid, flexible environment.
Gartner offers a hybrid, flexible environment, with remote work that allows associates great flexibility to work from home, and opportunities to connect with colleagues for moments that matter on-site. Candidates that apply should be located within a reasonable proximity to one of Gartner’s Centers of Excellence office locations.
About Gartner IT:
Join a world-class team of skilled engineers who build creative digital solutions to support our colleagues and clients. We make a broad organizational impact by delivering cutting-edge technology solutions that power Gartner. Gartner IT values its culture of nonstop innovation, an outcome-driven approach to success, and the notion that great ideas can come from anyone on the team.
About this role:
We are seeking a Principal SRE who is Network focused and will play a crucial role in supporting the production and operations of our conferences IT platforms and our enterprise network (both cloud & on prem). During live conferences, the candidate will work closely with network operations team members to ensure smooth operation of the network infrastructure by optimizing and maintaining its high performance and reliability. Additionally, during non-conference periods, the candidate will focus on ensuring the operational readiness of the network infrastructure. This includes Observability, performance and resiliency utilizing Chaos Engineering techniques.
What you’ll do
As part of SRE scrum team, troubleshoot and resolve complex network performance and reliability issues, working closely with network operations and engineering teams
Function as SME in utilizing NPM tools to drive forensics on network performance and reliability issues that impact optimal end user experience and / or application health
Work closely with the Conference Network team and also Enterprise Network team to maintain / enhance a comprehensive knowledge of the systems and infrastructure
Work closely with the Observability team to maintain / improve dashboard / alerting posture
Monitor operational dashboards and alerts during conferences and respond to alerts
Collaborate to develop / design chaos test cases that effectively simulate real-world scenarios, identify potential vulnerabilities and areas for improvement
Execute chaos tests, analyze using NPM, APM and other monitoring tools to identify performance and stability issues
Utilize breadth of knowledge and experience to accurately connect the dots between application and network performance issues
Utilize strong network forensics knowledge to cross train other IT engineers
Use data driven analysis to drive continuous improvement in network observability, performance, reliability and resilience
Perform analytics on previous incidents to understand root causes and use automation to detect problems faster, reduce the probability and/or impact of problem recurrence where possible
Support and drive advancement of our NPM tools and services
Available to work flexible hours as required for operational support and during select conferences to ensure coordination among globally distributed team
Participate in on-call schedule, ensuring that issues are addressed promptly and effectively
What you’ll need:
12+ years of information technology experience, with a desired 7+ years in a relevant network performance engineering role or similar position providing comprehensive support of critical multi-tier applications
Deep understanding of the following technologies and concepts:
AWS/Azure cloud, on-prem and it’s most commonly used network related services
OSI Framework, network protocols, network topologies, network components
OS and corresponding performance indicators; Windows Server & Linux
Front-end web analysis/optimizations including caching, CDNs, compression, SPOFs
Foundational understanding of N-tier architecture and its components
Proficiency in network packets analysis
5+ years’ experience in working with NPM tools like Extrahop, Riverbed AppResponse, LiveAction, etc.
Experience working with Network Monitoring tools like PRTG, Solarwinds, etc.
Stay updated with industry trends and emerging technologies related to network performance and reliability optimization
Experience in troubleshooting using APM tools such as Dynatrace or AppDynamics is desired
Exceptional analytical, diagnostic, and problem-solving skills
Ability to work
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s