SRO Lead
VersantAbout the role
Company Description
VERSANT (Nasdaq: VSNT) is an industry-changing media and entertainment business and home to trusted brands that shape culture, inform audiences, and build lasting connections. It operates across four core markets: political news and opinion, business news and personal finance, golf, and sports and genre entertainment. These markets are served through a powerful portfolio of iconic and innovative brands, including CNBC, MS NOW, USA Network, Golf Channel, Oxygen, E!, SYFY, and Versant's sports division USA Sports, along with complementary digital assets including Fandango, Rotten Tomatoes, GolfNow and GolfPass.
Job Description
The System Reliability Engineering (SRE) Lead is a hands-on technical leader responsible for improving the reliability, performance, and scalability of VERSANT’s software, production, and platform systems.
Reporting to the VP of Infrastructure, this role works closely with Software Engineering, Production Engineering, Platform Engineering, and Infrastructure teams to implement reliability best practices, drive end-to-end testing, and ensure systems perform under real-world conditions.
This is a player-coach role focused on execution—building testing frameworks, improving observability, and helping teams proactively identify and resolve system weaknesses before they impact production.
Key Responsibilities
Reliability & Engineering Practices
- Partner with engineering teams to improve system reliability, availability, and performance.
- Help define and implement SLIs, SLOs, and basic reliability standards across services.
- Identify reliability gaps and work with teams to address risks in system design and operations.
- Contribute directly to code, tooling, and automation that improves system resilience.
E2E & System Testing Execution
- Design and implement end-to-end (E2E) testing workflows across distributed systems.
- Build and maintain integration testing frameworks validating cross-service dependencies.
- Execute and scale load and performance testing to validate systems under peak conditions.
- Partner with teams to integrate automated testing into CI/CD pipelines.
- Help establish practical testing standards and ensure adoption across teams.
Performance & Capacity
- Support performance benchmarking and system capacity planning efforts.
- Analyze system performance and identify bottlenecks across application and infrastructure layers.
- Partner with infrastructure and platform teams to optimize system throughput and latency.
Observability & Operations
- Implement and improve monitoring, logging, and alerting across services.
- Help ensure systems are observable, debuggable, and well-instrumented.
- Participate in incident response and support root cause analysis efforts.
- Contribute to post-incident reviews and track follow-up actions to improve reliability.
Cross-Team Collaboration
- Work closely with software, platform, enterprise and production engineering teams to embed reliability practices into day-to-day development.
- Provide guidance and hands-on support for testing, observability, and performance improvements.
- Help standardize tools, frameworks, and processes used across teams.
- Mentor engineers on reliability engineering fundamentals and testing best practices.
Qualifications
Basic Requirements
- 5+ years of experience in Site Reliability Engineering, DevOps, Systems Engineering, or Infrastructure roles.
- Strong hands-on experience operating and troubleshooting production systems.
- Experience implementing integration testing, E2E testing, or performance/load testing frameworks.
- Familiarity with observability tools (metrics, logging, tracing) and monitoring systems.
- Experience with cloud platforms (AWS, GCP, or Azure) and distributed systems.
- Experience working with CI/CD pipelines and automation.
- Strong debugging and problem-solving skills in complex systems.
Desired Characteristics
- Experience supporting media, broadcast, or real-time production systems.
- Familiarity with high-throughput or low-latency systems.
- Exposure to SRE concepts such as SLIs/SLOs and incident management practices.
- Experience with containerized environments (Kubernetes, Docker) is a plus.
- Familiarity with infrastructure as code (Terraform, CloudFormation).
- Strong collaboration skills and ability to work across multiple engineering teams.
Additional Information
As part of our selection
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s