Jobs and Careers
RO

Senior Site Reliability Engineer - NYC

Rokt
New York City, United Statesfull_timeVerifiedPosted 6 Jul 2023
💰 $300,000/yr($200,000/yr$300,000/yr)

About the role

About Rokt

Rokt is the global leader in ecommerce technology, helping companies seize the full potential of every transaction moment to grow revenue and acquire new customers at scale. Live Nation, Groupon, Staples, Lands' End, UrbanStems, GoDaddy, Vistaprint and HelloFresh are among the more than 2,500 leading global businesses and advertisers that are using Rokt's solutions to drive more value through every transaction by offering highly relevant messages to their customers at the moment they are most likely to convert.

With our December 2021 Series E raise of USD$325M, Rokt is expanding rapidly and globally – operating in 19 countries across North America, Europe and the Asia-Pacific region with the largest office in NYC and a major R&D hub in Sydney. With annual revenues of more than US$200M and vibrant company culture, Rokt has been listed in ‘Great Places to Work’ in the US and Australia. Our award-winning culture is guided by our five core values: Smart with Humility, Own the Outcomes, Force for Good, Conquer New Frontiers, and Enjoy the Ride. These values help us attract, engage, and develop the right talent around the globe and ensure we have the right conditions to do our best work. Keen to join a fast-growing company and a vibrant culture? Learn more at rokt.com.

The Rokt engineering team builds best-in-class ecommerce technology that provides personalized and relevant experiences for customers globally and empowers marketers with sophisticated, AI-driven tooling to better understand consumers. Our bespoke platform handles millions of transactions per day and considers billions of data points which give engineers the opportunity to build technology at scale, collaborate across teams and gain exposure to a wide range of technology. We are expanding rapidly in our major R&D centers in NYC and Sydney. We are passionate about using intelligent systems to improve the transaction moment for retailers everywhere. Come join us and build the future!


The Role

As a Senior Site Reliability Engineer you will be part of a team responsible for designing and building high levels of availability, scalability and reliability into our systems. You will become intimate with the architecture of our systems and be responsible for diving deep into code, lead architecture and root cause analysis workshops working directly with feature teams.


Responsibilities

  • Evolve systems by pushing for changes that improve reliability and latency
  • Our day-to- day is driven by helping our product teams create robust software faster.
  • Introduce best practices into the teams around observability, SLOs and reliability.
  • Work in close collaboration with partner teams to shape the future roadmap to improve reliability and establish strong operational readiness across teams.
  • Participate in system design consulting, and capacity planning.
  • Identify areas for improvement across the organization and drive Engineering-wide technical change in the field of Site Reliability.
  • Share your knowledge by giving brown bags, tech talks, and evangelizing appropriate tech and engineering best practices.
  • Partner with the broader Rokt organization to build a culture of rigorously learning from incidents.
  • Contribute to Root Cause Analysis (RCA) investigations and follow up each incident to ensure the appropriate action items are in place and prioritized.
  • Designing tools to help our entire engineering organization be as productive as possible.
  • Lead development and roll out of new tools, technologies and processes that have high business impact and are used by multiple teams that improve reliability and velocity.
  • Contribute to documentation and uplifting of partner teams

Requirements

  • Bachelor’s degree or equivalent practical experience.
  • 5 years hands-on experience in Site Reliability and Observability Engineering, debugging, diagnosing and correcting errors and resolving high severity incidents
  • Commercial experience in one of the following languages Java, C#, Python or Go.
  • Think about systems - edge cases, failure modes, behaviors, specific implementations.
  • Experience building solutions in distributed systems for high volume transaction and/or developing support focused tooling.
  • Solid experience with cloud infrastructure and tooling (AWS, GCE, Azure, Kubernetes, Docker, CI/CD pipelines & Terraform).
  • Experience in Defensive programming, Circuit breakers, Resilience frameworks, Fault tolerance, and self-healing mechanisms of services.
  • Experience working on various monitoring, and alerting tools
  • An ability and desire to mentor and coach engineers.
  • Strong organizational and interpersonal skills, with experience devel

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Rokt

View company profile →