Site Reliability Engineer - Production Sandbox Reliability team
DeelAbout the role
Who we are is what we do.
Deel is the all-in-one payroll and HR platform for global teams. Our vision is to unlock global opportunity for every person, team, and business. Built for the way the world works today, Deel combines HRIS, payroll, compliance, benefits, performance, and equipment management into one seamless platform. With AI-powered tools and a fully owned payroll infrastructure, Deel supports every worker type in 150+ countries—helping businesses scale smarter, faster, and more compliantly.
Among the largest globally distributed companies in the world, our team of 6,000 spans more than 100 countries, speaks 74 languages, and brings a connected and dynamic culture that drives continuous learning and innovation for our customers.
Why should you be part of our success story?
As the fastest-growing Software as a Service (SaaS) company in history, Deel is transforming how global talent connects with world-class companies – breaking down borders that have traditionally limited both hiring and career opportunities. We're not just building software; we're creating the infrastructure for the future of work, enabling a more diverse and inclusive global economy. In 2024 alone, we paid $11.2 billion to workers in nearly 100 currencies and provided healthcare and benefits to workers in 109 countries—ensuring people get paid and protected, no matter where they are.
Our momentum is reflected in our achievements and customer satisfaction: CNBC Disruptor 50, Forbes Cloud 100, Deloitte Fast 500, and repeated recognition on Y Combinator’s top companies list – all while maintaining a 4.83 average rating from 15,000 reviews across G2, Trustpilot, Captera, Apple and Google.
Your experience at Deel will be a career accelerator. At the forefront of the global work revolution, you'll tackle complex challenges that impact millions of people's working lives. With our momentum—backed by a $12 billion valuation and $1 B in Annual Recurring Revenue (ARR) in just over five years—you'll drive meaningful impact while building expertise that makes you a sought-after leader in the transformation of global work.
About the Role
We’re looking for a Junior to Mid-level Site Reliability Engineer to join our Production Sandbox Reliability team – the first line of defense ensuring enterprise customer sandboxes remain stable, up-to-date, and well-monitored. You'll work in a globally distributed team and act as the operational guardian for these environments.
This role is ideal for someone who thrives in high-ownership environments, enjoys hands-on infrastructure work, and wants to grow their SRE skills while partnering closely with engineering and customer-facing teams.
Key Responsibilities
Maintain reliability: Monitor the health of enterprise customer sandbox environments and ensure high availability, uptime, and stability across all services
Stay up-to-date: Regularly roll out updates to microservices inside each sandbox to ensure alignment with the latest versions
Alert response & escalation: Triage infrastructure and application alerts, perform initial investigation and escalate incidents to the appropriate engineering teams with clear context
Improve observability: Enhance metrics, logs and tracing coverage using Datadog and the Grafana stack (Mimir, Loki, Tempo), identifying gaps and driving better alerting practices
Support incident workflows: Collaborate in post-incident reviews and ensure root cause analysis is followed up with actionable items and improvements by relevant teams
Communicate proactively: Act as the bridge between internal engineering teams and customer-facing teams, providing timely updates during incidents, maintenance and version upgrades
Participate in on-call rotation: Provide continuous coverage across APAC, EMEA and LATAM time zones as part of a rotating on-call schedule (follow the sun)
What We're Looking For
1-3 years of experience in SRE, DevOps or Infrastructure Engineering roles
Experience with Node.js or Go
Familiarity with AWS cloud services (EKS, S3, RDS)
Hand-on experience with Kubernetes, including Helm and ArgoCD
Experience with observability stacks: Datadog, Grafana, Mimir, Loki, Tempo, Zabbix
Strong verbal and written communication skills – able to interface effectively with both technical and non-technical stakeholders
Self-starter mindset with an eye for operational excellence and continuous improvement
Total Rewards<
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s