Senior Staff Engineer, Observability Engineer
CoupangAbout the role
We exist to wow our customers. We know we’re doing the right thing when we hear our customers say, “How did we ever live without Coupang?” Born out of an obsession to make shopping, eating, and living easier than ever, we’re collectively disrupting the multi-billion-dollar e-commerce industry from the ground up. We are one of the fastest-growing e-commerce companies that established an unparalleled reputation for being a dominant and reliable force in South Korean commerce.
We are proud to have the best of both worlds — a startup culture with the resources of a large global public company. This fuels us to continue our growth and launch new services at the speed we have been since our inception. We are all entrepreneurial surrounded by opportunities to drive new initiatives and innovations. At our core, we are bold and ambitious people that like to get our hands dirty and make a hands-on impact. At Coupang, you will see yourself, your colleagues, your team, and the company grow every day.
Our mission to build the future of commerce is real. We push the boundaries of what’s possible to solve problems and break traditional tradeoffs. Join Coupang now to create an epic experience in this always-on, high-tech, and hyper-connected world.
Summary: As a Sr. Staff Back-end Engineer within the Site Reliability organization, you will be working with large scale cloud infrastructure handling billions of metrics and peta-bytes of logs and metrics.
You will leverage this data to help internal teams to monitor service reliability and predict/prevent incidents. You have the opportunity to build the next generation Observability Platform based on Kubernetes and other OSS solutions, as well as building software components from scratch. You would work directly with various engineering teams in Coupang, influence them with SRE principles and best practices and see your impact directly.
Key Responsibilities:
- Design, implement, and maintain observability solutions such as monitoring, alerting, logging, and tracing across various platforms, applications, and infrastructure.
- Collaborate with cross-functional teams, including software engineers, SREs, and infrastructure teams, to identify and define observability requirements.
- Develop and implement best practices for creating and maintaining effective monitoring, alerting, and telemetry systems.
- Evaluate and recommend industry-leading observability tools and technologies to improve system visibility and reliability.
- Define and track key performance indicators (KPIs) and service-level objectives (SLOs) related to system availability, performance, and reliability.
- Assist in the troubleshooting and resolution of complex incidents and problems by analyzing data from observability tools.
- Provide guidance and mentorship to other engineers on observability principles, practices, and tools.
- Conduct ongoing evaluations of observability systems and identify opportunities for improvements and optimizations.
- Drive the standardization and simplification of observability processes, tools, and frameworks across the organization.
- Contribute to the development of training materials, documentation, and runbooks for observability systems and practices.
Essential Qualifications:
- Bachelor's Degree in Computer Science, Engineering, or a related technical field.
- Strong experience in implementing and managing observability solutions in large-scale, complex environments.
- Deep knowledge of monitoring, alerting, and logging systems and tools, such as Prometheus, Grafana, Elastic Stack, Datadog, or New Relic.
- Familiarity with distributed tracing technologies, such as Jaeger or Zipkin.
- Experience with cloud-based infrastructure, including AWS, Azure, or Google Cloud Platform.
- Strong understanding of DevOps and SRE practices, including continuous integration, continuous delivery, and infrastructure as code (IaC).
- Proficiency in scripting languages, such as Python, Bash, or Ruby.
- Excellent communication and collaboration skills, with the ability to work with teams across different functions and technical domains.
- Strong problem-solving and analytical skills, with a focus on data-driven decision-making.
- A proven track record of leading and delivering successful observability projects and initiatives.
Preferred Qualifications:
- Experience with containerization and orchestration technologies, such as Docker and Kubernetes.
- Familiarity with application performance management (APM) tools, such as Dynatrace or AppDynamics.
- Professional certifications in cloud platforms, monitoring tools, or related technologies.
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s