Jobs and Careers
BO

Senior Software Engineer, Reliability

Box
Redwood City, United Statesfull_timeVerifiedPosted 5 May 2025
💰 $187,000/yr

About the role

WHAT IS BOX? 

Box (NYSE:BOX) is the leader in Intelligent Content Management. Our platform enables organizations to fuel collaboration, manage the entire content lifecycle, secure critical content, and transform business workflows with enterprise AI. We help companies thrive in the new AI-first era of business. Founded in 2005, Box simplifies work for leading global organizations, including AstraZeneca, JLL, Morgan Stanley, and Nationwide. Box is headquartered in Redwood City, CA, with offices across the United States, Europe, and Asia.

By joining Box, you will have the unique opportunity to continue driving our platform forward. Content powers how we work. It’s the billions of files and information flowing across teams, departments, and key business processes every single day: contracts, invoices, employee records, financials, product specs, marketing assets, and more. Our mission is to bring intelligence to the world of content management and empower our customers to completely transform workflows across their organizations. With the combination of AI and enterprise content, the opportunity has never been greater to transform how the world works together and at Box you will be on the front lines of this massive shift.

Founded in 2005, Box is headquartered in Redwood City, CA, and we have offices across the United States, Europe, and Asia.

WHY BOX NEEDS YOU

The Reliability Team focuses on building frameworks and systems to enhance availability, reliability and resilience of Box systems. We design high-performance, low-latency, high-throughput services, promote best practices, and engage in architectural design to embed reliability into every layer of our products.

At Box, reliability and customer experience are top priorities, driving a strong need for deep observability and robust distributed system design. We seek your expertise in distributed systems, resilience engineering, and large-scale production operations — to identify gaps, design and build solutions, and guide product teams towards building highly available and resilient services. Your work will directly strengthen our SRE strategy, operational excellence, system performance, and reliability culture.

We are seeking innovative problem-solvers passionate about large-scale distributed systems and eager to grow their skills in modern SRE practices. As a small team tackling complex challenges at scale, we offer the opportunity to make significant technical contributions while driving observability culture across the organization.

WHAT YOU'LL DO

  • You will build software, frameworks, and tools required for reliable operations of Box's services across multiple cloud environments
  • You will manage the stability and operation of several of Box's most critical production applications through application reviews, capacity planning, and performance tuning
  • You will be constantly developing automations / frameworks / tools for better platform reliability/resilience/availability
  • You will collaborate with other engineers on the team as well as cross functionally to foster solid software engineering principles and represent our engineering values
  • You will participate in various POCs on new projects and frameworks being evaluated for the product/platforms
  • You will improve our observability as both a developer/maintainer of systems/frameworks, and a mentor to our product development teams
  • You will work with modern cloud-native technologies including container orchestration (Kubernetes, Docker), service mesh solutions (Istio, Linkerd), and cloud platforms (AWS, GCP)
  • You will participate in product design reviews and architectural discussions to ensure reliability is considered early in the development lifecycle of product/services
  • You will participate in a team on-call rotation

WHO YOU ARE 

  • 5+ years of working experience designing, developing, and operating large-scale, customer-facing products or services
  • Experience coding in higher-level languages (e.g., Java, Scala, Go, Python) is preferred
  • A strong interest in solving challenging problems using innovative and data-driven approaches
  • An SRE-centric mindset — you build and manage systems with reliability, scalability, availability, and security as core principles
  • Experience designing complex systems and frameworks using proven system design principles, such as NALSD (Non-Abstract Large System Design) methodologies
  • Experience troubleshooting issues across distributed Linux environments, with comfort tracing problems across applications, systems, and networks
  • Proficient with modern cloud technologies such as GCP, AWS, and Kubernetes
  • Experienced in service observability practices and tools (e.g., Prometheus, OpenTelemetry, S

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Box

View company profile →