Jobs and Careers
TW

Site Reliability Engineer

Twilio
Remote - US, United StatesRemotefull_timeVerifiedPosted 16 Jul 2025
💰 $224,200/yr($161,500/yr$224,200/yr)

About the role

Who we are 

At Twilio, we’re shaping the future of communications, all from the comfort of our homes. We deliver innovative solutions to hundreds of thousands of businesses and empower millions of developers worldwide to craft personalized customer experiences.

Our dedication to remote-first work, and strong culture of connection and global inclusion means that no matter your location, you’re part of a vibrant team with diverse experiences making a global impact each day. As we continue to revolutionize how the world interacts, we’re acquiring new skills and experiences that make work feel truly rewarding. Your career at Twilio is in your hands.

See yourself at Twilio

Join the team as Twilio’s next Site Reliability Engineer on our Data Infrastructure Platform.

About the job

We are looking for a talented and experienced Software Engineer to join our Data Platform team. In this role, you will play a crucial part in designing, building, and optimizing our platform to support a wide range of data-driven initiatives. You will work closely with cross-functional teams to understand business requirements, architect scalable solutions, and implement data solutions and infrastructure for our Data Platform. The ideal candidate will have a passion for leveraging data to drive business impact, strong technical skills, and experience with modern data technologies.

Responsibilities

In this role, you’ll:

  • Design, build, and maintain infrastructure and scalable frameworks to support data ingestion, processing, and analysis.
  • Collaborate with stakeholders, analysts, and product teams to understand business requirements and translate them into technical solutions.
  • Architect and implement data streaming solutions using modern data technologies such as Kafka, AWS MSK, Terraform, Hive, Hudi, Presto, Airflow, and cloud-based services like AWS EKS, Lakeformation, Glue and Athena. 
  • Design and implement frameworks and solutions for performance, reliability, and cost-efficiency.
  • Ensure data quality, integrity, and security throughout the data lifecycle.
  • Stay current with emerging technologies and best practices in big data technologies
  • Mentor early in career engineers and contribute to a culture of continuous learning and improvement

Qualifications 

Twilio values diverse experiences from all kinds of industries, and we encourage everyone who meets the required qualifications to apply. If your career is just starting or hasn't followed a traditional path, don't let that stop you from considering Twilio. We are always looking for people who will bring something new to the table!

*Required:

  • Bachelor's or Master's degree in Computer Science, Engineering, or a related field.
  • 8+ years of experience in Site Reliability Engineering, DevOps, or Software Engineering roles with a focus on infrastructure or backend systems.
  • Strong production experience, including operational management, scaling, partitioning strategies, and tuning for performance and reliability.
  • Hands-on experience with Kubernetes (preferably EKS), including deploying and managing stateful services and operators in Kubernetes environments.
  • Deep understanding of AWS cloud services, particularly those relevant to data infrastructure (e.g., EC2, EBS, S3, IAM, MSK, CloudWatch, VPC, ALB/NLB).
  • Proficiency in infrastructure-as-code tools, such as Terraform or CloudFormation, for managing and automating infrastructure.
  • Expertise in observability tools (e.g., Prometheus, Grafana, OpenTelemetry, Datadog) to monitor distributed systems and set up alerting for reliability and latency.
  • Proficient in at least one programming language (e.g., Go, Python, Java, or similar) for building automation, tooling, and contributing to platform services.
  • Experience designing and implementing incident response processes, SLOs/SLIs, runbooks, and participating in on-call rotations.
  • Strong understanding of distributed systems principles, including consensus, durability, throughput, and availability tradeoffs.
  • Proven track record of driving reliability improvements in high-scale, data-intensive systems and collaborating with platform and data engineering teams.
  • Excellent problem-solving and analytical skills.
  • Strong verbal & written communication skills, with the ability to work effectively in a cross-functional team environment.

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Twilio

View company profile →