Jobs and Careers
VE

Senior Software Engineer - Site Reliability

Vertex Inc.
Remote - OR, United States, United StatesRemotefull_timeVerifiedPosted 23 May 2025

About the role

Job Description:

The Software Site Reliability Engineer at Vertex ensures enterprise-wide systems are reliable, scalable, and performant by relentlessly measuring and improving environments. They lead and guide teams to implement new software and system capabilities, enhance code, and optimize processes and tools. Leveraging deep infrastructure and software engineering expertise, they build reliable solutions from inception or refactor legacy systems for improved reliability. Success is driven by data, customer satisfaction, and empowering teams to achieve excellence. 

ESSENTIAL JOB FUNCTIONS AND RESPONSIBILITIES 

  • Drive Reliability: Drive initiatives that enhance system reliability and operational efficiency, guiding teams in implementing code and system design reliability improvements and efficiencies. 

  • Design Optimization: Guide teams in designing, developing, implementing optimized and efficient systems and environments ensuring performance, reliability, and scalability. 

  • Observation and Alerting: Influence teams in designing and implementing applications and systems that put reliability, monitoring, alerting, and analytics first. 

  • Performance Metrics: Guide teams in measuring the health and performance of environments using observability tools, ensuring accurate and actionable metrics. 

  • Culture of Reliability: Foster a culture of reliability and operational excellence through mentorship and training, ensuring consistent implementation of SRE principles. 

  • Incident Management: Guide teams in triaging, isolating, and resolving environmental issues expediently and openly according with incident response protocols and procedures. 

  • Proactive Resolution: Guide teams to anticipate and correct production issues, including outages, processing slowdowns, errors, and failures, using incident management best practices. Ensure teams minimize downtime and ensure rapid recovery. 

  • Technical Leadership: Provide technical leadership for projects, ensuring solutions align with reliability best practices and organizational goals. 

  • Standards & Practices: Develop and publish standards and best practices, guiding teams to implement observability and monitor system performance effectively. 

  • Reliability Feedback: Capture and document engineering and operations case studies to refine published SRE software policies and best practices. 

  • CI/CD Reliability: Guide teams in building and delivering reliability starting from Continuous Integration (CI) and Continuous Deployment (CD) processes, ensuring robust and reliable software delivery pipelines. 

  • Agile Practices: Participate in the plan, prioritization, and breakdown of team deliverables to ensure that they deliver on reliability and quality organizational outcomes. 

  • Mentorship: Guide and mentor organizational software engineering staff, developing their technical skills and knowledge of Site Reliability patterns and practices. 

KNOWLEDGE, SKILLS AND ABILITIES 

Candidate must possess Advanced proficiency of the following: 

Technical 

  • Design and delivery of highly reliable SaaS solutions hosted in AWS, Azure, OCI, or GCP 

  • Software Development frameworks using Java, Spring Boot, .NET Core, MVC, JavaScript  

  • Designing and delivering highly observed, reliable and recoverable enterprise event-driven systems 

  • Deep observability and monitoring experience with Open Telemetry, Datadog, and CloudWatch 

  • Infrastructure, application and synthetic monitoring and alerting techniques and patterns 

  • Institutionalization of application and system metrics with KPIs, SLIs and SLOs 

  • Observable and reliable relational storage solutions with Postgres, MSQL, or similar 

  • Observable and reliable non-relational database technologies and cloud storage like AWS S3 

  • Observable and reliable containerization apps in Kubernetes, ArgoCD, Helm and TF 

  • CI powered performance and synthetics augmenting shift-left testing strategy methods  

  • CD experience using GitHub Actions, Terraform, Go, PowerShell and/or Python  

  • Exposure to AI automation paired programing with GitHub Copilot or similar tools 

  • Scaling application optimization for Network, Memory and IO performance concerns 

Interpersonal 

  • Results-oriented and customer-focused, acting with urgency and purpose. 

  • Ability to make data-driven decisions guided by commitment to customer outcomes. 

  • Strong time manageme

Apply for this role

Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.

Apply Now →Generate Application Kit

Free account required — sign up in 30s

Company

Vertex Inc.

View company profile →