Director, Infrastructure & SRE
TailorCareAbout the role
<p><strong>About the Role</strong></p> <p>The Director of Infrastructure & SRE owns the function end-to-end: reliability, security, scalability, and operational governance of TailorCare’s infrastructure, plus the team that delivers it. You will be a peer to the Director of Software Engineering, Director of Data Engineering, and Director of Data Science, own the Infrastructure & SRE scorecard in front of the executive team, and lead vendor escalations with Salesforce, AWS, and Cresta, among others, at the Director level.</p> <p>This is a player-coach role. In year one you will spend roughly 60% of your time hands-on (writing Terraform, leading incidents, doing architecture work) and 40% building the team and the practice. As the team scales, that ratio shifts toward leadership, but you will never stop being technical.</p> <p>This is not a slideware role. We are not hiring a manager who reviews architecture diagrams from a distance. We are hiring an operator who codes, runs incidents, owns the platform, and ships</p> <p><strong>Primary Responsibilities</strong></p> <p><strong>Infrastructure as Code</strong></p> <ul> <li>Converge all AWS resources to Terraform; eliminate manual provisioning</li> <li>Establish reproducible environments (dev, staging, production) with proper isolation and parity</li> <li>Standardize CI/CD pipelines across all engineering teams</li> </ul> <p><strong>Site Reliability</strong></p> <ul> <li>Define and operate SLOs, SLIs, and error budgets for all production systems (web/mobile applications, Salesforce, data processing, telephony stack)</li> <li>Build observability (metrics, logs, traces, alerting) across AWS, Salesforce, telephony/omni-channel, and Cresta integrations</li> <li>Stand up the infrastructure on-call rotation, incident management, and post-incident review discipline, including RCAs</li> <li>Own uptime, MTTR, and incident-volume trends as published metrics</li> </ul> <p><strong>Disaster Recovery & Business Continuity</strong></p> <ul> <li>Design and implement a tested DR strategy with documented RPO/RTO commitments</li> <li>Validate recovery procedures on a recurring cadence</li> <li>Align DR posture with HITRUST and HIPAA expectations</li> </ul> <p><strong>Integration Reliability</strong></p> <ul> <li>Stabilize Salesforce, telephony/omni-channel, and Cresta integrations; close persistent gaps in skills-based routing, warm transfers, and telephony data parity</li> <li>Partner with Data Engineering on the reliability of data ingest paths (Fivetran, SFTP, S3) and Salesforce bulk API flows.</li> </ul> <p><strong>S
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s