Senior SRE Engineer
AkoyaAbout the role
The Role
This is a 6-month contract role, with a potential to convert to full-time on the Platform Engineering team, working closely with global development, operations, and delivery teams. Platform Engineering owns the reliability, availability, and performance of our infrastructure and services across a multi-account, multi-tenant AWS environment spanning multiple regions. The role balances hands-on technical work with team leadership: driving initiatives, mentoring engineers, owning SLA/SLO adherence, leading incident management, and automating toil out of day-to-day operations.
Ready to lead reliability at scale? If you thrive in fast-paced environments and are passionate about system reliability, security, and operational excellence, we invite you to join our Platform Engineering team.
As a Senior SRE Engineer, you will:
- Design, implement, and maintain scalable systems for uptime, resilience, and performance, while defining and enforcing SLOs, SLIs, and SLAs with product teams.
- Own incident detection, escalation, and resolution, developing playbooks for failure scenarios and leading post-mortem analyses with corrective actions.
- Build monitoring, logging, and alerting systems, and develop automation for deployments, maintenance, and toil reduction.
- Drive CI/CD improvements and maintain operational documentation, runbooks, and best practices.
- Partner with developers on reliability by design, fostering shared responsibility for observability and fault tolerance.
- Enforce security policies and standards, collaborating on vulnerability identification and mitigation strategies.
Required Experience & Skills
- AWS Cloud Services: Deep working knowledge across compute, networking, security, storage, and data services in multi-account, multi-region environments.
- Kubernetes/EKS: Experience deploying and maintaining Kubernetes with multi-AZ, multi-region, and multi-tenant topologies.
- Hybrid Networking: Expertise in VPC design, VPN connectivity, on-prem/hybrid topology, and edge solutions.
- CI/CD Pipelines: Experience building and maintaining pipelines using CodePipeline, CodeBuild, ECR, and ArgoCD.
- Infrastructure Design: Ability to design scalable, resilient infrastructure with security, observability, and automation built in.
- Cert/PKI Management: Knowledge of certificate and PKI lifecycle management, including HSM-backed key protection and mTLS.
- Security & Cryptography: Strong understanding of encryption, cryptographic protocols, and HSM integration.
- Scripting & Automation: Proficiency in Python, Go, and Bash for automation and infrastructure-as-code.
- Compliance as Code: Experience using AWS Config to translate policy requirements into automated compliance checks.
- Observability Platforms: Hands-on experience with Datadog for metrics, distributed tracing, log management, and APM.
- Chaos Engineering: Experience applying chaos engineering practices to proactively validate system resilience.
Preferred Skills & Attributes
- Tactical & Operational: Ability to drive team initiatives, own delivery outcomes, and act as a reliability champion.
- System Reliability Advocacy: Champion robust system design and operational excellence across engineering and product teams.
- Analytical Problem-Solving: Skilled at diagnosing and resolving complex, multi-layer issues in distributed, cloud-native environments.
- Cross-Functional Communication: Effective collaborator across development, security, and operations teams in a global organization.
- Incident Response Leadership: Experience leading swift resolution of high-severity incidents to minimize downtime.
- Culture of Ownership: Fosters accountability and a blameless culture focused on reliability and continuous improvement.
- Adaptability: Comfortable problem-solving creatively in fast-moving environments with shifting priorities.
The Environment You’ll Work In
The Platform Environment
- Multi-account AWS architecture: Per-client account isolation for highest-tier tenants; shared-account segregation for standard tenants.
- Multi-region deployment: Active-active or active-passive topology for disaster recovery and geographic resilien
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s