Senior SRE
Blackpoint CyberAbout the role
Blackpoint Cyber is the leading provider of world-class cybersecurity threat hunting, detection and remediation technology. Founded by former National Security Agency (NSA) cyber operations experts who applied their learnings to bring national security-grade technology solutions to commercial customers around the world, Blackpoint Cyber is in hyper-growth mode, fueled by a recent $190m series C round.
Job Overview:
We are seeking an experienced Senior Site Reliability Engineer to join our dynamic team. As a Senior SRE Engineer, you will be responsible for designing, implementing, and maintaining our Cloud, On-Premise infrastructure and CI/CD pipelines, with a focus on automation, scalability, and performance.
You will collaborate with cross-functional teams to ensure the reliability and efficiency of our systems while fostering a culture of continuous improvement.
Key Responsibilities:
Infrastructure Management & Automation: Design, develop, and maintain highly scalable infrastructure utilizing Infrastructure as Code (IaC) methodologies, with primary focus on Terraform and Terragrunt for automated cloud resource provisioning and orchestration.
Cloud Platform Administration: Oversee and optimize cloud environments, with a specialized focus on Amazon Web Services (AWS), ensuring adherence to cost optimization strategies, security best practices, and high-availability standards.
Container Orchestration & Continuous Delivery: Manage and optimize Kubernetes cluster environments utilizing Helm, ArgoCD, Istio, and Kustomize to support continuous delivery pipelines and infrastructure-as-code practices.
Data Streaming Platform Operations: Administer and scale data streaming infrastructure using Confluent Cloud and Apache Kafka to support enterprise-level data processing requirements.
Caching & Real-Time Data Solutions: Deploy, configure, and maintain Redis instances to facilitate caching mechanisms and real-time data processing capabilities.
Observability & Incident Management: Implement and maintain comprehensive monitoring, alerting, and incident response frameworks utilizing Prometheus, Grafana, Alert Manager, and OpsGenie/PagerDuty to ensure optimal system reliability and performance.
Feature Management & Release Engineering: Facilitate controlled feature deployments and progressive rollouts through LaunchDarkly/PostHog platform integration and management.
Cross-Functional Collaboration: Partner with software development teams to ensure seamless integration of new services, applications, and features into existing infrastructure ecosystems.
Technical Issue Resolution: Diagnose and resolve complex system-level issues, implementing solutions that maintain high performance standards and maximize system uptime.
Process Optimization & Enhancement: Drive continuous improvement initiatives for automation tools, operational processes, and engineering methodologies to enhance system scalability, reliability, and maintainability.
Technical Innovation & Knowledge Management: Maintain current knowledge of emerging Site Reliability Engineering trends, tools, and technologies, ensuring organizational adoption of relevant industry advancements and best practices.
Skills & Qualifications:
Professional Experience: Minimum of eight (8) years of demonstrated experience in a Senior Site Reliability Engineer role or equivalent position, with substantial emphasis on cloud infrastructure management and automation technologies.
Infrastructure as Code Proficiency: Expertise in Infrastructure as Code (IaC) methodologies, specifically utilizing Terraform and Terragrunt for enterprise-scale deployments.
Cloud Architecture Expertise: Comprehensive knowledge of Amazon Web Services (AWS) cloud platform, including demonstrated proficiency in designing, implementing, and maintaining secure, scalable, and resilient cloud architectures aligned with industry best practices.
Distributed Streaming Systems: Extensive hands-on experience architecting and managing distributed data streaming solutions utilizing Confluent Cloud and Apache Kafka platforms.
Data Storage & Caching Technologies: Proven experience implementing and managing Redis for high-performance caching sol
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s