Principal Cloud Engineer
HarnessAbout the role
Position Summary
As a Cloud Engineer (or Site Reliability Engineer), you'll be responsible for designing, building, debugging, tuning, and supporting existing and new pieces of infrastructure that are core to Split. We expect you to have strong problem-solving skills, be independent and a self-starter, be a great teammate, and have an insatiable appetite to learn more. A Cloud Engineer plays a key role in our software development process, ensuring Harness has a strong foundation to run on and that our customers enjoy a fast and reliable experience when using Harness, no matter where they are in the world.About the role
-
Co-own our overall infrastructure with the rest of our infrastructure team.
-
Design and implement the infrastructure needed to support our product teams.
-
Define and document infrastructure and operational patterns.
-
Scale our team and infrastructure through automation and proactive improvements.
-
Contribute to meeting our security standards and compliance requirements.
-
Measure and improve reliability, availability, and performance metrics across systems.
-
Advocate for and implement best practices in reliability and observability.
-
Participate in on-call rotations to respond to incidents and drive root cause analysis.
About you
Core Skills:
-
Experience with cloud platforms: Strong experience with AWS, GCP, or Azure.
-
Infrastructure as Code (IaC): Experience with tools like Terraform, CloudFormation, or Opentofu.
-
Configuration management
-
Containerization and orchestration: Proficiency with Docker, Kubernetes, and managed solutions like EKS or GKE.
-
Scripting and automation: Proficiency in Bash, Python, or similar languages for automating tasks.
Reliability Engineering Expertise:
-
Observability: Experience with monitoring and logging tools (e.g., Prometheus, Grafana, or DataDog).
-
Incident response: Familiarity with incident management tools like PagerDuty or OpsGenie, and experience in post-mortem/root cause analysis.
-
Performance tuning: Knowledge of analyzing and optimizing system performance (e.g., CPU, memory, storage, and network).
Networking and Web Services:
-
Networking knowledge: Excellent understanding of Internet technologies and protocols (TCP/IP, DNS, HTTP, SSL, etc.).
-
CDNs and DNS: Familiarity with tools like Fastly, Cloudflare, Akamai, or similar.
Cultural Fit and Collaboration:
-
Passion for automating repetitive tasks to improve efficiency.
-
A proactive approach to improving systems and processes.
-
A desire to learn, teach, and collaborate across teams.
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s
Similar roles
Principal Member of Technical Staff – Full Stack Development - Oracle Health
Oracle
$223,400/yr
Senior/ Principal Member of Technical Staff - Backend Developer - Remote
Oracle
$223,400/yr
Principal Software Engineer - Dept 20 - 325
Que Technology Group