Fullstack Engineer, Observability & SRE - (Remote)
Archesys IncAbout the role
Archesys is a technology firm specializing in innovative cloud solutions and services for clients across various industries. We pride ourselves on our cutting-edge technologies, exceptional customer service, and collaborative work environment.
We're looking for a highly motivated and skilled Fullstack Engineer to join our team, focusing on Observability and Site Reliability Engineering (SRE). In this critical role, you'll be at the forefront of designing, developing, deploying, and ensuring the operational excellence of our Grafana dashboards and the vital data pipelines that feed them. You'll bridge the gap between intuitive front-end visualizations and robust back-end data engineering, ensuring our monitoring solutions are scalable, resilient, and provide actionable insights for our engineering and operations teams. You'll also play a key role in building and automating our DevOps pipelines, all while leveraging the power of AWS.
This position demands a blend of strong software engineering prowess, a deep understanding of SRE principles, expertise in leveraging Grafana to its fullest potential, and significant experience with AWS cloud services and DevOps build automation. You'll be instrumental in enhancing our system visibility, enabling proactive issue detection, and driving continuous improvement in our service reliability.
This is a fully remote, full-time position.
Key Responsibilities:
Grafana Dashboard Development:
Ā Design, develop, and maintain comprehensive, intuitive, and real-time Grafana dashboards that visualize key operational metrics, business KPIs, and application logs.
- Collaborate with SRE, development, and product teams to gather requirements and translate complex data into clear, actionable visualizations.
- Optimize Grafana dashboards for performance, scalability, and usability, ensuring quick loading times and effective data presentation.
- Implement alerting rules within Grafana to proactively notify teams of anomalies and potential issues.
Data Pipeline Engineering (Backend Focus):
- Design and implement robust ETL/ELT pipelines to extract, transform, and load data from various sources (e.g., Prometheus, Splunk, CloudWatch, RDS, OpenTelemetry, custom APIs) into data stores consumable by Grafana.
- Write and optimize complex queries (SQL, PromQL, Splunk SPL, etc.) to ensure data accuracy and efficiency.
- Develop and maintain APIs to facilitate data exchange and integration between different system components and monitoring tools.
- Implement data quality checks, performance tuning (indexing, partitioning), and backup/restore strategies for data sources.
AWS Infrastructure Management:
- Design, deploy, and manage scalable and resilient AWS infrastructure to support Grafana instances, data sources, and related services.
- Utilize AWS services such as EC2, ECS/EKS, Lambda, S3, RDS, CloudWatch, Kinesis, DynamoDB, and others to build and optimize our observability platform.
- Implement security best practices within the AWS environment, including IAM roles, security groups, and network configurations.
DevOps Build Automation:
- Design, implement, and maintain robust CI/CD pipelines for automating the build, testing, and deployment of Grafana dashboards, underlying data pipelines, and infrastructure as code.
- Utilize tools like AWS CodePipeline, Jenkins, GitLab CI, or similar for continuous integration and continuous deployment.
- Develop and maintain Infrastructure as Code (IaC) using Terraform, CloudFormation, or Ansible for managing all AWS resources.
- Automate operational tasks, monitoring deployments, and testing processes to improve efficiency and reliability.
Site Reliability Engineering (SRE) Practices:
- Apply SRE principles to ensure the reliability, scalability, and performance of our monitoring and observability infrastructure.
- Participate in on-call rotations, responding to alerts and incidents related to dashboard functionality, data accuracy, and performance.
- Conduct root cause analysis (RCA) for incidents and implement corrective actions to prevent recurrence.
- Define and track Service Level Objectives (SLOs) and Service Level Indicators (SLIs) for key services, ensuring dashboards reflect these metrics accurately.
Documentation and Collaboration:
- Work closely with cross-functional teams (development, operations, product) to understand monitoring needs and provide expert guidance on observability best practices.
- Create and maintain comprehensive documentation detailing dashboard designs, data sources, query logic, AWS architecture, and operational procedures.
- Contribute
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights ā in under 60 seconds.
Apply Now āGenerate Application KitFree account required ā sign up in 30s