Software Engineer, Webservice Platform
lucidlinkAbout the role
The job
We're looking for a TypeScript Engineer with Site Reliability Engineering (SRE) mentality to join our Webservices Platform team. You'll help build and operate a central cloud service that handles customer authentication and authorization, storage, and filespace provisioning — critical infrastructure that other engineering teams and functions like Legal and Finance, depend on for uptime and SLA commitments.
We follow an SRE model where the engineers who write a service also monitor and operate it in production. You'll split your time between writing backend code in Node.js/TypeScript, and applying reliability engineering practices — observability, incident response, SLI/SLO — to keep that code healthy in production. Over time, you'll take ownership of initiatives across the platform, including its next major architecture generations (Web Service 2x/3x), as we modernize toward high availability and disaster recovery.
What you'll do:
- Write and maintain backend services in Node.js and TypeScript, and manage the cloud infrastructure (AWS) they run on.
- Set up and maintain observability — metrics, logging, distributed tracing, and alerting — using tools such as Cloudwatch, Elastic, Prometheus, InfluxDB, and Grafana.
- Participate in on-call rotations: triage incidents, drive rapid mitigation, lead root-cause analysis, and write postmortems. (You'll ramp into on-call gradually — first learning the service, its alerts and runbooks)
- Define and track SLIs, SLOs, and SLAs for the services you own, and use them to prioritize reliability work.
- Reduce operational toil by building internal tooling, CI/CD pipelines, and infrastructure-as-code.
- Contribute to modernization efforts aimed at high availability, disaster recovery, and high uptime, and help remove architectural obstacles blocking that work.
- Support provisioning and platform work for new customers alongside reliability and feature work.
- Collaborate closely with the rest of the Web Service Platform team, Platform Engineering, and other engineering teams whose services depend on this platform.
Key responsibilities:
- Design and develop reliable, scalable TypeScript/Node.js services that power authentication, storage, filesystem/filespace provisioning, billing and other services for our customers.
- Own what you build in production — set up monitoring, logging, and alerting, and take part in on-call rotations.
- Challenge technical decisions to drive the platform toward high availability, disaster recovery, and a high uptime SLA.
- Define and track SLIs/SLOs, and reduce operational toil by building automation, CI/CD pipelines, and infrastructure-as-code.
- Triage production incidents, drive rapid mitigation, lead root-cause analysis, and write postmortems.
- Reproduce and investigate complex production scenarios — performance issues, failure modes, edge cases — not just surface-level bugs.
Job requirements:
- 3-5+ years of backend software development experience, ideally with TypeScript/Node.js.
- Strong backend architecture skills, with experience designing for scale and fault tolerance.
- Deep technical understanding of Linux/systems fundamentals
- SRE experience or strong SRE practices — observability, incident management, SLO-driven work.
- Experience operating and monitoring cloud services in production (AWS or similar); familiarity with tools like Cloudwatch, Elastic, Prometheus, InfluxDB, and Grafana.
- Experience with RDBMS and NoSQL systems like MariaDB, Mongo, Redis or others.
- Experience with message broker software, like RabbitMQ, Redis, Amazon SQS.
- Comfortable working independently, communicating blockers openly, and collaborating across teams.
Interview process:
Here’s what you can expect when you apply:
1. Initial interview – a chat with our People & Culture team to get to know you and your background.
2. Technical interview – a conversation with some of our engineers to dive into your skills and experience.
3. Take-home task – an assignment to work on in your own time, showcasing your problem-solving approach.
4. Task review – a follow-up interview where you’ll walk us through your solution and thought process.
5. Final interview – a discussion with our VP of Engineering, Engineering Manager, or both, to make sure we’re the right fit for each other.
REASONS TO JOIN LUCIDLINK:
- Tackle big challenges: You’ll have the chance to solve complex, high-stakes problems that redefine how teams collaborate globally. By starting with the Media & Entertainment industry and expanding into data-intensive sectors, you’ll gain deep insight into cutting-edge technologies and play a role in shaping the future of global workflows.
- Values-led culture: Our values don’t just exist on paper—they guide every decision and interaction. You’ll thrive in an environment where integrity, innovation, and empathy are at the core of how we operate, empowering you to grow personally and profes
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s