Site Reliability Engineer III, (IoT Observability)
Flock SafetyAbout the role
Who is Flock?
Flock Safety is the leading safety technology platform, helping communities thrive by taking a proactive approach to crime prevention and security. Our hardware and software suite connects cities, law enforcement, businesses, schools, and neighborhoods in a nationwide public-private safety network. Trusted by over 5,000 communities, 4,500 law enforcement agencies, and 1,000 businesses, Flock delivers real-time intelligence while prioritizing privacy and responsible innovation.
We’re a high-performance, low-ego team driven by urgency, collaboration, and bold thinking. Working at Flock means tackling big challenges, moving fast, and continuously improving. It’s intense but deeply rewarding for those who want to make an impact.
With nearly $700M in venture funding and a $7.5B valuation, we’re scaling intentionally and seeking top talent to help build the impossible. If you value teamwork, ownership, and solving tough problems, Flock could be the place for you.
The Opportunity
The Device SRE team owns deployment of software to our fleet of IoT devices, as well as observability during and after deployment. This involves working in the entire stack from telemetry/event generation in the firmware, data ingestion into the cloud, visualization of data, and alarms that indicate issues. We are expanding our scope beyond Cameras to include Aviation devices. Our goal is to provide consistent infrastructure, processes, and tools across the organization that enable the development team to roll out new software rapidly and reliably, while enabling them to identify and respond to issues quickly, resulting in fast rollouts of new features with minimal downtime.
The Skillset
Must Have
Experience developing software for embedded systems, especially large or complex IoT/edge devices
Hands-on experience instrumenting metrics and telemetry directly on-device (e.g., firmware counters, health signals, performance instrumentation, event/event-generation)
Strong coding skills (ex: C/C++, Python, R, JS, Java, Groovy) and understanding of common algorithms
Strongly Valued
Proficiency in scripting languages (e.g., Bash, Python) to automate processes
Experience with data ingestion pipelines (telemetry from device --> cloud, batching, retries, reliability patterns)
Data visualization & dashboarding (ex: Grafana, Sigma)
Broad experience with databases & logging (SQL: PostgreSQL, NoSQL, Time Series like Prometheus / DataDog), even if not deep SQL expertise
Cloud computing experience (e.g., AWS) and distributed systems fundamentals
Site Reliability Engineering experience for IoT devices (on-call, monitoring, & alerting)
Strong grasp of software development workflows (CI/CD, test automation, semantic versioning, branching strategies)
Nice-to-Have
Experience with infrastructure-as-code (IaC) tools (Ansible, Terraform)
Experience with volume data processing (pipelines, modeling, large-scale storage)
Experience with distributed telemetry architectures or large-fleet device management patterns
Feeling uneasy that you haven’t ticked every box? That’s okay; we’ve felt that way too. Studies have shown women and minorities are less likely to apply unless they meet all qualifications. We encourage you to break the status quo and apply to roles that would make you excited to come to work every day.
90 Days at Flock
We prescribe 90 day plans and believe that good days lead to good weeks, which lead to good months. This serves as a preview of the 90 day plan you will receive if you were to be hired as a Senior SRE, Devices Observability at Flock Safety.
The First 30 Days
Maintain and improve data collected from devices for monitoring fleet health.
Improve ability to identify and root-cause issues in the fleet, including automated alarms.
Maintain and build out dashboards that provide insight into fleet health.
Streamline software rollout workflows and monitoring to reduce manual steps and improve response times.
Define and build out on-call rotations for monitoring fleet health.
Perform ad-hoc analysis of unexpected issues by analyzing data collected from the fleet.
Participate in on-call rotations to serve as first responder for outages and general requests from stakeholders.
The First 60 Days
Improve or create a Grafana dashboard.
Improve or create a Sigma dashboard.
I
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s