Senior AI Engineer - APM Experiences
DatadogAbout the role
<h2><strong>The opportunity</strong></h2> <p>Datadog’s APM Experiences team owns the core product experience for <a href="https://docs.datadoghq.com/tracing/">Application Performance Monitoring</a> — including distributed tracing, service representation, and more. We’re building a new wave of AI-powered capabilities that help customers <em>detect, resolve, and prevent</em> performance issues faster. In this role, you will lead end‑to‑end development of LLM- and Agent‑based features that can:</p> <ul> <li>Debug and investigate application performance issues down to the root cause, as both a developer assistant and a fully autonomous agent</li> <li>Proactively recommend performance and reliability-based optimizations to prevent the next incident</li> <li>Automatically create intelligent monitors and SLOs for the most important business flows and critical paths</li> </ul> <p>This is a highly product‑minded engineering role: you’ll work from problem discovery and UX all the way to reliable, scalable production systems.</p> <h2><strong>What you’ll do</strong></h2> <ul> <li><strong>Shape AI experiences for APM.</strong> Design and ship LLM/agentic workflows that analyze traces, metrics, logs, and other telemetry to generate diagnoses, explanations, and guided fixes.</li> <li><strong>Own the full loop.</strong> Prototype quickly, define success metrics and evals, run experiments, iterate, and ultimately productionize for scale and reliability.</li> <li><strong>Build robust agent systems.</strong> Develop tools, retrieval and planning strategies, and guardrails; manage prompts/evals; design fallbacks and human‑in‑the‑loop paths.</li> <li><strong>Integrate with Datadog’s platform.</strong> Leverage surfaces like Trace Explorer, Service Catalog, monitors, and workflows to deliver end‑to‑end value in the APM UI.</li> <li><strong>Partner deeply.</strong> Collaborate with PM, Design, and partner teams to build cohesive experiences.</li> <li><strong>Raise the bar on engineering.</strong> Write performant, maintainable backend code, own services in production, and improve reliability for high‑throughput, low‑latency data systems.</li> </ul> <h2><strong>Who you are</strong></h2> <p><strong>Product‑minded engineer who ships AI to production</strong></p> <ul> <li>4+ years building backend or real-time ML systems; you value simplicity, correctness, and performance</li> <li>Proven experience delivering LLM/agent features to production (prompting, tooling, evals, safety/guardrails)</li> <li>Comfortable owning user journeys, iterating from prototype → alpha → GA, and measuring impa
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s