Staff Software Development Engineer - Enterprise AI Infrastructure - #4898
GRAILAbout the role
The Staff Software Development Engineer - Enterprise AI Infrastructure is a senior technical role responsible for leading the design, development, and scaling of a centralized, highly governed enterprise AI platform. This position serves as a technical expert focused on AWS and Kubernetes-based (EKS) AI infrastructure, agentic development, and multi-agent orchestration operating within a regulated environment. The Staff Engineer partners closely with cross-functional stakeholders across Software Engineering, Data Science, Security, Regulatory Affairs, and Product to build a unified control plane that securely connects large language models with enterprise tools and company knowledge.
This role is expected to drive technical excellence in cloud infrastructure, container orchestration, AI governance, identity-scoped integrations, and agentic workflows while mentoring engineering teams and advancing the organization's enterprise AI strategy.
This role is based in Menlo Park, California, and will move to Sunnyvale, California in Fall 2026 OR Durham NC. It offers a flexible work arrangement, with the ability to work from GRAIL's office or from home. Our current flexible work arrangement policy requires that a minimum of 60%, or 24 hours, of your total work week be on-site. Your specific schedule, determined in collaboration with your manager, will align with team and business needs and could exceed the 60% requirement for the site. At our Menlo Park campus, Tuesdays and Thursdays are the key days where we encourage on-site presence to engage in events and on-site activities.
Responsibilities
-
Lead the end-to-end design, development, deployment, and monitoring of a scalable, governed enterprise AI platform leveraging Amazon EKS and AWS native services (e.g., Bedrock, OpenSearch Serverless, KMS, VPC).
-
Design and implement agentic AI workflows, specialized autonomous agents, and multi-agent systems using advanced LLM orchestration techniques and agent frameworks.
-
Architect and manage secure integrations using the Model Context Protocol (MCP) to connect the AI platform with internal systems, vector databases, and third-party SaaS applications (e.g., Google Workspace, Slack).
-
Build and enforce strict identity, authorization, and zero-trust token brokering flows leveraging Okta, Auth0, and custom JWT authorizers to ensure secure, least-privilege tool execution.
-
Implement deterministic policy controls (e.g., Cedar policy engine) to enforce role-based access, approval gates, and human-in-the-loop checks at the API gateway level.
-
Develop and maintain highly isolated, scalable containerized runtime environments (e.g., Kubernetes pods on Amazon EKS) for secure AI model execution, tool usage, and knowledge retrieval.
-
Establish and maintain comprehensive audit trails and observability for all AI interactions, utilizing AWS CloudTrail and GenAI observability tools (e.g., OpenTelemetry) to track cost, latency, and tool calls.
-
Collaborate with Product Management, Security, Regulatory, and business stakeholders to translate enterprise requirements into scalable, compliant AI infrastructure solutions.
-
Troubleshoot and resolve complex technical issues involving cloud infrastructure, Kubernetes networking, network isolation (PrivateLink), and agentic workflows.
-
Contribute to technology roadmaps, AI infrastructure strategy, and long-term platform evolution initiatives.
-
Mentor engineers, software developers, and technical teams while promoting engineering excellence, infrastructure-as-code (IaC) best practices, and continuous improvement.
-
Partner with Quality, Regulatory, Privacy, Security, and Compliance func
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s