Sr. Manager/Staff Engineer, AI Infrastructure and Operations
PfizerAbout the role
ROLE SUMMARY
The Senior Manager/Staff Engineer, AI Infrastructure & MLOps Engineering is a senior individual contributor position reporting directly to the Director, AI Infrastructure and Operations Lead. This role is highly technical, hands-on, and focused on building the core automation, tooling, and infrastructure that power the internal AI platform and services.
In this position, the Senior Manager/Staff Engineer is responsible for designing and implementing systems that enable scientists and engineers to rapidly build, deploy, and monitor machine learning models in production. Work will span Python-based automation, containerization with Docker, CI/CD pipelines, AWS cloud infrastructure, microservices, and high-performance model serving frameworks.
The role plays a critical part in advancing the organization’s MLOps capabilities by creating reusable components and internal developer platforms that increase velocity, reliability, and scalability of AI/ML delivery.
ROLE RESPONSIBILITIES
Core Engineering & Automation
Design, build, and maintain Python-based tooling, SDKs, and automation frameworks to support model development, deployment, and monitoring workflows.
Develop containerized solutions using Docker and orchestrate them using Kubernetes (including Kubeflow or similar MLOps platforms).
Build and maintain CI/CD pipelines to streamline ML model integration, testing, and deployment into production environments.
Implement robust automation for provisioning, configuring, and managing cloud resources using Infrastructure-as-Code (Terraform, Pulumi, AWS CDK, etc.).
Cloud Infrastructure & Platform Engineering
Architect and manage scalable, secure, and high-availability AWS infrastructure to support AI workloads.
Develop and optimize microservices architectures for AI/ML serving, ensuring high throughput and low latency.
Build and maintain APIs and services for model management, feature stores, and inference pipelines.
Implement monitoring, logging, and observability tools to ensure performance, availability, and reliability of AI services.
Model Serving & MLOps Enablement
Operationalize ML model serving at scale using frameworks such as TensorFlow Serving, TorchServe, KServe, Seldon Core, or custom inference services.
Create reusable MLOps components for data preprocessing, training orchestration, model validation, and deployment.
Develop automation to reduce ML model deployment time, enforce versioning, and enable rollback/upgrade capabilities.
Work closely with data scientists to translate research workflows into production-grade, scalable services.
Collaboration & Best Practices
Partner with AI researchers, data engineers, and platform engineers to deliver integrated solutions.
Champion engineering excellence by promoting design documentation, code reviews, CI/CD best practices, and testing automation.
Contribute to a culture of shared ownership, transparency, and internal open-source development.
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s