Manager, Core Infrastructure Engineering
OracleAbout the role
For a single team delivering components of distributed systems. Translates goals into a 1–2 quarter execution plan, sets coding, testing, and scalability practices, and provides hands-on oversight of performance tuning and load/perf testing. Guides the team in building fault-tolerant, in-service-upgradable components (redundancy, replication, failover) and in applying resiliency patterns (retries, circuit breakers, timeouts). Ensures robust observability (tests, alarms, dashboards, telemetry) and operational readiness via reviewed runbooks and standard procedures. Manages delivery of scoped features and correctness testing (including fault-injection/brownouts), and directs implementation of data replication/synchronization to maintain integrity and availability. Leads team incident response and root-cause efforts, enforces no-customer-downtime practices, and drives use of automation/IaC for troubleshooting. Oversees team security implementation (encryption, access controls), tracks remediation plans, verifies compliance documentation, and coaches adherence to change-management plans for safe patching, updates, and rollbacks.
Key Responsibilities
System Design & Architecture – System Scalability:
- Provides oversight to the team on the development of components of distributed systems, including the use of distributed state management tools.
- Monitors and enables optimization of code and/or systems for large-scale data processing.
- Guides team to implement scalability requirements for assigned components.
- Coaches team to leverage data plane platforms to effectively handle large-scale data retrieval, storage, and processing.
- Ensures team accurately implements and executes performance and load testing.
System Design & Architecture – System Reliability Design:
- Provides guidance to the team in building fault-tolerant systems capable of withstanding in-service updates by leading the implementation of redundancy, replication, and automatic failover mechanisms.
- Guides the design of components to effectively handle service disruptions.
- Coaches team on various approaches to handle network unreliability, including retry mechanisms, circuit breakers, and timeouts.
System Design & Architecture – System Reliability Performance:
- Guides team to implement testing and alarming configurations for detecting and addressing issues/failures.
- Ensures team effectively supports recovery efforts by reviewing and aligning on runbooks and operational procedures.
- Coaches team on building and customizing dashboards, telemetry systems, and alerting mechanisms to monitor component health.
System Design & Architecture – Correctness / Availability:
- Manages the implementation and design of functional requirements and testing for features within an existing system.
- Coaches team on implementing test scenarios (e.g., fault-injection, brown-out) to evaluate system correctness.
- Guides the implementation of data replication and synchronization techniques within the team to maintain data integrity and availability.
Operational Troubleshooting & Incident Management:
- Leads team efforts in diagnosing, debugging, and resolving issues in system components to support ongoing operation.
- Ensures teams are following protocols for preventing interruptions, ensuring no maintenance windows are required for customers and users when resolving issues.
- Manages team in effectively implementing automation scripts and tooling when troubleshooting operational issues.
- Creates schedules and manages operational support rotations.
Compliance & Security:
- Manages team implementation of robust security measures to protect data and applications in multi-tenant environments, overseeing encryption techniques and access controls.
- Manages execution of remediation plans to address identified security gaps, ensuring continuous improvement of security measures.
- Reviews documentation and ensures cloud infrastructure compliance with industry standards and regulations.
Automation & Change Management:
- Leads the maintenance of automation scripts and tools (e.g., Infrastructure as Code (IaC)) to manage cloud infrastructure.
- Coaches team on change management plans for patching, updating, and rolling back applications.
Core Responsibilities
Planning & Execution:
- Creates and owns the execution plan for the team’s work and multiple projects or initiatives, monitoring timelines and budgets (when applicable) to ensure projects are completed on time and in adherence with requirements.
Apply for this role
Generate a tailored application kit with a matched cover letter, interview prep, and CV highlights — in under 60 seconds.
Apply Now →Generate Application KitFree account required — sign up in 30s