Oracle Cloud Infrastructure (OCI) Networking is the foundation that enables customers to run mission-critical workloads at global cloud scale.
Within the OCI Networking Data Plane organization, the Fleet Management and workload orchestration team builds the foundational platform for deploying, operating, and scaling OCI Networking services across Oracle's global cloud infrastructure. The platform provides common capabilities for fleet lifecycle management, workload placement, deployment automation, capacity expansion, and infrastructure orchestration, enabling networking services to operate consistently and efficiently at scale.
As a Software Engineering Director, you will lead the engineering organization responsible for this strategic platform. You'll define technical strategy, drive execution, and grow a high-performing engineering organization while partnering with architects and engineering leaders across OCI. You'll help shape a standardized platform that simplifies service operations, accelerates infrastructure expansion, and enables networking teams to innovate faster.
Internal Responsibilities
Responsibilities
As a Software Engineering Director, you will:
- Lead and grow the engineering organization responsible for Fleet Management and workload orchestration within the OCI Networking Data Plane organization.
- Define and execute the engineering strategy and roadmap for the platform that manages the lifecycle of OCI Networking fleets, including workload placement, deployment automation, capacity expansion, and infrastructure orchestration.
- Build standardized platform capabilities that simplify how OCI Networking services are deployed, operated, and scaled across Oracle Cloud Infrastructure.
- Own organizational planning, execution, staffing, and delivery across multiple strategic initiatives while balancing customer commitments, operational excellence, and long-term platform investments.
- Drive engineering effectiveness through modern software engineering practices, automation, AI-enabled engineering, CI/CD, and developer tooling.
- Provide technical leadership by guiding architecture, establishing engineering standards, mentoring engineering managers and technical leaders, and raising the engineering bar.
- Deliver reliable, scalable, and observable platform services with a strong focus on operational excellence, automation, and reliability engineering.
- Partner across OCI Networking and OCI Infrastructure organizations to align roadmaps, influence platform strategy, and deliver foundational cloud capabilities.
Preferred Qualifications
- 12+ years of experience designing, building, and operating large-scale distributed systems, cloud infrastructure, or platform services, including experience leading high-performing software engineering organizations.
- Strong technical background in distributed systems, cloud platforms, infrastructure orchestration, deployment automation, fleet management, workload scheduling, or networking infrastructure.
- Proven experience defining technical strategy and delivering complex, mission-critical platforms through cross-functional collaboration.
- Strong computer science fundamentals, including distributed systems, concurrency, software architecture, and system design.
- Experience building and operating reliable, production-grade services with a strong focus on scalability, automation, operational excellence, and customer experience.
- Demonstrated success improving engineering effectiveness through modern software engineering practices, automation, CI/CD, and developer tooling.
- BS or MS in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
Key Responsibilities
System Design & Architecture – System Scalability:
Implements strategies across multiple teams or groups for the architecture and design of interdependent scalable distributed systems, including the use of distributed state management tools, ensuring organizational and system demands are met.
Spearheads code and/or system optimization initiatives for large-scale data processing and high-throughput requirements across multiple areas, driving improvements that support hyper-scale systems.
Facilitates collaborations to define system scalability requirements, ensuring the defined requirements meet customer expectations.
Oversees the design of interdependent systems to scale with elasticity (e.g., effectively scaling both up and down).
Drives the effective use and implementation of data plane platforms for large-scale data operations.
System Design & Architecture – System Reliability Design:
Provides strategic oversight for the architecture of fault-tolerant interdependent systems capable of withstanding in-service updates by overseeing implementation across teams of redundancy, replication, and automatic failover mechanisms.
Influences and sets direction for designing systems to effectively handle service disruptions (e.g., network partitions) by prioritizing consistency, availability, or partition tolerance.
Leads strategic optimization initiatives for handling network unreliability, including directing the design of load-shedding, throttling, and rate-limiting techniques.
Holds teams accountable for leveraging formal verification techniques to verify system designs and conduct peer reviews across teams.
Drives the design of systems that are durable and adhere to service level objectives (SLOs), developing standards for availability and durability of other computing services across the department.
System Design & Architecture – System Reliability Performance:
Drives strategies for defining key performance indicators (KPIs) and telemetry to identify risks, gaps, or cyclical dependencies in running systems, ensuring alignment with organizational goals.
Directs the creation and customization of complex dashboards, telemetry systems, and alerting mechanisms that proactively monitor and ensure optimal system health across teams.
System Design & Architecture – Correctness / Availability:
Implements strategies to effectively determine if systems are meeting functional and correctness requirements, and encourages teams to identify improvement opportunities.
Provides thought leadership on processes for formally verifying complex features to ensure system design correctness.
Oversees the implementation of data replication and synchronization techniques, ensuring data integrity and availability across the organization.
Operational Troubleshooting & Incident Management:
Provides strategic oversight for diagnosing, debugging, and resolving issues in active systems to support ongoing operation.
Directs strategies within teams to prevent interruptions, ensuring no maintenance windows are required for customers and users when resolving issues.
Drives alignment across teams for operational readiness protocol and standard operating procedures.
Provides expert guidance for complex incident response and root cause investigations.
Compliance & Security:
Provides strategic guidance in architecting robust security measures to protect data and applications in multi-tenant environments, ensuring encryption techniques and access controls are implemented.
Oversees execution of remediation plans to address identified security gaps, promoting significant improvements and continuous advancement of security measures.
Drives documentation efforts and ensures cloud infrastructure compliance with industry standards and regulations.
Automation & Change Management:
Provides strategic guidance across teams on developing and maintaining automation scripts and tools (e.g., Infrastructure as Code (IaC)) to manage cloud infrastructure.
Drives strategic alignment of change management plans for patching, updating, and rolling back applications, and oversees that system designs allow for automation of these processes.
Core Responsibilities
Planning & Execution:
Oversees and guides multiple teams on managing complex projects or initiatives, monitoring timelines, deliverables, and budgets (when applicable) to ensure strategic objectives are met.
Serves as a role model for appropriately delegating work, setting priorities, and ensuring alignment with business needs.
Coaches others on adjusting resources or project timelines in anticipation of business changes.
Collaboration & Partnership:
Role models leading cross-functional collaborative efforts to ensure alignment of expectations and strategic objectives.
Empowers teams to build and maintain partnerships with business leaders, stakeholders, and/or customers to address barriers and contribute to organizational success.
Drives transparency and inclusivity by modeling actively seeking, listening to, and leveraging diverse perspectives.
Problem Solving:
Shares problem-solving strategies across teams, providing oversight on complex operational and/or technical issues, as needed.
Coaches teams on analyzing highly complex data and/or information to identify solutions to ambiguous issues.
Provides direction on identifying root causes to prevent recurrence of issues.
Continuous Learning:
Pursues strategic learning opportunities to maintain expertise and apply best practices at the organizational level.
Creates opportunities for team members and leaders to build their expertise in new areas, coaching them to build innovative skills.
Identifies skill gap trends across the organization and upholds a culture that places significant emphasis on sharing knowledge and pursuing learning opportunities that advance the organization.
Evaluates the efficiency of learning strategies and recommends adjustments as needed.
Continuous Improvement:
Empowers teams to own the development and implementation of ideas that increase the efficiency and effectiveness of processes, protocols, and workflows across the department.
Coaches teams to gain buy-in for ideas and to seek feedback on approaches and methods for continued improvement.
Prioritizes and reviews the roadmap of improvement initiatives to ensure alignment with strategic direction and maximize return on investments.
Performance and Development:
Serves as a role model for driving performance across teams through tailored feedback and coaching in alignment with performance management processes, guidelines, and expectations.
Drives consistency in the application of talent development procedures and socializes performance expectations across the organization.
Ensures that individual development goals are aligned with organizational strategic initiatives.
Collaborates with HR to implement talent strategy through hiring and promotion processes.
External Responsibilities
Responsibilities
As a Software Engineering Director, you will:
- Lead and grow the engineering organization responsible for Fleet Management and workload orchestration within the OCI Networking Data Plane organization.
- Define and execute the engineering strategy and roadmap for the platform that manages the lifecycle of OCI Networking fleets, including workload placement, deployment automation, capacity expansion, and infrastructure orchestration.
- Build standardized platform capabilities that simplify how OCI Networking services are deployed, operated, and scaled across Oracle Cloud Infrastructure.
- Own organizational planning, execution, staffing, and delivery across multiple strategic initiatives while balancing customer commitments, operational excellence, and long-term platform investments.
- Drive engineering effectiveness through modern software engineering practices, automation, AI-enabled engineering, CI/CD, and developer tooling.
- Provide technical leadership by guiding architecture, establishing engineering standards, mentoring engineering managers and technical leaders, and raising the engineering bar.
- Deliver reliable, scalable, and observable platform services with a strong focus on operational excellence, automation, and reliability engineering.
- Partner across OCI Networking and OCI Infrastructure organizations to align roadmaps, influence platform strategy, and deliver foundational cloud capabilities.
Preferred Qualifications
- 12+ years of experience designing, building, and operating large-scale distributed systems, cloud infrastructure, or platform services, including experience leading high-performing software engineering organizations.
- Strong technical background in distributed systems, cloud platforms, infrastructure orchestration, deployment automation, fleet management, workload scheduling, or networking infrastructure.
- Proven experience defining technical strategy and delivering complex, mission-critical platforms through cross-functional collaboration.
- Strong computer science fundamentals, including distributed systems, concurrency, software architecture, and system design.
- Experience building and operating reliable, production-grade services with a strong focus on scalability, automation, operational excellence, and customer experience.
- Demonstrated success improving engineering effectiveness through modern software engineering practices, automation, CI/CD, and developer tooling.
- BS or MS in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
Key Responsibilities
System Design & Architecture – System Scalability:
Implements strategies across multiple teams or groups for the architecture and design of interdependent scalable distributed systems, including the use of distributed state management tools, ensuring organizational and system demands are met.
Spearheads code and/or system optimization initiatives for large-scale data processing and high-throughput requirements across multiple areas, driving improvements that support hyper-scale systems.
Facilitates collaborations to define system scalability requirements, ensuring the defined requirements meet customer expectations.
Oversees the design of interdependent systems to scale with elasticity (e.g., effectively scaling both up and down).
Drives the effective use and implementation of data plane platforms for large-scale data operations.
System Design & Architecture – System Reliability Design:
Provides strategic oversight for the architecture of fault-tolerant interdependent systems capable of withstanding in-service updates by overseeing implementation across teams of redundancy, replication, and automatic failover mechanisms.
Influences and sets direction for designing systems to effectively handle service disruptions (e.g., network partitions) by prioritizing consistency, availability, or partition tolerance.
Leads strategic optimization initiatives for handling network unreliability, including directing the design of load-shedding, throttling, and rate-limiting techniques.
Holds teams accountable for leveraging formal verification techniques to verify system designs and conduct peer reviews across teams.
Drives the design of systems that are durable and adhere to service level objectives (SLOs), developing standards for availability and durability of other computing services across the department.
System Design & Architecture – System Reliability Performance:
Drives strategies for defining key performance indicators (KPIs) and telemetry to identify risks, gaps, or cyclical dependencies in running systems, ensuring alignment with organizational goals.
Directs the creation and customization of complex dashboards, telemetry systems, and alerting mechanisms that proactively monitor and ensure optimal system health across teams.
System Design & Architecture – Correctness / Availability:
Implements strategies to effectively determine if systems are meeting functional and correctness requirements, and encourages teams to identify improvement opportunities.
Provides thought leadership on processes for formally verifying complex features to ensure system design correctness.
Oversees the implementation of data replication and synchronization techniques, ensuring data integrity and availability across the organization.
Operational Troubleshooting & Incident Management:
Provides strategic oversight for diagnosing, debugging, and resolving issues in active systems to support ongoing operation.
Directs strategies within teams to prevent interruptions, ensuring no maintenance windows are required for customers and users when resolving issues.
Drives alignment across teams for operational readiness protocol and standard operating procedures.
Provides expert guidance for complex incident response and root cause investigations.
Compliance & Security:
Provides strategic guidance in architecting robust security measures to protect data and applications in multi-tenant environments, ensuring encryption techniques and access controls are implemented.
Oversees execution of remediation plans to address identified security gaps, promoting significant improvements and continuous advancement of security measures.
Drives documentation efforts and ensures cloud infrastructure compliance with industry standards and regulations.
Automation & Change Management:
Provides strategic guidance across teams on developing and maintaining automation scripts and tools (e.g., Infrastructure as Code (IaC)) to manage cloud infrastructure.
Drives strategic alignment of change management plans for patching, updating, and rolling back applications, and oversees that system designs allow for automation of these processes.
Core Responsibilities
Planning & Execution:
Oversees and guides multiple teams on managing complex projects or initiatives, monitoring timelines, deliverables, and budgets (when applicable) to ensure strategic objectives are met.
Serves as a role model for appropriately delegating work, setting priorities, and ensuring alignment with business needs.
Coaches others on adjusting resources or project timelines in anticipation of business changes.
Collaboration & Partnership:
Role models leading cross-functional collaborative efforts to ensure alignment of expectations and strategic objectives.
Empowers teams to build and maintain partnerships with business leaders, stakeholders, and/or customers to address barriers and contribute to organizational success.
Drives transparency and inclusivity by modeling actively seeking, listening to, and leveraging diverse perspectives.
Problem Solving:
Shares problem-solving strategies across teams, providing oversight on complex operational and/or technical issues, as needed.
Coaches teams on analyzing highly complex data and/or information to identify solutions to ambiguous issues.
Provides direction on identifying root causes to prevent recurrence of issues.
Continuous Learning:
Pursues strategic learning opportunities to maintain expertise and apply best practices at the organizational level.
Creates opportunities for team members and leaders to build their expertise in new areas, coaching them to build innovative skills.
Identifies skill gap trends across the organization and upholds a culture that places significant emphasis on sharing knowledge and pursuing learning opportunities that advance the organization.
Evaluates the efficiency of learning strategies and recommends adjustments as needed.
Continuous Improvement:
Empowers teams to own the development and implementation of ideas that increase the efficiency and effectiveness of processes, protocols, and workflows across the department.
Coaches teams to gain buy-in for ideas and to seek feedback on approaches and methods for continued improvement.
Prioritizes and reviews the roadmap of improvement initiatives to ensure alignment with strategic direction and maximize return on investments.
Performance and Development: