Oracle Cloud Infrastructure (OCI) is building the next generation of hyperscale cloud and AI infrastructure that powers some of the world's most demanding workloads, including large-scale GPU Superclusters, AI training platforms, and commercial cloud services. Our mission is to deliver highly available, scalable, automated, and observable network infrastructure that enables customers to innovate with confidence.
As Oracle continues to expand its AI and cloud footprint, we are investing heavily in Network Automation and Observability platforms that simplify operations, improve reliability, accelerate service delivery, and provide deep operational insights across one of the world's largest cloud networks.
We are looking for a Principal Technical Program Manager (TPM) to lead cross-functional programs focused on Network Automation, Monitoring, and Operational Excellence for OCI GPU Superclusters and Commercial Cloud Infrastructure. This role is ideal for someone who thrives in large-scale distributed systems, enjoys solving complex operational challenges, and has a passion for driving automation, reliability, and engineering excellence.
Working closely with Network Engineering, Cloud Infrastructure, SRE, Observability, Automation Platform, Product Management, and Operations teams, you will lead strategic initiatives that improve service scalability, operational efficiency, deployment velocity, and customer experience across OCI's global infrastructure.
This is an opportunity to influence the evolution of next-generation cloud networking platforms that power AI, HPC, and enterprise cloud services at massive scale.
Internal Responsibilities
- Lead complex cross-organizational programs focused on network automation, infrastructure observability, operational tooling, and platform engineering across OCI Cloud Infrastructure and GPU Superclusters.
- Partner with engineering, architecture, SRE, network operations, capacity planning, and product teams to define strategic roadmaps and execution plans.
- Drive initiatives that improve infrastructure reliability, operational efficiency, deployment automation, service scalability, and operational visibility.
- Define program scope, milestones, dependencies, success metrics, KPIs, and execution plans across multiple engineering organizations.
- Identify technical and program risks early, develop mitigation strategies, and communicate program health effectively to senior leadership.
- Champion engineering best practices around automation, telemetry, monitoring, incident reduction, operational readiness, and service health.
- Drive cross-functional execution for large infrastructure initiatives including network automation frameworks, observability platforms, AI infrastructure deployments, and operational tooling.
- Develop executive dashboards, scorecards, and reporting mechanisms to measure operational excellence, automation adoption, service health, reliability, and engineering productivity.
- Partner with engineering leaders to improve development processes, release management, operational readiness, and continuous improvement initiatives.
- Facilitate technical discussions and help resolve complex cross-team dependencies across geographically distributed organizations.
- Lead Rhythm of Business (RoB), roadmap planning, quarterly execution reviews, and executive program reporting.
- Promote a culture of operational excellence, automation-first thinking, data-driven decision making, and continuous improvement.
Minimum Qualifications
- 5+ years of Technical Program Management experience leading large-scale infrastructure or cloud platform programs.
- Experience supporting engineering organizations building cloud infrastructure, networking, distributed systems, platform services, or AI infrastructure.
- Strong understanding of cloud infrastructure, networking fundamentals, distributed systems, or infrastructure operations.
- Experience managing complex cross-functional technical programs involving multiple engineering teams.
- Experience driving automation initiatives, operational excellence programs, or platform modernization efforts.
- Strong analytical, organizational, and problem-solving skills with the ability to manage competing priorities.
- Excellent written, verbal, and executive-level communication skills.
- Experience using Atlassian tools (Jira, Confluence) and project management methodologies.
- Demonstrated ability to influence engineering leaders without direct authority.
- Ability to translate complex technical concepts into clear business outcomes for executive stakeholders.
Preferred Qualifications
- Experience with hyperscale cloud infrastructure, large-scale networking, AI infrastructure, GPU clusters, HPC environments, or cloud networking platforms.
- Experience leading programs involving infrastructure automation, CI/CD, Infrastructure as Code, or cloud platform engineering.
- Familiarity with network automation technologies, software-defined networking, service reliability engineering (SRE), or DevOps practices.
- Experience managing programs involving service reliability, operational readiness, incident management, or production operations.
- Bachelor's degree in computer science, engineering, Information Systems, or a related technical discipline; advanced degree preferred.
What You'll Bring
- Passion for building highly automated, reliable cloud infrastructure at hyperscale.
- Strong program leadership across highly technical engineering organizations.
- Ability to navigate ambiguity while driving clarity, accountability, and execution.
- Data-driven mindset with a focus on operational excellence and measurable outcomes.
- Experience influencing engineering strategy and organizational change.
- Customer-first mentality with a continuous improvement mindset.
- Ability to thrive in a fast-paced, rapidly evolving cloud engineering environment.
External Responsibilities
- Lead complex cross-organizational programs focused on network automation, infrastructure observability, operational tooling, and platform engineering across OCI Cloud Infrastructure and GPU Superclusters.
- Partner with engineering, architecture, SRE, network operations, capacity planning, and product teams to define strategic roadmaps and execution plans.
- Drive initiatives that improve infrastructure reliability, operational efficiency, deployment automation, service scalability, and operational visibility.
- Define program scope, milestones, dependencies, success metrics, KPIs, and execution plans across multiple engineering organizations.
- Identify technical and program risks early, develop mitigation strategies, and communicate program health effectively to senior leadership.
- Champion engineering best practices around automation, telemetry, monitoring, incident reduction, operational readiness, and service health.
- Drive cross-functional execution for large infrastructure initiatives including network automation frameworks, observability platforms, AI infrastructure deployments, and operational tooling.
- Develop executive dashboards, scorecards, and reporting mechanisms to measure operational excellence, automation adoption, service health, reliability, and engineering productivity.
- Partner with engineering leaders to improve development processes, release management, operational readiness, and continuous improvement initiatives.
- Facilitate technical discussions and help resolve complex cross-team dependencies across geographically distributed organizations.
- Lead Rhythm of Business (RoB), roadmap planning, quarterly execution reviews, and executive program reporting.
- Promote a culture of operational excellence, automation-first thinking, data-driven decision making, and continuous improvement.
Minimum Qualifications
- 5+ years of Technical Program Management experience leading large-scale infrastructure or cloud platform programs.
- Experience supporting engineering organizations building cloud infrastructure, networking, distributed systems, platform services, or AI infrastructure.
- Strong understanding of cloud infrastructure, networking fundamentals, distributed systems, or infrastructure operations.
- Experience managing complex cross-functional technical programs involving multiple engineering teams.
- Experience driving automation initiatives, operational excellence programs, or platform modernization efforts.
- Strong analytical, organizational, and problem-solving skills with the ability to manage competing priorities.
- Excellent written, verbal, and executive-level communication skills.
- Experience using Atlassian tools (Jira, Confluence) and project management methodologies.
- Demonstrated ability to influence engineering leaders without direct authority.
- Ability to translate complex technical concepts into clear business outcomes for executive stakeholders.
Preferred Qualifications
- Experience with hyperscale cloud infrastructure, large-scale networking, AI infrastructure, GPU clusters, HPC environments, or cloud networking platforms.
- Experience leading programs involving infrastructure automation, CI/CD, Infrastructure as Code, or cloud platform engineering.
- Familiarity with network automation technologies, software-defined networking, service reliability engineering (SRE), or DevOps practices.
- Experience managing programs involving service reliability, operational readiness, incident management, or production operations.
- Bachelor's degree in computer science, engineering, Information Systems, or a related technical discipline; advanced degree preferred.
What You'll Bring
- Passion for building highly automated, reliable cloud infrastructure at hyperscale.
- Strong program leadership across highly technical engineering organizations.
- Ability to navigate ambiguity while driving clarity, accountability, and execution.
- Data-driven mindset with a focus on operational excellence and measurable outcomes.
- Experience influencing engineering strategy and organizational change.
- Customer-first mentality with a continuous improvement mindset.
- Ability to thrive in a fast-paced, rapidly evolving cloud engineering environment.