Oracle Cloud Infrastructure (OCI) delivers mission-critical applications for leading enterprises worldwide. Our cloud offers hyperscale, multi-tenant services deployed across more than 50 regions globally. OCI continues to expand beyond traditional public-cloud boundaries to support dedicated, hybrid, and multicloud solutions, edge computing, and more.
As a Principal Core Infrastructure Engineer, you will lead the design and evolution of foundational distributed systems behind OCI. You will build highly scalable, elastic, and fault-tolerant services for high-volume data retrieval, storage, and processing, and set the technical direction for their reliability, correctness, security, and operational readiness.
This position is office-based and requires onsite presence in Nashville, Tennessee. Relocation assistance may be available in accordance with Oracle's relocation policies.
Internal Responsibilities
What You'll Do
- Lead the design, implementation, and ongoing evolution of core distributed systems and data-plane services at hyperscale.
- Define scalability, elasticity, durability, and availability requirements for owned components and ensure designs meet them.
- Optimize high-throughput data paths for large-scale retrieval, storage, and processing using distributed state, replication, and synchronization patterns.
- Design fault-tolerant systems that support in-service updates through redundancy, automatic failover, and recovery-oriented design.
- Apply sound distributed-systems tradeoffs for network partitions and reliability, including load shedding, throttling, rate limiting, retries, and timeouts.
- Establish service-level objectives, key performance indicators, telemetry, dashboards, and proactive alerting for critical systems.
- Design and lead performance, load, fault-injection, and brownout testing to validate correctness, resilience, and operational readiness.
- Lead production incident diagnosis and recovery, guide root-cause analysis, and mentor engineers in operational excellence.
- Build and improve Infrastructure as Code and operational automation that enable safe patching, updates, rollbacks, and change management.
Apply robust security controls and remediation practices for multi-tenant cloud infrastructure, including encryption, access controls, and compliance readiness.
What You'll Bring
Preferred Qualifications
- Experience with Oracle Cloud, AWS, Azure, Google Cloud, or other large-scale cloud platforms.
- Experience with data-plane platforms, distributed storage, microservices, replication, state management, or high-throughput data processing.
- Experience defining SLOs, building observability systems, and operating services in a 24x7 production environment.
- Experience with Infrastructure as Code, service automation, security controls, and compliance requirements for cloud infrastructure.
External Responsibilities
What You'll Do
- Lead the design, implementation, and ongoing evolution of core distributed systems and data-plane services at hyperscale.
- Define scalability, elasticity, durability, and availability requirements for owned components and ensure designs meet them.
- Optimize high-throughput data paths for large-scale retrieval, storage, and processing using distributed state, replication, and synchronization patterns.
- Design fault-tolerant systems that support in-service updates through redundancy, automatic failover, and recovery-oriented design.
- Apply sound distributed-systems tradeoffs for network partitions and reliability, including load shedding, throttling, rate limiting, retries, and timeouts.
- Establish service-level objectives, key performance indicators, telemetry, dashboards, and proactive alerting for critical systems.
- Design and lead performance, load, fault-injection, and brownout testing to validate correctness, resilience, and operational readiness.
- Lead production incident diagnosis and recovery, guide root-cause analysis, and mentor engineers in operational excellence.
- Build and improve Infrastructure as Code and operational automation that enable safe patching, updates, rollbacks, and change management.
Apply robust security controls and remediation practices for multi-tenant cloud infrastructure, including encryption, access controls, and compliance readiness.
What You'll Bring
Preferred Qualifications
- Experience with Oracle Cloud, AWS, Azure, Google Cloud, or other large-scale cloud platforms.
- Experience with data-plane platforms, distributed storage, microservices, replication, state management, or high-throughput data processing.
- Experience defining SLOs, building observability systems, and operating services in a 24x7 production environment.
- Experience with Infrastructure as Code, service automation, security controls, and compliance requirements for cloud infrastructure.