As Oracle Cloud Infrastructure (OCI) continues its rapid expansion, we are seeking a skilled Software Engineer to join our newly established Cloud Performance Organization. This team plays a key role in addressing service inefficiencies, reducing cloud expenses, improving customer experience, and ensuring scalability. Your work will focus on optimizing the performance of OCI’s critical components, internal tools, and applications while fostering a culture of performance engineering.
This is a greenfield opportunity to design and build new cloud services from the ground up. We are growing fast, still at an early stage, and working on ambitious new initiatives. You will be part of a team of smart, motivated, diverse people, and given the autonomy as well as support to do your best work. It is a dynamic and flexible workplace where you’ll belong and be encouraged.
Leads development and architecture of scalable, elastic distributed systems for high-throughput, hyperscale workloads. Designs fault-tolerant, highly available systems with robust observability, testing, replication, and resilience mechanisms to meet SLOs. Drives operational readiness, production troubleshooting, and peer mentorship while implementing security, compliance, IaC, and automation for safe patching, upgrades, and rollbacks.
Internal Responsibilities
Key Responsibilities
System Design & Architecture
- Lead development and architecture of scalable, elastic distributed systems for high-throughput, hyperscale workloads.
- Define scalability requirements, optimize performance, and leverage distributed state management and data-plane platforms.
- Design fault-tolerant, highly available systems using redundancy, replication, failover, load shedding, throttling, and rate limiting.
- Establish SLOs, KPIs, telemetry, dashboards, and alerts to ensure reliability and performance.
- Design performance, load, fault-injection, and brownout testing, and implement replication and synchronization for correctness and availability.
Operational Excellence
- Proactively diagnose production issues, guide incident response and root cause analysis, and ensure operational readiness.
- Enable in-service maintenance and upgrades with minimal customer impact.
- Mentor engineers in troubleshooting and operational practices.
Security & Compliance
- Implement encryption, access controls, and security remediation for multi-tenant environments.
- Ensure compliance with applicable standards and maintain required documentation.
Automation & Change Management
- Develop and maintain IaC and automation for cloud infrastructure.
- Enable safe and repeatable patching, updates, and rollbacks through effective change-management practices.
Core Responsibilities
Planning & Execution
- Manage moderately complex initiatives, prioritizing work, timelines, resources, and deliverables while providing technical oversight.
Collaboration & Partnership
- Collaborate across teams and stakeholders to align objectives and deliver solutions that meet business and customer needs.
- Promote inclusive collaboration and diverse perspectives.
Problem Solving
- Analyze and resolve moderately complex issues, escalating critical concerns with clear assessments and recommended solutions.
- Document and share effective problem-solving practices.
Continuous Learning
- Stay current with industry trends and continuously develop technical skills.
- Coach and mentor junior engineers and promote knowledge sharing.
Continuous Improvement
- Identify and implement improvements to processes, workflows, and team effectiveness.
- Evaluate outcomes and incorporate stakeholder feedback.
Performance & Development
- Support talent development through candidate interviews, assessments, and hiring recommendations.
External Responsibilities
Key Responsibilities
System Design & Architecture
- Lead development and architecture of scalable, elastic distributed systems for high-throughput, hyperscale workloads.
- Define scalability requirements, optimize performance, and leverage distributed state management and data-plane platforms.
- Design fault-tolerant, highly available systems using redundancy, replication, failover, load shedding, throttling, and rate limiting.
- Establish SLOs, KPIs, telemetry, dashboards, and alerts to ensure reliability and performance.
- Design performance, load, fault-injection, and brownout testing, and implement replication and synchronization for correctness and availability.
Operational Excellence
- Proactively diagnose production issues, guide incident response and root cause analysis, and ensure operational readiness.
- Enable in-service maintenance and upgrades with minimal customer impact.
- Mentor engineers in troubleshooting and operational practices.
Security & Compliance
- Implement encryption, access controls, and security remediation for multi-tenant environments.
- Ensure compliance with applicable standards and maintain required documentation.
Automation & Change Management
- Develop and maintain IaC and automation for cloud infrastructure.
- Enable safe and repeatable patching, updates, and rollbacks through effective change-management practices.
Core Responsibilities
Planning & Execution
- Manage moderately complex initiatives, prioritizing work, timelines, resources, and deliverables while providing technical oversight.
Collaboration & Partnership
- Collaborate across teams and stakeholders to align objectives and deliver solutions that meet business and customer needs.
- Promote inclusive collaboration and diverse perspectives.
Problem Solving
- Analyze and resolve moderately complex issues, escalating critical concerns with clear assessments and recommended solutions.
- Document and share effective problem-solving practices.
Continuous Learning
- Stay current with industry trends and continuously develop technical skills.
- Coach and mentor junior engineers and promote knowledge sharing.
Continuous Improvement
- Identify and implement improvements to processes, workflows, and team effectiveness.
- Evaluate outcomes and incorporate stakeholder feedback.
Performance & Development
- Support talent development through candidate interviews, assessments, and hiring recommendations.