This is an on-site opportunity in Nashville, TN.
Join and manage the team building Oracle Cloud Infrastructure's state of the art observability platform, powering visibility and operational intelligence for both OCI's internal cloud services and customers running mission-critical workloads on OCI.
OCI Observability serves as a foundational platform used by OCI engineering teams to operate and troubleshoot hundreds of cloud services while also enabling customers to monitor, analyze, and gain insights into their own applications and infrastructure.
This unique position offers the opportunity to build observability solutions that operate at massive scale, serving the demanding needs of OCI's own services as well as a global customer base.
Our team tackles some of the industry's most challenging distributed systems problems, including high-throughput telemetry ingestion, large-scale data processing, cost-efficient storage, low-latency query execution, multi-tenant reliability, and operational excellence.
If you are passionate about leading a team that builds cloud-native observability platforms that power both the cloud itself and the customers who depend on it, we'd love to talk to you.
Internal Responsibilities
- Own the design, development, and operation of a high-scale, distributed telemetry platform that processes billions of datapoints and petabytes of time-series data.
- Ensure the reliability, availability, and operational excellence of services responsible for Observability.
- Provide technical leadership, direction, and strategic vision for a team of senior and principal engineers, fostering a culture of innovation, accountability, and continuous improvement.
- Define and execute a clear, prioritized roadmap of features, platform investments, and operational improvements delivering on commitments on time and with high quality.
- Collaborate cross-functionally with Product Management, other OCI service teams, and Oracle-wide stakeholders to align goals, manage dependencies, and drive integrated solutions.
- Drive and mature engineering processes, including design reviews, operational readiness reviews, quality standards, and incident postmortems.
- Represent the team in executive-level updates and strategic planning discussions, articulating technical direction, risks, and delivery status.
- Proactively monitor the health and performance of services in the global OCI fleet, identifying trends, mitigating risks, and ensuring fault-tolerant, scalable telemetry infrastructure.
- Perform periodic talent reviews and goal setting for your team.
- Responsible for building and managing a cloud service team including recruiting, hiring, and performance management.
Minimum Qualifications
- BS or MS in Computer Science or equivalent experience
- 10+ years of IT; 5+ years of people-management experience managing engineers in software development environment.
- 7+ years’ experience delivering and operating large scale, highly available distributed systems, preferably as a service owner
- Strong knowledge of Java, GoLang, C++ or C and experience with scripting languages such as Python.
- Proficient in data structures, algorithms, operating systems, and distributed systems fundamentals.
- Working familiarity with cloud infrastructure.
- Strong understanding of ingestion, storage and distributed persistence technologies.
- Experience building scalable, cloud native distributed systems.
Preferred Qualifications
- Experience developing services on a public cloud platform (e.g. AWS, Azure, GCP or OCI)
- Experience building multi-tenant, high throughput, low latency cloud scale systems.
External Responsibilities
- Own the design, development, and operation of a high-scale, distributed telemetry platform that processes billions of datapoints and petabytes of time-series data.
- Ensure the reliability, availability, and operational excellence of services responsible for Observability.
- Provide technical leadership, direction, and strategic vision for a team of senior and principal engineers, fostering a culture of innovation, accountability, and continuous improvement.
- Define and execute a clear, prioritized roadmap of features, platform investments, and operational improvements delivering on commitments on time and with high quality.
- Collaborate cross-functionally with Product Management, other OCI service teams, and Oracle-wide stakeholders to align goals, manage dependencies, and drive integrated solutions.
- Drive and mature engineering processes, including design reviews, operational readiness reviews, quality standards, and incident postmortems.
- Represent the team in executive-level updates and strategic planning discussions, articulating technical direction, risks, and delivery status.
- Proactively monitor the health and performance of services in the global OCI fleet, identifying trends, mitigating risks, and ensuring fault-tolerant, scalable telemetry infrastructure.
- Perform periodic talent reviews and goal setting for your team.
- Responsible for building and managing a cloud service team including recruiting, hiring, and performance management.
Minimum Qualifications
- BS or MS in Computer Science or equivalent experience
- 10+ years of IT; 5+ years of people-management experience managing engineers in software development environment.
- 7+ years’ experience delivering and operating large scale, highly available distributed systems, preferably as a service owner
- Strong knowledge of Java, GoLang, C++ or C and experience with scripting languages such as Python.
- Proficient in data structures, algorithms, operating systems, and distributed systems fundamentals.
- Working familiarity with cloud infrastructure.
- Strong understanding of ingestion, storage and distributed persistence technologies.
- Experience building scalable, cloud native distributed systems.
Preferred Qualifications
- Experience developing services on a public cloud platform (e.g. AWS, Azure, GCP or OCI)
- Experience building multi-tenant, high throughput, low latency cloud scale systems.