Oracle Cloud Infrastructure’s (OCI) Developer Platform Builder Tools organization builds the next generation of developer productivity platforms, intelligent development workflows, and AI-powered engineering tools that accelerate software delivery across the enterprise.
We are seeking a Principal Platform Software Engineer to define the architecture and technical direction of scalable services that improve the developer experience and increase engineering productivity. You will work across platform engineering, distributed systems, cloud-native technologies, machine learning, large language models (LLMs), and developer tooling.
A key focus of this role is building a new AI-native testing platform that learns real-world service behavior, generates realistic test traffic, safely validates workloads, and detects regressions with minimal manual test authoring. Its closed-loop workflow observes changing traffic patterns, executes and evaluates tests, and continuously improves future coverage.
As a principal engineer, you will shape the platform from foundational architecture through production deployment and broad adoption across OCI. You will solve ambiguous, organization-wide technical problems, establish engineering standards, and provide technical leadership to senior engineers, product managers, data scientists, and OCI service teams.
Internal Responsibilities
Key Responsibilities
Define the technical vision, architecture, and multiyear evolution of OCI’s AI-native developer platforms.
Lead the design and delivery of secure, highly available services that operate reliably across large-scale, distributed cloud environments.
Architect systems that analyze service behavior, model traffic patterns, generate realistic workloads, and evaluate functional and non-functional outcomes.
Establish agent-assisted workflows that plan, execute, and evaluate canary, functional, integration, load, and performance tests within clearly defined safety controls.
Guide the application of machine learning and LLM capabilities to service telemetry, API changes, incidents, test results, and engineering knowledge.
Develop architectural approaches for failure diagnosis, regression detection, workload modeling, and the conversion of production incidents into reusable test coverage.
Define platform APIs, data models, extension points, and integration patterns that enable adoption across diverse OCI services and engineering environments.
Set engineering standards for scalability, reliability, observability, security, privacy, performance, and responsible AI.
Define success metrics and feedback loops that measure test effectiveness, traffic-pattern coverage, developer effort saved, platform reliability, and adoption.
Provide hands-on technical leadership, mentor senior engineers, resolve cross-team architectural challenges, and raise engineering standards across the organization.
Required Qualifications
Bachelor’s or Master’s degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
Eight or more years of professional software engineering experience, including significant experience designing and operating large-scale production systems.
Proficiency in one or more modern programming languages, such as Java, Go, Python, or C++.
Deep expertise in distributed systems, system design, APIs, concurrency, data modeling, and scalable service architectures.
Demonstrated success defining architecture and delivering complex platforms, backend services, microservices, event-driven systems, or large-scale data pipelines.
Experience developing and operating cloud-native software using CI/CD, containerization, orchestration, and infrastructure automation.
Strong understanding of service reliability, observability, performance engineering, security, and production operations.
Experience integrating machine learning, generative AI, or other intelligent capabilities into production software systems.
Proven ability to lead high-impact technical initiatives across multiple teams, influence without direct authority, and navigate ambiguous requirements.
Excellent technical judgment and communication skills, including the ability to explain complex architectural decisions to engineering, product, and executive stakeholders.
A track record of mentoring senior engineers and improving engineering practices across teams or organizations.
Preferred Qualifications
Experience architecting developer platforms, CI/CD systems, testing infrastructure, or engineering productivity tools at enterprise scale.
Experience developing production solutions using machine learning, generative AI, or LLMs.
Familiarity with AI agents, tool-using models, retrieval-augmented generation, prompt engineering, LLM-as-a-judge approaches, or model evaluation.
Experience with machine learning frameworks such as PyTorch, TensorFlow, Hugging Face Transformers, or equivalent technologies.
Experience with traffic replay, workload modeling, canary analysis, load testing, performance testing, or chaos engineering.
Deep experience with Kubernetes, containers, service meshes, and OCI or another major public cloud.
Experience designing extensible, multi-tenant platforms used by multiple engineering organizations.
Experience analyzing API schemas, code changes, service dependencies, production incidents, and traffic patterns.
Familiarity with AI safety controls, model monitoring, privacy, and secure enterprise data handling.
Why Join Us
Define the architecture and technical direction of a new AI-native developer platform with broad impact across OCI.
Solve complex, high-ambiguity problems spanning distributed systems, cloud infrastructure, machine learning, and generative AI.
Build capabilities that help thousands of engineers reduce manual testing, detect regressions earlier, diagnose failures faster, and release with greater confidence.
Influence engineering strategy across teams while remaining deeply involved in technical design and implementation.
Work with experienced engineers, data scientists, product managers, and OCI service teams to establish a new standard for intelligent software testing at enterprise scale.
#LI-SP1
External Responsibilities
Key Responsibilities
Define the technical vision, architecture, and multiyear evolution of OCI’s AI-native developer platforms.
Lead the design and delivery of secure, highly available services that operate reliably across large-scale, distributed cloud environments.
Architect systems that analyze service behavior, model traffic patterns, generate realistic workloads, and evaluate functional and non-functional outcomes.
Establish agent-assisted workflows that plan, execute, and evaluate canary, functional, integration, load, and performance tests within clearly defined safety controls.
Guide the application of machine learning and LLM capabilities to service telemetry, API changes, incidents, test results, and engineering knowledge.
Develop architectural approaches for failure diagnosis, regression detection, workload modeling, and the conversion of production incidents into reusable test coverage.
Define platform APIs, data models, extension points, and integration patterns that enable adoption across diverse OCI services and engineering environments.
Set engineering standards for scalability, reliability, observability, security, privacy, performance, and responsible AI.
Define success metrics and feedback loops that measure test effectiveness, traffic-pattern coverage, developer effort saved, platform reliability, and adoption.
Provide hands-on technical leadership, mentor senior engineers, resolve cross-team architectural challenges, and raise engineering standards across the organization.
Required Qualifications
Bachelor’s or Master’s degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
Eight or more years of professional software engineering experience, including significant experience designing and operating large-scale production systems.
Proficiency in one or more modern programming languages, such as Java, Go, Python, or C++.
Deep expertise in distributed systems, system design, APIs, concurrency, data modeling, and scalable service architectures.
Demonstrated success defining architecture and delivering complex platforms, backend services, microservices, event-driven systems, or large-scale data pipelines.
Experience developing and operating cloud-native software using CI/CD, containerization, orchestration, and infrastructure automation.
Strong understanding of service reliability, observability, performance engineering, security, and production operations.
Experience integrating machine learning, generative AI, or other intelligent capabilities into production software systems.
Proven ability to lead high-impact technical initiatives across multiple teams, influence without direct authority, and navigate ambiguous requirements.
Excellent technical judgment and communication skills, including the ability to explain complex architectural decisions to engineering, product, and executive stakeholders.
A track record of mentoring senior engineers and improving engineering practices across teams or organizations.
Preferred Qualifications
Experience architecting developer platforms, CI/CD systems, testing infrastructure, or engineering productivity tools at enterprise scale.
Experience developing production solutions using machine learning, generative AI, or LLMs.
Familiarity with AI agents, tool-using models, retrieval-augmented generation, prompt engineering, LLM-as-a-judge approaches, or model evaluation.
Experience with machine learning frameworks such as PyTorch, TensorFlow, Hugging Face Transformers, or equivalent technologies.
Experience with traffic replay, workload modeling, canary analysis, load testing, performance testing, or chaos engineering.
Deep experience with Kubernetes, containers, service meshes, and OCI or another major public cloud.
Experience designing extensible, multi-tenant platforms used by multiple engineering organizations.
Experience analyzing API schemas, code changes, service dependencies, production incidents, and traffic patterns.
Familiarity with AI safety controls, model monitoring, privacy, and secure enterprise data handling.
Why Join Us
Define the architecture and technical direction of a new AI-native developer platform with broad impact across OCI.
Solve complex, high-ambiguity problems spanning distributed systems, cloud infrastructure, machine learning, and generative AI.
Build capabilities that help thousands of engineers reduce manual testing, detect regressions earlier, diagnose failures faster, and release with greater confidence.
Influence engineering strategy across teams while remaining deeply involved in technical design and implementation.
Work with experienced engineers, data scientists, product managers, and OCI service teams to establish a new standard for intelligent software testing at enterprise scale.
#LI-SP1