Oracle Cloud Infrastructure’s (OCI) Developer Platform Tools and Services organization builds the next generation of developer productivity platforms, intelligent development workflows, and AI-powered engineering tools and platforms that accelerate software delivery across the enterprise.
We are seeking a Senior Platform Software Engineer to design, build, and operate scalable services that improve the developer experience and increase engineering productivity. You will work across platform engineering, distributed systems, cloud-native technologies, machine learning, large language models (LLMs), and developer tooling.
A key focus of this role is building a new AI-native testing platform that learns real-world service behavior, generates realistic test traffic, safely validates workloads, and detects regressions with minimal manual test authoring. Its closed-loop workflow observes changing traffic patterns, executes and evaluates tests, and continuously improves future coverage. You will help take this platform from early architecture through production deployment and enterprise-wide adoption, working with engineers, product managers, data scientists, and OCI service teams.
Internal Responsibilities
Key Responsibilities
Design, build, deploy, and operate highly available platform services for developer productivity and AI-assisted engineering workflows.
Lead the architecture and delivery of scalable systems that analyze service behavior, model traffic patterns, generate realistic workloads, and evaluate results.
Develop agent-assisted workflows that plan, execute, and evaluate canary, functional, integration, load, and performance tests within defined safety controls.
Apply machine learning and LLM capabilities to service telemetry, API changes, incidents, test results, and engineering knowledge.
Build safe traffic-generation and workload-modeling systems that reflect production usage, service dependencies, and changing traffic patterns.
Develop evaluation and diagnostic capabilities for functional correctness, latency, error rates, reliability, scalability, and regression detection.
Provide actionable explanations of failures, including what changed, likely sources of impact, and how to reproduce an issue.
Automate the conversion of incidents, failed canaries, and detected regressions into reusable test coverage.
Design APIs, data pipelines, and integrations that make intelligent testing capabilities accessible across OCI engineering teams.
Establish strong practices for reliability, observability, security, privacy, performance, and responsible AI while mentoring engineers and influencing technical direction.
Required Qualifications
Bachelor’s or Master’s degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
Four or more years of professional software engineering experience building and operating production systems.
Proficiency in one or more modern programming languages, such as Java, Go, Python, or C++.
Strong knowledge of system design, algorithms, data structures, concurrency, APIs, and scalable architectures.
Experience building backend services, distributed systems, microservices, event-driven applications, or large-scale data pipelines.
Experience developing and operating software in cloud environments using CI/CD, containerization, and cloud-native architectures.
Experience with service testing, reliability, observability, performance engineering, or production operations.
Ability to troubleshoot complex systems using logs, metrics, traces, telemetry, and production data.
Demonstrated ability to lead technical initiatives across teams, communicate architectural decisions clearly, and deliver high-quality production software.
Preferred Qualifications
Experience building developer platforms, CI/CD systems, testing frameworks, or engineering productivity tools.
Experience developing or integrating production solutions using machine learning, generative AI, or LLMs.
Familiarity with AI agents, tool-using models, retrieval-augmented generation, prompt engineering, or LLM evaluation techniques.
Experience with machine learning frameworks such as PyTorch, TensorFlow, Hugging Face Transformers, or equivalent technologies.
Experience with traffic replay, workload modeling, canary analysis, load testing, performance testing, or chaos engineering.
Hands-on experience with Kubernetes, containers, service meshes, and OCI or another major public cloud.
Experience analyzing API schemas, code changes, service dependencies, incidents, and production traffic.
Experience mentoring engineers and influencing architecture across teams.
Why Join Us
Shape a new AI-native developer platform from foundational architecture through production adoption.
Build systems that improve how thousands of engineers develop, test, and release software.
Solve complex, real-world problems spanning distributed systems, cloud infrastructure, machine learning, and generative AI.
Help engineering teams reduce manual testing, detect regressions earlier, diagnose failures faster, and release with greater confidence.
Work with experienced engineers, data scientists, product managers, and OCI service teams while expanding your technical influence and expertise.
#LI-SP1
External Responsibilities
Key Responsibilities
Design, build, deploy, and operate highly available platform services for developer productivity and AI-assisted engineering workflows.
Lead the architecture and delivery of scalable systems that analyze service behavior, model traffic patterns, generate realistic workloads, and evaluate results.
Develop agent-assisted workflows that plan, execute, and evaluate canary, functional, integration, load, and performance tests within defined safety controls.
Apply machine learning and LLM capabilities to service telemetry, API changes, incidents, test results, and engineering knowledge.
Build safe traffic-generation and workload-modeling systems that reflect production usage, service dependencies, and changing traffic patterns.
Develop evaluation and diagnostic capabilities for functional correctness, latency, error rates, reliability, scalability, and regression detection.
Provide actionable explanations of failures, including what changed, likely sources of impact, and how to reproduce an issue.
Automate the conversion of incidents, failed canaries, and detected regressions into reusable test coverage.
Design APIs, data pipelines, and integrations that make intelligent testing capabilities accessible across OCI engineering teams.
Establish strong practices for reliability, observability, security, privacy, performance, and responsible AI while mentoring engineers and influencing technical direction.
Required Qualifications
Bachelor’s or Master’s degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
Four or more years of professional software engineering experience building and operating production systems.
Proficiency in one or more modern programming languages, such as Java, Go, Python, or C++.
Strong knowledge of system design, algorithms, data structures, concurrency, APIs, and scalable architectures.
Experience building backend services, distributed systems, microservices, event-driven applications, or large-scale data pipelines.
Experience developing and operating software in cloud environments using CI/CD, containerization, and cloud-native architectures.
Experience with service testing, reliability, observability, performance engineering, or production operations.
Ability to troubleshoot complex systems using logs, metrics, traces, telemetry, and production data.
Demonstrated ability to lead technical initiatives across teams, communicate architectural decisions clearly, and deliver high-quality production software.
Preferred Qualifications
Experience building developer platforms, CI/CD systems, testing frameworks, or engineering productivity tools.
Experience developing or integrating production solutions using machine learning, generative AI, or LLMs.
Familiarity with AI agents, tool-using models, retrieval-augmented generation, prompt engineering, or LLM evaluation techniques.
Experience with machine learning frameworks such as PyTorch, TensorFlow, Hugging Face Transformers, or equivalent technologies.
Experience with traffic replay, workload modeling, canary analysis, load testing, performance testing, or chaos engineering.
Hands-on experience with Kubernetes, containers, service meshes, and OCI or another major public cloud.
Experience analyzing API schemas, code changes, service dependencies, incidents, and production traffic.
Experience mentoring engineers and influencing architecture across teams.
Why Join Us
Shape a new AI-native developer platform from foundational architecture through production adoption.
Build systems that improve how thousands of engineers develop, test, and release software.
Solve complex, real-world problems spanning distributed systems, cloud infrastructure, machine learning, and generative AI.
Help engineering teams reduce manual testing, detect regressions earlier, diagnose failures faster, and release with greater confidence.
Work with experienced engineers, data scientists, product managers, and OCI service teams while expanding your technical influence and expertise.
#LI-SP1