Designs, implements, and optimizes components in distributed systems with an emphasis on scalability, resiliency, and operability. Delivers features and load/performance tests; leverages data plane platforms and distributed state tools for high-volume retrieval, storage, and processing; and reviews peers’ implementations for scalability compliance. Builds fault-tolerant paths (redundancy, replication, automatic failover), applies recovery‑oriented principles, and implements retries, circuit breakers, and timeouts.
Proactively detects and mitigates issues via tests, alarms, dashboards, and telemetry; authors runbooks and participates in incident response and RCAs. Implements standard replication and synchronization, develops automation/IaC for troubleshooting and maintenance, and applies advanced security controls (encryption, access, remediation) while ensuring change, compliance, and documentation standards are met.
Internal Responsibilities
At Oracle Cloud Infrastructure (OCI), we build the future of the cloud for enterprises. We combine the pace and ownership of a start-up with the scale, reliability, and customer focus required to run one of the world’s leading enterprise cloud platforms.
We are seeking a Senior Core Infrastructure Engineer to join our DNS Control Plane team. You will design, build, and operate the systems that safely configure and govern DNS at global scale—turning customer and platform intent into correct, auditable, and resilient DNS behavior across a distributed cloud.
What You’ll Do
- Design and implement control-plane services and APIs for DNS zones, records, views, policies, and routing configuration.
- Build reliable workflows for validation, change orchestration, propagation, reconciliation, and lifecycle management across distributed DNS infrastructure.
- Develop production services and libraries primarily in Java, with supporting automation in Python where needed, and a strong focus on API design and correctness.
- Use AI-assisted engineering tools and agentic workflows to accelerate implementation, testing, incident investigation, and operational automation, with appropriate human review, security, and quality controls.
- Improve safety and operational resilience through idempotent workflows, optimistic concurrency, auditability, rollback paths, canary releases, and feature flags.
- Own observability and operational readiness: metrics, logs, tracing, SLOs, runbooks, capacity planning, and thoughtful on-call participation.
- Partner with DNS data plane, security, networking, and SRE teams to deliver secure, end-to-end platform improvements.
- Contribute technical leadership through design reviews, code reviews, documentation, and mentoring of other engineers.
Required Qualifications
- 8+ years of professional software engineering experience, with a track record of shipping and operating resilient production services.
- 4+ years of hands-on experience building production-grade distributed systems, cloud services, and/or control-plane platforms.
- Strong proficiency in Java, including production service development, concurrent programming, and performance troubleshooting.
- Demonstrated ability to use AI-assisted development tools and agentic workflows to improve design, code generation, testing, troubleshooting, and process automation while maintaining strong engineering judgment.
- Strong API design skills, including RESTful service design and OpenAPI/Swagger or comparable API contracts.
- Experience designing reliable stateful workflows, including asynchronous processing, retries, idempotency, concurrency control, and failure recovery.
- Solid understanding of Linux systems, networking fundamentals, and highly available distributed systems.
- Experience with containers and orchestration technologies such as Docker and Kubernetes.
- Experience with infrastructure-as-code and configuration automation, especially Terraform and tools such as Ansible or Chef.
- Strong testing discipline across unit, integration, and end-to-end testing, with a commitment to secure and safe rollout practices.
- Clear written and verbal communication skills and the ability to collaborate effectively across organizational boundaries.
Preferred Qualifications
- DNS expertise, including authoritative DNS, zone and record management, DNS delegation, TTLs, DNSSEC, and DNS change propagation.
- Experience with cloud control planes, resource lifecycle management, or multi-tenant platform APIs.
- Knowledge of internet protocols and security standards, including DNS, TLS, HTTP, and REST.
- Experience with relational or distributed data stores and schema evolution for highly available services.
- Familiarity with modern CI/CD platforms such as OCI DevOps, GitHub Actions, or GitLab CI/CD, and engineering productivity tools including Bitbucket, Jira, and Confluence.
- Experience operating mission-critical services in a 24/7 environment.
What Success Looks Like
- You deliver control-plane capabilities that make DNS changes safer, faster, and easier for customers and internal operators.
- You improve the reliability, auditability, and operability of DNS configuration workflows at scale.
- You raise engineering standards through pragmatic design, rigorous testing, and collaborative technical leadership.
External Responsibilities
At Oracle Cloud Infrastructure (OCI), we build the future of the cloud for enterprises. We combine the pace and ownership of a start-up with the scale, reliability, and customer focus required to run one of the world’s leading enterprise cloud platforms.
We are seeking a Senior Core Infrastructure Engineer to join our DNS Control Plane team. You will design, build, and operate the systems that safely configure and govern DNS at global scale—turning customer and platform intent into correct, auditable, and resilient DNS behavior across a distributed cloud.
What You’ll Do
- Design and implement control-plane services and APIs for DNS zones, records, views, policies, and routing configuration.
- Build reliable workflows for validation, change orchestration, propagation, reconciliation, and lifecycle management across distributed DNS infrastructure.
- Develop production services and libraries primarily in Java, with supporting automation in Python where needed, and a strong focus on API design and correctness.
- Use AI-assisted engineering tools and agentic workflows to accelerate implementation, testing, incident investigation, and operational automation, with appropriate human review, security, and quality controls.
- Improve safety and operational resilience through idempotent workflows, optimistic concurrency, auditability, rollback paths, canary releases, and feature flags.
- Own observability and operational readiness: metrics, logs, tracing, SLOs, runbooks, capacity planning, and thoughtful on-call participation.
- Partner with DNS data plane, security, networking, and SRE teams to deliver secure, end-to-end platform improvements.
- Contribute technical leadership through design reviews, code reviews, documentation, and mentoring of other engineers.
Required Qualifications
- 8+ years of professional software engineering experience, with a track record of shipping and operating resilient production services.
- 4+ years of hands-on experience building production-grade distributed systems, cloud services, and/or control-plane platforms.
- Strong proficiency in Java, including production service development, concurrent programming, and performance troubleshooting.
- Demonstrated ability to use AI-assisted development tools and agentic workflows to improve design, code generation, testing, troubleshooting, and process automation while maintaining strong engineering judgment.
- Strong API design skills, including RESTful service design and OpenAPI/Swagger or comparable API contracts.
- Experience designing reliable stateful workflows, including asynchronous processing, retries, idempotency, concurrency control, and failure recovery.
- Solid understanding of Linux systems, networking fundamentals, and highly available distributed systems.
- Experience with containers and orchestration technologies such as Docker and Kubernetes.
- Experience with infrastructure-as-code and configuration automation, especially Terraform and tools such as Ansible or Chef.
- Strong testing discipline across unit, integration, and end-to-end testing, with a commitment to secure and safe rollout practices.
- Clear written and verbal communication skills and the ability to collaborate effectively across organizational boundaries.
Preferred Qualifications
- DNS expertise, including authoritative DNS, zone and record management, DNS delegation, TTLs, DNSSEC, and DNS change propagation.
- Experience with cloud control planes, resource lifecycle management, or multi-tenant platform APIs.
- Knowledge of internet protocols and security standards, including DNS, TLS, HTTP, and REST.
- Experience with relational or distributed data stores and schema evolution for highly available services.
- Familiarity with modern CI/CD platforms such as OCI DevOps, GitHub Actions, or GitLab CI/CD, and engineering productivity tools including Bitbucket, Jira, and Confluence.
- Experience operating mission-critical services in a 24/7 environment.
What Success Looks Like
- You deliver control-plane capabilities that make DNS changes safer, faster, and easier for customers and internal operators.
- You improve the reliability, auditability, and operability of DNS configuration workflows at scale.
- You raise engineering standards through pragmatic design, rigorous testing, and collaborative technical leadership.