As a member of the software engineering division, you will lead the design, development, debugging, and improvement of software supporting OCI’s hardware repair platform. You will lead delivery and contribute to the architecture of distributed-system components, defining scalability, reliability, and performance requirements for the components you own. You will provide technical guidance to other engineers, evolve software architecture to meet changing needs, and ensure services are ready to operate reliably in production.
Minimum Qualifications:
Meet one of the following education and software development experience combinations:
- 11 years of experience in software development; or
- A bachelor’s degree, including a B.Tech, in a relevant field and 7 years of experience in software development; or
- A master’s degree, including an M.Tech, in a relevant field and 5 years of experience in software development; or
- A doctorate in a relevant field and 3 years of experience in software development.
Relevant fields include Computer Science, Computer Engineering, Software Engineering, Electrical/Electronics Engineering, Computer Information Systems, Information Systems, Information Technology, Telecommunications, Mathematics, Physics, or a related field.
In addition:
- 3 years of academic or professional experience with cloud platforms such as AWS, Azure, Google Cloud, or OCI.
- 5 years of experience in testing and automation at the system level.
- 4 years of experience delivering and operating large-scale distributed systems.
- Proficiency in Java and strong knowledge of data structures, algorithms, operating systems, and distributed systems fundamentals.
- Strong troubleshooting and performance tuning skills, including using operational data to identify issues and guide improvements.
- Experience developing, testing, and integrating secure REST APIs and services, with knowledge of modern web application development and prototyping.
- Experience designing automation, writing and maintaining scripts such as Python or Perl, and developing software test plans and quality assurance practices.
- Experience conducting in-depth code reviews, using source control, estimating software effort, and delivering software through agile development practices.
- Ability to assess security vulnerabilities, apply security and compliance requirements, and provide technical guidance while collaborating across teams.
Preferred Qualifications:
- Experience building multi-tenant, virtualized infrastructure.
- Strong understanding of databases, NoSQL systems, storage, and distributed persistence technologies.
Internal Responsibilities
- Lead development of owned components and contribute to their architecture, defining requirements for elastic scaling, throughput, availability, and durability.
- Design fault-tolerant services that support in-service updates, make informed consistency and availability trade-offs during network partitions, and use load shedding, throttling, and rate limiting to meet service level objectives.
- Define operational metrics and telemetry, design performance and failure-scenario tests, and establish operational readiness.
- Serve in support rotations and guide incident response and root-cause investigations. Develop infrastructure automation and safe update and rollback plans; implement security controls, remediation, and compliance documentation.
- Coordinate delivery of moderately complex initiatives, mentor engineers, improve practices across teams, and contribute to candidate interviews and hiring recommendations.
External Responsibilities
- Lead development of owned components and contribute to their architecture, defining requirements for elastic scaling, throughput, availability, and durability.
- Design fault-tolerant services that support in-service updates, make informed consistency and availability trade-offs during network partitions, and use load shedding, throttling, and rate limiting to meet service level objectives.
- Define operational metrics and telemetry, design performance and failure-scenario tests, and establish operational readiness.
- Serve in support rotations and guide incident response and root-cause investigations. Develop infrastructure automation and safe update and rollback plans; implement security controls, remediation, and compliance documentation.
- Coordinate delivery of moderately complex initiatives, mentor engineers, improve practices across teams, and contribute to candidate interviews and hiring recommendations.