Designs, implements, and optimizes components in distributed systems with an emphasis on scalability, resiliency, and operability. Delivers features and load/performance tests; leverages data plane platforms and distributed state tools for high-volume retrieval, storage, and processing; and reviews peers’ implementations for scalability compliance. Builds fault-tolerant paths (redundancy, replication, automatic failover), applies recovery‑oriented principles, and implements retries, circuit breakers, and timeouts. Proactively detects and mitigates issues via tests, alarms, dashboards, and telemetry; authors runbooks and participates in incident response and RCAs. Implements standard replication and synchronization, develops automation/IaC for troubleshooting and maintenance, and applies advanced security controls (encryption, access, remediation) while ensuring change, compliance, and documentation standards are met.
Internal Responsibilities
Qualifications:
- 5+ years’ experience delivering and operating large scale, highly available distributed systems.
- Strong knowledge and interest in AI adoption including prompt engineering and agentic programming, with ChatGPT and Codex experience a plus.
- Strong knowledge of a base language such as Java, with a preference for functional programming language such as Scala.
- Strong knowledge of data structures, algorithms, operating systems, and distributed systems fundamentals.
- Experience with tools such as Terraform for Infrastructure as Code.
- Deep knowledge with networking protocols (TCP/IP, HTTP) and network architectures.
- Ability to design, troubleshoot and maintain networking infrastructure for high throughput use cases.
- Strong understanding of databases, storage, and distributed persistence technologies.
- Strong troubleshooting and performance tuning skills.
- Experience building multi-tenant, virtualized infrastructure a strong plus.
External Responsibilities
Qualifications:
- 5+ years’ experience delivering and operating large scale, highly available distributed systems.
- Strong knowledge and interest in AI adoption including prompt engineering and agentic programming, with ChatGPT and Codex experience a plus.
- Strong knowledge of a base language such as Java, with a preference for functional programming language such as Scala.
- Strong knowledge of data structures, algorithms, operating systems, and distributed systems fundamentals.
- Experience with tools such as Terraform for Infrastructure as Code.
- Deep knowledge with networking protocols (TCP/IP, HTTP) and network architectures.
- Ability to design, troubleshoot and maintain networking infrastructure for high throughput use cases.
- Strong understanding of databases, storage, and distributed persistence technologies.
- Strong troubleshooting and performance tuning skills.
- Experience building multi-tenant, virtualized infrastructure a strong plus.