Leads development and begins architecting components of scalable, elastic distributed systems. Defines and enforces scalability requirements for owned components; optimizes code and data paths for high‑throughput, hyper‑scale workloads; and leverages data plane platforms for large‑scale retrieval, storage, and processing. Designs fault‑tolerant, in‑service‑upgradable systems using redundancy, replication, failover, and policies for partitions, applying load‑shedding, throttling, and rate‑limiting to handle network unreliability while meeting SLOs. Establishes KPIs and telemetry; builds proactive dashboards and alerts; and designs complex validation (fault injection, brownouts), replication, and synchronization for correctness and durability. Proactively diagnoses and resolves production issues, mentors peers, and ensures operational readiness. Implements robust security controls, executes remediation, maintains compliance documentation, and develops IaC and automation that enable safe patching, updates, and rollbacks within change‑management plans.
Internal Responsibilities
System Scalability
Develops distributed-system components supporting horizontal and vertical scaling through distributed-state-management tools.
Optimizes code and systems for large-scale data processing; implements scalability requirements and reviews team implementations.
Uses data-plane platform components for large-scale retrieval, storage, and processing.
Implements performance and load testing; designs, develops, tests, deploys, and maintains production software and cloud services.
Reliability, Correctness, and Availability
Operations, Security, and Change Management
- Diagnoses, debugs, and resolves component issues; implements strategies to prevent interruptions and avoid customer maintenance windows.
- Designs, implements, and maintains automation scripts, tooling, and IaC for troubleshooting and cloud-infrastructure management.
- Participates in operational-support rotations, incident response, root-cause investigations, and follow-up improvements.
- Applies advanced multi-tenant security measures, including encryption and access controls; implements remediation plans.
Ensures compliance with relevant industry standards and regulations, maintains current documentation, and follows change-management plans for patching, updates, and rollbacks.
Collaboration and Delivery
- Partners with product managers, architects, and engineering teams to translate requirements into technical solutions.
- Creates technical and operational documentation; drives projects forward and mentors junior team members.
Education and/or Experience
8 years of software-development experience; or
Bachelor’s degree in Computer Science, Computer Engineering, Software Engineering, Electrical/Electronics Engineering, Computer/Information Systems, Information Technology, Telecommunications, Mathematics, Physics, or related field, plus 4 years of software-development experience; or
Master’s degree in those fields, plus 2 years of software-development experience; or
Doctorate in those fields.
Skills and Experience
Demonstrated ability or knowledge in distributed systems; prototyping; computer-science programming; software engineering; web development; innovation; cross-functional collaboration; information vulnerabilities; operating systems; API development and integration; applied algorithm engineering; and source control.
Proficiency in Java, Python, Go, C++, C#, or a similar language.
Strong data structures, algorithms, object-oriented design, software-engineering, problem-solving, communication, and collaboration skills.
Experience building, testing, and debugging production-quality software; familiarity with REST APIs, databases, cloud applications, and agile methodologies.
Cloud-platform experience (AWS, Azure, Google Cloud, or Oracle Cloud); System-level testing and automation; Delivering and operating large-scale distributed systems.
External Responsibilities
System Scalability
Develops distributed-system components supporting horizontal and vertical scaling through distributed-state-management tools.
Optimizes code and systems for large-scale data processing; implements scalability requirements and reviews team implementations.
Uses data-plane platform components for large-scale retrieval, storage, and processing.
Implements performance and load testing; designs, develops, tests, deploys, and maintains production software and cloud services.
Reliability, Correctness, and Availability
Operations, Security, and Change Management
- Diagnoses, debugs, and resolves component issues; implements strategies to prevent interruptions and avoid customer maintenance windows.
- Designs, implements, and maintains automation scripts, tooling, and IaC for troubleshooting and cloud-infrastructure management.
- Participates in operational-support rotations, incident response, root-cause investigations, and follow-up improvements.
- Applies advanced multi-tenant security measures, including encryption and access controls; implements remediation plans.
Ensures compliance with relevant industry standards and regulations, maintains current documentation, and follows change-management plans for patching, updates, and rollbacks.
Collaboration and Delivery
- Partners with product managers, architects, and engineering teams to translate requirements into technical solutions.
- Creates technical and operational documentation; drives projects forward and mentors junior team members.
Education and/or Experience
8 years of software-development experience; or
Bachelor’s degree in Computer Science, Computer Engineering, Software Engineering, Electrical/Electronics Engineering, Computer/Information Systems, Information Technology, Telecommunications, Mathematics, Physics, or related field, plus 4 years of software-development experience; or
Master’s degree in those fields, plus 2 years of software-development experience; or
Doctorate in those fields.
Skills and Experience
Demonstrated ability or knowledge in distributed systems; prototyping; computer-science programming; software engineering; web development; innovation; cross-functional collaboration; information vulnerabilities; operating systems; API development and integration; applied algorithm engineering; and source control.
Proficiency in Java, Python, Go, C++, C#, or a similar language.
Strong data structures, algorithms, object-oriented design, software-engineering, problem-solving, communication, and collaboration skills.
Experience building, testing, and debugging production-quality software; familiarity with REST APIs, databases, cloud applications, and agile methodologies.
Cloud-platform experience (AWS, Azure, Google Cloud, or Oracle Cloud); System-level testing and automation; Delivering and operating large-scale distributed systems.