Core responsibilities:
System Design & Architecture
- Develop and optimize scalable distributed systems using distributed state-management and data-plane technologies.
- Implement scalability requirements and performance/load testing for large-scale workloads.
- Build fault-tolerant components using redundancy, replication, failover, retries, circuit breakers, and timeouts.
- Implement telemetry, dashboards, alerts, and fault testing to ensure reliability, correctness, and availability.
- Apply replication and synchronization techniques to maintain data integrity.
Operational Excellence
- Diagnose and resolve production issues and automate troubleshooting and maintenance.
- Participate in incident response, operational support, and root cause analysis.
- Support in-service updates with minimal customer impact.
Security & Compliance
- Implement encryption, access controls, and security remediation.
- Ensure compliance with applicable standards and documentation requirements.
Automation & Change Management
- Maintain Infrastructure as Code (IaC) and operational automation.
- Follow change-management processes for patching, updates, and rollbacks.
Planning & Execution
- Track timelines, prioritize work, and adjust plans to meet project requirements.
Collaboration & Partnership
- Collaborate across teams and understand business, stakeholder, and customer needs.
- Seek and incorporate diverse perspectives.
Problem Solving
- Independently troubleshoot issues, escalate complex problems, and share best practices.
Continuous Learning
- Build new skills, stay current with industry trends, and promote knowledge sharing.
Continuous Improvement
- Identify and recommend improvements to processes, workflows, and team practices.