Key Responsibilities
System Installation &
Configuration – Software Administration:
-Monitors and
maintains operating systems to ensure effective installation and performance.
-Administers
middleware products in environments.
-Supports the deployment,
maintenance, and operation of internal applications as needed.
-Performs
regular administration and conducts standard performance trend analyses and
manages server capacity to ensure service performance meets standards.
-Leverages
working knowledge of application monitoring tools to ensure efficiency.
System Installation &
Configuration – Installation and Configuration:
-Installs and
configures servers, cloud infrastructure, and all software and environments.
-Performs system
configurations and backups independently to ensure optimization of server
infrastructure.
-Supports
hardware maintenance, auditing, installation, and provisioning as necessary.
-Collaborates
with internal technical experts and third-party vendors to resolve integration
challenges.
System Installation &
Configuration – Identity & Access Management:
-Administers
access privileges in the identity and access management system to ensure
accurate access to IT resources.
-Monitors user
activity in the identity and access management system and reports on system
activity.
-Learns to
provision access management systems.
Service Lifecycle Management
– Batch Processing:
-Monitors the
batch process to ensure the updates are applied and triages any issues that
arise.
-Implements
batch management using different work schedulers to configure jobs and job
streams, define dependencies, and report job performance.
-Ensures
scheduling and budgets of batch monitoring services comply with Service Level
Agreements (SLA).
Service Lifecycle Management
– Security Maintenance:
-Follows
existing procedures to provide assurance that compute and storage devices are
secure.
-Maintains
privileged accounts/secrets integrity of systems and compute and file system
security for the compute and storage environment.
-Monitors and evaluates
high-level service and infrastructure dashboards and takes action to address
identified anomalies.
-Implements
monthly, quarterly, or hotfix patches to address security vulnerabilities or
bugs across the technical stack.
Service Lifecycle Management
– System & Security Improvements:
-Deploys
enhancements to improve the performance, reliability, and security of systems
and environments.
Incident Management &
Support – Incident Management:
-Supports the
end-to-end incident management lifecycle to ensure systems are stable, secure,
and performing accurately.
-Collates
incident-based data for team metrics and key performance indicators (KPIs) by
assisting with system and network incidents to identify patterns, root causes,
and solutions.
-Participates in
incident review meetings to provide feedback for operational performance and
solution implementation.
-Partners with
third party vendors and cross-functional teams (e.g., Development, Cloud
Engineering, Product Engineering, other IT teams) to drive collaboration for
implementation and/or resolution for high-severity incidents, risks, or
migrations.
-Investigates
standard system issues and triages high-severity incidents by implementing
Corrective and Preventative Action plans (CAPA) as instructed.
Incident Management &
Support – Escalation Cases:
-Coordinates
escalated support cases by collaborating with internal technical teams and
third party vendors to drive issue resolution for a wide range of production
environment problems (e.g., immense growth, scaling, leveraging the cloud,
extremely high performance, high availability requirements).
Incident Management &
Support – Technical Support:
-Proactively
monitors the production environment by checking system error logs, monitoring
ticket queues, and consulting with other teams involved in maintaining the
environments.
-Adheres to team
schedule to contribute to ongoing technical support and service objectives.
-Resolves
complex, critical customer system issues and implements and documents technical
solutions.
Incident Management &
Support – Backups and Disaster Recovery:
-Leverages
working knowledge of systems to perform backup, restore, and disaster recovery
processes.
-Participates in
disaster recovery drills to ensure preparedness and compliance.
-Implements
disaster recovery solutions to ensure preparedness and regulatory compliance.
Communication &
Documentation – Technical Communication:
-Communicates
technical information to both technical and nontechnical personnel.
-Serves as a
technical liaison and provides domain-specific expertise to cross-organization
projects, programs, and activities.
Communication &
Documentation – Documentation & Reporting:
-Maintains
documentation on ticket updates, code contributions, infrastructure,
configurations, processes, and procedures (e.g., Disaster Recovery plans,
Standard Operating Procedures, Corrective and Preventative Action Plans).
-Generates
weekly and monthly reports on system performance and incident progress to
support operational and management outcomes, and builds an awareness of the
business impacts.
-Researches,
proofs, and authors technical documentation in the area of standards and best
practices for internal use.
Additional Responsibilities
(as needed)
Cloud Infrastructure
Support:
-Collaborates
with DevOps and Site Reliability Engineer (SRE) teams to provide support for
large-scale infrastructure.
-Implements
continuous integration and continuous deployment (CI/CD) pipelines
independently.
-Performs
patching and version upgrades independently to support cloud infrastructure.
Automation:
-Supports
Workload Automation tools through design support, administration, and
optimization efforts.
-Troubleshoots
issues with automation tools, agents, and other connectivity to 3rd party
applications.
-Maintains and
supports cloud technologies.
Core Responsibilities
Planning & Execution:
-Independently
manages work, monitoring timelines and deliverables to ensure projects or
initiatives stay on track and meet requirements.
-Proactively
prioritizes work and adapts to resource or timeline shifts, suggesting
adjustments to maintain project efficiency.
Collaboration &
Partnership:
-Collaborates
across teams to align on expectations and achieve shared objectives.
-Builds and
maintains a comprehensive understanding of business, stakeholder, and/or
customer needs to build and support effective partnerships.
-Actively
listens to diverse perspectives and asks questions to ensure understanding of
others.
Problem Solving:
-Independently
identifies and addresses standard and non-standard issues in accordance with
standard practices, escalating more complex issues as appropriate.
-Analyzes data
and/or information from multiple sources to troubleshoot standard and
non-standard errors.
-Contributes to
knowledge sharing and best practices.
Continuous Learning:
-Embraces
continuous learning by actively seeking to build knowledge and new skills
and/or tools and staying current with industry trends and best practices.
-Seeks out and
leverages feedback and training to improve skills.
-Contributes to
a culture of continuous learning and knowledge sharing with team members.
Continuous Improvement:
--Develops ideas
and recommends updates to increase the efficiency and effectiveness of
processes, protocols, and workflows within a team.
-Seeks input
from team members on alternative approaches and methods for improving work.