Takes proactive steps to design and architect infrastructure and service to ensure reliability and functionality. Forecasts demands and responds to capacity needs. Collaborates with software development teams to develop reliable and scalable infrastructures. Performs data collection to maintain and optimize operations and reliability. Leverages knowledge to perform incident response and/or maintenance tasks. Provides health and performance reporting. Identifies opportunities for automation. Communicates about services and identifies and explains the potential impact of changes. Provides support for technology and documents incidents. Experiments with new tools and assesses potential impact and develops knowledge of site reliability trends.
Internal Responsibilities
Key Responsibilities
- Administer and support Windows Server and Linux systems in production and non-production environments.
- Perform post-deployment configuration, application installation, service validation, and environment-readiness checks.
- Execute structured build and deployment activities using approved runbooks, deployment guides, and change procedures.
- Run, review, and troubleshoot scripts used for system configuration, deployments, patching, and validation.
- Investigate operating system, service, application, installation, patching, startup, and connectivity issues.
- Review logs, Windows Event Viewer, service status, permissions, processes, ports, certificates, and configuration files to identify and resolve problems.
- Apply operating system, middleware, and application patches, including pre-maintenance checks, post-patch validation, and rollback support.
- Support vulnerability remediation, system hardening, STIG implementation, and other compliance-driven configuration activities.
- Troubleshoot network and service-connectivity issues involving DNS, routing, firewalls, load balancers, ports, and certificates.
- Use diagnostic tools such as ping, traceroute, telnet, netcat, curl, nslookup, and netstat.
- Support cloud-hosted infrastructure involving compute, storage, networking, identity, access management, and environment provisioning.
- Participate in incident response, root-cause analysis, and service-restoration activities.
- Create and maintain technical documentation, runbooks, implementation records, validation results, and escalation notes.
- Identify opportunities to automate repetitive work, improve operational processes, and reduce recurring incidents.
- Participate in scheduled maintenance windows, after-hours support, or an on-call rotation as required.
Required Qualifications
- Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field, or equivalent practical experience.
- Three or more years of experience in site reliability engineering, systems administration, infrastructure operations, DevOps, cloud operations, production support, or a related technical field.
- Hands-on experience administering Windows Server, Linux, or both.
- Experience installing, configuring, validating, and troubleshooting applications in hosted environments.
- Experience completing structured build, deployment, configuration, or maintenance activities from technical runbooks.
- Ability to run, modify, or troubleshoot scripts using one or more of the following:
- PowerShell
- Bash
- Python
- Ansible
- Chef
- Experience applying operating system, middleware, or application patches.
- Understanding of maintenance windows, change control, rollback planning, and post-maintenance validation.
- Working knowledge of networking concepts, including DNS, routing, firewalls, load balancers, ports, and certificates.
- Familiarity with cloud-hosted infrastructure and services.
- Experience with ticketing, incident-management, or change-management systems.
- Strong troubleshooting, documentation, and problem-solving skills.
- Clear written and verbal communication skills.
- Ability to collaborate effectively across technical and client-facing teams.
Preferred Qualifications
- Experience with Oracle Cloud Infrastructure.
- Experience supporting Oracle Health Millennium or Cerner applications and environments.
- Experience supporting federal clients, government-hosted systems, or regulated environments.
- Familiarity with STIGs, vulnerability remediation, cybersecurity hardening, and federal compliance workflows.
- Familiarity with OCI, Scripting, federal compliance, healthcare or Millennium environments
- Experience with Citrix technologies.
- Experience supporting legacy systems or complex application architectures.
- Experience with infrastructure-as-code or configuration-management tools.
- Familiarity with production monitoring, alerting, centralized logging, and observability practices.
- Experience with incident response, root-cause analysis, and SRE operational practices.
- Knowledge of Oracle Database, SQL, middleware, or related Oracle technologies.
- Relevant certifications in Oracle Cloud Infrastructure, Windows Server, Linux, networking, cybersecurity, or cloud computing.
- Experience with creating CI/CD pipelines
External Responsibilities
Key Responsibilities
- Administer and support Windows Server and Linux systems in production and non-production environments.
- Perform post-deployment configuration, application installation, service validation, and environment-readiness checks.
- Execute structured build and deployment activities using approved runbooks, deployment guides, and change procedures.
- Run, review, and troubleshoot scripts used for system configuration, deployments, patching, and validation.
- Investigate operating system, service, application, installation, patching, startup, and connectivity issues.
- Review logs, Windows Event Viewer, service status, permissions, processes, ports, certificates, and configuration files to identify and resolve problems.
- Apply operating system, middleware, and application patches, including pre-maintenance checks, post-patch validation, and rollback support.
- Support vulnerability remediation, system hardening, STIG implementation, and other compliance-driven configuration activities.
- Troubleshoot network and service-connectivity issues involving DNS, routing, firewalls, load balancers, ports, and certificates.
- Use diagnostic tools such as ping, traceroute, telnet, netcat, curl, nslookup, and netstat.
- Support cloud-hosted infrastructure involving compute, storage, networking, identity, access management, and environment provisioning.
- Participate in incident response, root-cause analysis, and service-restoration activities.
- Create and maintain technical documentation, runbooks, implementation records, validation results, and escalation notes.
- Identify opportunities to automate repetitive work, improve operational processes, and reduce recurring incidents.
- Participate in scheduled maintenance windows, after-hours support, or an on-call rotation as required.
Required Qualifications
- Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field, or equivalent practical experience.
- Three or more years of experience in site reliability engineering, systems administration, infrastructure operations, DevOps, cloud operations, production support, or a related technical field.
- Hands-on experience administering Windows Server, Linux, or both.
- Experience installing, configuring, validating, and troubleshooting applications in hosted environments.
- Experience completing structured build, deployment, configuration, or maintenance activities from technical runbooks.
- Ability to run, modify, or troubleshoot scripts using one or more of the following:
- PowerShell
- Bash
- Python
- Ansible
- Chef
- Experience applying operating system, middleware, or application patches.
- Understanding of maintenance windows, change control, rollback planning, and post-maintenance validation.
- Working knowledge of networking concepts, including DNS, routing, firewalls, load balancers, ports, and certificates.
- Familiarity with cloud-hosted infrastructure and services.
- Experience with ticketing, incident-management, or change-management systems.
- Strong troubleshooting, documentation, and problem-solving skills.
- Clear written and verbal communication skills.
- Ability to collaborate effectively across technical and client-facing teams.
Preferred Qualifications
- Experience with Oracle Cloud Infrastructure.
- Experience supporting Oracle Health Millennium or Cerner applications and environments.
- Experience supporting federal clients, government-hosted systems, or regulated environments.
- Familiarity with STIGs, vulnerability remediation, cybersecurity hardening, and federal compliance workflows.
- Familiarity with OCI, Scripting, federal compliance, healthcare or Millennium environments
- Experience with Citrix technologies.
- Experience supporting legacy systems or complex application architectures.
- Experience with infrastructure-as-code or configuration-management tools.
- Familiarity with production monitoring, alerting, centralized logging, and observability practices.
- Experience with incident response, root-cause analysis, and SRE operational practices.
- Knowledge of Oracle Database, SQL, middleware, or related Oracle technologies.
- Relevant certifications in Oracle Cloud Infrastructure, Windows Server, Linux, networking, cybersecurity, or cloud computing.
- Experience with creating CI/CD pipelines