Job DescriptionInfluences and leads regular administration and conducts highly complex performance trend analyses and manages server capacity. Facilitates the optimization of system configurations and backups. Analyzes system performance data and insights to drive enhancements to improve the performance reliability, and security, of systems and environments. Serves as a subject matter expert in the investigation of highly complex system issues and facilitates high-severity incident triage. Designs disaster recovery solutions.
Responsibilities Key Responsibilities System Installation & Configuration – Software Administration:
- * Evaluates and sets standards for the performance and installation requirements of operating systems to ensure optimal installation and performance.
- Provides guidance on the administration of middleware products in environments.
- Facilitates and implements strategies for the deployment, maintenance, and operation of internal applications, ensuring the efficiency and performance of these systems.
- Influences and leads regular administration and conducts highly complex performance trend analyses and manages server capacity to ensure service performance meets and exceeds standards.
- Serves as a subject matter expert in utilizing application monitoring tools to optimize and ensure efficiency.
System Installation & Configuration – Installation and Configuration:
- * Leads the team installing and configuring servers, cloud infrastructure, and all software and environments.
- Facilitates the optimization of system configurations and backups to ensure the optimal performance and stability of the server infrastructure.
- Leads hardware maintenance, auditing, installation, and provisioning, ensuring all tasks are performed efficiently and effectively.
- Partners with internal technical experts and third-party vendors to resolve integration challenges, providing expert guidance and industry insights for innovative solutions.
System Installation & Configuration – Identity & Access Management:
- * Provides additional support and guidance for the administration of access privileges in the identity and access management system, ensuring accurate and secure access to IT resources.
- Interprets user activity data in the identity and access management system and leverages expertise to provide insights on system activity and recommend improvements.
- Designs and optimizes access management systems.
Service Lifecycle Management – Batch Processing:
- * Drives the monitoring and assessment of the batch process to ensure updates are applied and proactively resolves any issues that arise.
- Leverages industry insights to drive improvements in batch management techniques using different work schedulers to configure jobs and job streams, define dependencies, and report job performance.
- Influences and collaborates with teams to ensure scheduling and budgets of batch monitoring services align with and support Service Level Agreements (SLA).
Service Lifecycle Management – Security Maintenance:
- * Drives strategic improvements to procedures to ensure that compute and storage devices are secure.
- Influences and collaborates with teams to maintain privileged accounts/secrets integrity of systems and compute and file system security for the compute and storage environment.
- Analyzes and evaluates highly complex service and infrastructure dashboards, taking the lead in addressing identified anomalies.
- Coordinates and manages long-term implementation strategies for monthly, quarterly, or hotfix patches to address security vulnerabilities or bugs.
Service Lifecycle Management – System & Security Improvements:
- * Analyzes system performance data and insights to drive enhancements to improve the performance, reliability, and security of systems and environments.
- Influences and collaborates with Service teams to proactively identify, address, and predict gaps in operational capabilities, enhancing scalability and resiliency.
Incident Management & Support – Incident Management:
- * Leads end-to-end incident management lifecycle to ensure systems are stable, secure, and performing accurately.
- Evaluates results from incident-based data analyses for team metrics and key performance indicators (KPIs) to identify patterns, root causes, and solutions to prevent system and network incidents.
- Facilitates incident review meetings to provide strategic oversight for operational performance and long-term solution implementation.
- Influences third party vendors and cross-functional teams (e.g., Development, Cloud Engineering, Product Engineering, other IT teams) to develop and implement long-term solutions for high-severity incidents, risks, or migrations.
- Serves as a subject matter expert in the investigation of highly complex system issues and facilitates high-severity incident triage by designing Corrective and Preventative Action plans (CAPA) to drive incident resolution and prevention.
Incident Management & Support – Escalation Cases:
- * Provides expertise for escalated support cases by collaborating with internal technical teams and third party vendors to drive issue resolution for a wide range of production environment problems (e.g., immense growth, scaling, leveraging the cloud, extremely high performance, high availability requirements).
Incident Management & Support – Technical Support:
- * Drives and evaluates the production environment by analyzing system error logs and ticket queues, and coordinating with multiple teams involved in maintaining the environments.
- Adheres to team schedule to drive ongoing technical support and service objectives.
- Facilitates strategic resolution and long-term solutions for highly complex, critical customer system issues and develops and documents comprehensive technical solutions.
Incident Management & Support – Backups and Disaster Recovery:
- * Drives the execution and effectiveness of backup, restore, and disaster recovery processes.
- Influences the strategic planning and coordination of disaster recovery drills.
- Designs disaster recovery solutions to ensure preparedness and regulatory compliance.
Communication & Documentation – Technical Communication:
- * Communicates highly complex technical information to both technical and nontechnical personnel including management.
- Develops and implements training programs to ensure personnel are well-versed in domain-specific knowledge and practices.
- Drives technical strategies and solutions for cross-organization projects, programs, and activities by leveraging domain-specific expertise and interpreting highly complex technical information.
Communication & Documentation – Documentation & Reporting:
- * Drives the creation of documentation on ticket updates, code contributions, infrastructure, configurations, processes, and procedures (e.g., Disaster Recovery plans, Standard Operating Procedures, Corrective and Preventative Action Plans).
- Analyzes and interprets weekly and monthly reports on system performance and incident progress to provide insights on operational and management outcomes and business impacts.
- Reviews and refines technical documentation standards and best practices for internal use.
Additional Responsibilities (as Needed) Cloud Infrastructure Support:
- * Serves as a subject matter expert in collaborations with DevOps and Site Reliability Engineer (SRE) teams to manage large-scale infrastructure.
- Manages continuous integration and continuous deployment (CI/CD) pipelines.
- Outlines and plans for patching and version upgrades to support cloud infrastructure.
Automation:
- * Approves recommendations and determines plans for improvements to reduce incidents and problems with automation and simplify server management.
- Owns and leads the implementation of reusable frameworks, standards, and automation to support Oracle Cloud Infrastructure.
- Establishes and oversees Workload Automation tools through design support, administration, and optimization efforts.
- Champions troubleshooting complex issues with automation tools, agents, and other connectivity to 3rd party applications.
- Manages cloud technologies.
Core Responsibilities Planning & Execution:
- * Manages and provides direction on timelines, deliverables, and budgets when applicable for critical high-impact projects or initiatives that impact the line of business, ensuring timely completion and adherence to requirements.
- Anticipates and plans for shifts in resources or timelines based on changing business priorities, ensuring optimal outcomes.
Collaboration & Partnership:
- * Influences cross-functional leaders and external stakeholders to gain alignment on strategic objectives.
- Fosters partnerships with key business leaders, stakeholders, and/or customers, identifying opportunities for expanding partnerships and promoting long-term organizational success.
- Champions transparency and inclusivity by actively seeking, listening to, and incorporating diverse perspectives.
Problem Solving:
- * Leads specialized, advanced problem-solving efforts, serving as an escalation point for complex issues.
- Guides others to leverage innovative data-driven techniques to address ambiguous or novel issues, identify root causes, and drives the implementation of solutions that prevent future issues.
Continuous Learning:
- * Leverages deep industry knowledge and expertise to serve as a thought leader within the organization.
- Contributes to the advancement of the field or industry through thought leadership (e.g., conference presentations, white papers, research contributions).
- Maintains and evolves expertise in relevant areas by proactively monitoring emerging trends, technologies, and industry standards, ensuring the organization remains current with best practices.
- Champions continuous learning and knowledge sharing, promoting professional development across teams. Applies new knowledge to drive advancement and mentors others to do the same.
Continuous Improvement:
- * Develops innovative solutions and drives the implementation of ideas that increase the efficiency and effectiveness of processes, protocols, and workflows across the organization.
- Evaluates effectiveness of updated approaches and methods for continued improvement to enhance efficiencies and ensure changes align with organizational goals.
- Designs and develops metrics to measure success of improvement initiatives.
Performance and Development:
- * Serves as a subject matter expert regarding talent needs and organizational talent strategy.
- Imparts leadership and expert knowledge throughout the talent development pipeline including candidate interviews, candidate assessment, and hiring decisions, ensuring alignment with organizational talent strategy.
About UsOnly Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives.
True innovation starts when everyone is empowered to contribute. That’s why we’re committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs.
We’re committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing accommodation-request_mb@oracle.com or by calling 1-888-404-2494 in the United States.
Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.