purpose of the job:
Lead the design, deployment, and operation of scalable, secure HPC and AI infrastructure across on-premise and cloud environments, serving as the senior technical authority and escalation point for multiple client projects.
Mentor and manage a team of systems engineers while owning client relationships, project delivery, and pre-sales support to ensure SLA adherence and long-term account growth.
Requirements :
B.Sc. in Computer and/or Communications Engineering or similar degree that provides a strong knowledge of computer science.
3 – 6 years of hands-on experience in Linux systems engineering, HPC, cloud solution architect or infrastructure/DevOps roles, including experience leading engineers or technical workstreams.
Basic Qualifications:
Technical:
- Expert knowledge of Linux/Unix based systems administration and internals
- Proven hands-on experience with Kubernetes in production (deployment, upgrades, networking, storage, RBAC, Helm) and container technologies (Docker, containerd)
- Strong experience with configuration management and IaC tools (Ansible, Puppet, Terraform)
- Strong programming concepts and scripting skills [bash - python]
- Good knowledge of IP Networks (TCP/IP, VLANs, routing, DNS, load balancing)
- Experience with monitoring and observability tools
- Excellent troubleshooting capabilities & problem-solving skills with a structured, root-cause approach
Leadership & Management:
- Demonstrated experience leading, mentoring or supervising engineers
- Ability to delegate, prioritize and make sound technical decisions under pressure
- Basic knowledge of project management concepts (scope, planning, risk, stakeholder management) and methodologies (Agile/Scrum, Waterfall)
- Basic understanding of account management: managing client expectations, relationship building and identifying client needs
- Ability to explain technical topics clearly to non-technical stakeholders
Personal:
- We are looking for candidates who are smart, passionate, hard workers, with a high ability of self-learning, and who feel comfortable working on the edge of the unknown.
- Strong sense of ownership and accountability; leads by example
- Customer-oriented mindset with professional and diplomatic client handling
- Excellent oral and written communication and presentation skills
- Excellent oral and written English skills.
- Candidates are ready to travel 50% of the time in the GULF region.
Following will be considered as a plus:
- Kubernetes certifications (CKA, CKAD, CKS)
- Experience with GPU clusters and AI/ML infrastructure (e.g., NVIDIA GPU Operator, CUDA, DGX systems)
- Good knowledge of Virtualization technologies (i.e VMware, KVM, OpenStack)
- Good knowledge of Cloud Computing platforms/technologies (AWS, GCP, Azure,...) including managed Kubernetes (EKS, GKE, AKS)
- Project management or Agile certifications (PMP, CAPM, Scrum Master)
- Linux certifications (e.g., RHCE) and familiarity with ITIL / IT service management practices
- Experience with HPC environments: job schedulers (Slurm, PBS), parallel file systems, MPI and high-speed interconnects (InfiniBand)
- Experience/ knowledge with AI and Agentic AI fundamentals