Senior Solutions Architect, Post-Sales (EMEA) – AI/ML Infrastructure - EMEA Remote
Role:
A high-growth GPU cloud and AI infrastructure platform provider is looking for a Senior Solutions Architect to join its EMEA post-sales team. This is a customer-facing, highly technical role helping enterprise clients design, deploy, and scale production AI/ML workloads on a Kubernetes-native, GPU-accelerated Platform-as-a-Service. You'll sit at the intersection of platform engineering, MLOps, and infrastructure strategy, acting as a trusted technical advisor throughout the customer lifecycle.
You will be one of the primary technical voices shaping how enterprise customers architect and scale some of the most demanding AI infrastructure workloads in the market today, working directly with cutting-edge GPU fabric and large-scale distributed training environments.
This is a business riding the wave of explosive demand for AI compute, offering genuine scale, investment, and a seat at the table as the platform's capabilities and customer base expand rapidly.
You'll join a small, senior Solutions Architecture function with real autonomy, direct access to product and engineering leadership, and the flexibility of a role built around trust rather than micromanagement.
Responsibilities:
- Design end-to-end AI/ML platform architectures spanning inference, training, and data pipelines for enterprise customers
- Develop reference architectures for GPU cluster deployment, LLM serving, and multi-tenant ML infrastructure
- Advise on GPU fabric topology, including NVLink, InfiniBand, and RoCEv2, for distributed training environments
- Act as the primary technical advisor and escalation point for assigned customers, leading root cause analysis on complex production issues
- Deliver technical workshops, proof-of-concept engagements, and executive-level presentations on AI infrastructure strategy
- Design observability strategies across DCGM, OpenTelemetry, eBPF, and GPU metrics pipelines
- Partner with customer platform, MLOps, and data science stakeholders to translate workload requirements into scalable architecture
- Feed customer insights back into the product and engineering roadmap, and mentor junior members of the Solutions Architecture team
Skills/Must have:
- Experience: 8+ years in infrastructure, platform, or solutions engineering, including 3+ years focused specifically on AI/ML infrastructure or MLOps
- Core Tech/Domain: Deep, hands-on Kubernetes expertise (cluster lifecycle, workloads, operators, RBAC) plus direct experience with NVIDIA GPU infrastructure (H100/H200/B200 preferred)
- Methodology/Protocols: Working knowledge of distributed training concepts (NCCL, tensor and pipeline parallelism), LLM inference serving and optimisation (vLLM, NIM, TGI), and GPU networking (GPU Operator, MIG, SR-IOV, IB/RoCEv2)
- Soft Skills: Confident, credible communicator able to run technical discussions with engineers through to executive stakeholders; strong troubleshooting mindset and customer advisory presence
Nice to haves:
- Experience with Run:AI or Slurm for GPU scheduling and workload optimisation
- Familiarity with PyTorch or TensorFlow
- Public cloud experience (AWS, Azure, or GCP) across networking, IAM, and managed Kubernetes
- Certifications such as CKA, CKAD, AWS Solutions Architect, Azure Solutions Architect, or GCP Professional Cloud Architect
- Understanding of multi-tenant GPU isolation (SR-IOV VFs, DPU offload)
Benefits:
- Competitive base salary with performance-related bonus
- Remote-first/flexible working across EMEA
- Private healthcare and pension contribution
- Training and certification allowance
- Genuine exposure to cutting-edge GPU and AI infrastructure at enterprise scale
Salary:
£150,000