We are building an enterprise AI platform in which AI agents operate business applications through their user interface, the way a person does. It runs on Kubernetes across Microsoft Azure, on-premises, and air-gapped environments, delivered through GitOps. Because it is deployed inside customers' own perimeters, it must install and run with no dependency on cloud services in the critical path. We are a small, senior team working in two-week sprints towards a first production release in December 2026.
You will own the target architecture, hosting, and platform: infrastructure as code, cluster management, packaging and delivery, GPU model serving, sandboxed browser workloads, observability, security, and cost. You will lead platform engineering and mentor a junior DevOps engineer.
What you will do
- Define and own the target architecture and hosting. Produce the architecture and security documentation needed for internal approvals and security clearance.
- Design and run Kubernetes platforms across Azure (AKS) and on-premises clusters (e.g. RKE2, OpenShift, Rancher, or upstream Kubernetes), with consistent tooling and policies across both.
- Package the platform as Helm charts that install into customer environments, including air-gapped installs: offline registries (e.g. Harbor), image and chart mirroring, and signed release bundles.
- Build GitOps delivery with Argo CD or Flux: repository structure, environment promotion, progressive delivery, and drift detection.
- Define all infrastructure as code with Terraform / OpenTofu and/or Bicep.
- Run GPU workloads: GPU node pools and scheduling (NVIDIA GPU Operator), and self-hosted serving of open-weight multimodal models (e.g. vLLM, SGLang) behind an OpenAI-compatible gateway.
- Run fleets of sandboxed headless browsers for AI agents at scale, with strong isolation (e.g. gVisor, Kata Containers, network policies) and controlled egress.
- Run stateful services on Kubernetes without managed cloud equivalents: PostgreSQL (e.g. CloudNativePG), S3-compatible object storage, and secrets management (e.g. OpenBao / Vault).
- Own platform security: policy as code (OPA Gatekeeper / Kyverno), image scanning, SBOMs, and supply-chain security.
- Build observability: metrics, logs, and traces (Prometheus, Grafana, Loki, OpenTelemetry), LLM tracing (e.g. Langfuse), SLOs, and alerting.
- Own reliability: backup and disaster recovery, capacity planning, incident response, and blameless post-mortems.
- Lead and mentor the DevOps Engineer – CI/CD & Environments.
What we are looking for
Must have
- 10+ years in software, infrastructure, DevOps, SRE, or platform engineering, including production Kubernetes at scale.
- Experience defining target architectures and getting them through security and architecture reviews.
- Deep hands-on Azure experience (AKS, networking, identity, Key Vault, Monitor, ACR).
- Experience running on-premises Kubernetes and connecting cloud and on-premises environments.
- Experience delivering software into air-gapped or highly restricted environments.
- Production experience with GitOps (Argo CD or Flux), and writing and maintaining Helm charts.
- Strong infrastructure as code skills with Terraform / OpenTofu and/or Bicep.
- Solid Linux, networking (TCP/IP, DNS, TLS, load balancing), and scripting (Bash plus Python or Go).
- A security-first mindset: least privilege, secrets management, policy as code, and supply-chain security.
- A track record of technical leadership: owning architecture decisions, mentoring engineers, and raising engineering standards across a team.
Nice to have
- GPU workloads on Kubernetes and LLM serving (vLLM, SGLang, Triton, KServe).
- Running headless browser fleets (Chromium, Playwright) or other untrusted workloads in sandboxes.
- Service mesh (Istio, Linkerd, Cilium) and eBPF-based networking.
- CKA / CKS and/or Azure certifications (AZ-104, AZ-305, AZ-400).
- Experience in regulated or data-sovereign environments.
What we will assess
- System design: a platform that runs on Azure and in an air-gapped on-premises site, covering packaging, networking, identity, security, GPU serving, observability, and disaster recovery.
- Technical exercise: given a service, design its GitOps repository layout, IaC, and promotion flow across environments.
- Incident scenario: troubleshoot a realistic production failure.
- Leadership: how you would guide and grow a junior engineer.
Why join
- Own the architecture of a platform that installs and runs anywhere, from Azure to fully air-gapped sites.
- Run infrastructure for production AI workloads: GPUs, self-hosted models, and browser sandboxes.
- Lead platform engineering from day one.