About Gaia
Gaia is building a sovereign enterprise AI platform that connects an organization’s knowledge, systems, and workflows through enterprise search, AI assistants, and agents.
Our deployments are not limited to a single cloud model. Depending on the customer, Gaia may run in Gaia Cloud, the customer’s own cloud environment (BYOC), a private VPC, or fully on-premises infrastructure.
We are looking for a Senior DevOps Engineer who thinks beyond a single cloud provider and can design infrastructure that is secure, portable, automated, observable, and production-ready across very different environments.
The Role
This is not a role for someone whose infrastructure knowledge ends with clicking through AWS or Azure consoles. You will be responsible for helping us build a cloud-agnostic deployment platform that allows Gaia to be deployed consistently across AWS, Azure, Google Cloud, OCI, and customer-controlled on-prem environments. You should be comfortable moving between managed cloud services and self-hosted open-source equivalents, understanding the trade-offs between them, and designing infrastructure that does not unnecessarily lock Gaia into one provider.
You will work closely with engineering, AI, security, and customer technical teams to take Gaia from code to reliable enterprise production environments.
What You’ll Own
- Design and operate cloud-agnostic infrastructure architectures across AWS, Azure, GCP, OCI, and on-premises environments.
- Build repeatable BYOC and on-prem deployment patterns for enterprise customers.
- Own Kubernetes infrastructure across both managed Kubernetes services and self-managed/vanilla Kubernetes environments.
- Build reusable Infrastructure-as-Code and deployment automation so environments can be provisioned consistently across providers.
- Design deployment abstractions that minimize provider-specific dependencies and make Gaia portable between environments.
- Manage production, staging, development, sandbox, and customer-specific environments.
- Build and maintain CI/CD and GitOps workflows.
- Own Kubernetes deployment patterns, Helm charts, manifests, secrets, configuration, upgrades, autoscaling, and rollback strategies.
- Operate and optimize self-hosted infrastructure components such as databases, caching, search/vector infrastructure, observability, and supporting services.
- Design secure network architectures including VPCs, private subnets, ingress/egress controls, VPN/private connectivity, load balancers, firewalls, DNS, and certificates.
- Implement proper tenant and environment isolation.
- Build observability across infrastructure and applications using metrics, logs, traces, dashboards, and actionable alerting.
- Define infrastructure SLOs, health checks, disaster recovery strategies, backup policies, and recovery procedures.
- Implement infrastructure security controls including IAM, RBAC, least privilege, secrets management, image scanning, patching, vulnerability management, and audit logging.
- Support deployment into highly restricted customer environments, including environments with limited or no internet connectivity.
- Work with customer infrastructure and cybersecurity teams during enterprise deployments and technical assessments.
- Troubleshoot complex production issues spanning networking, Kubernetes, storage, databases, GPUs, applications, and cloud infrastructure.
- Perform capacity planning and infrastructure right-sizing based on users, indexed data, workloads, and expected AI usage.
- Continuously optimize infrastructure cost without compromising availability or security.
- Build tooling that allows Gaia to provision, upgrade, monitor, pause, and manage customer deployments throughout their lifecycle.
- Document architecture, deployment procedures, runbooks, disaster recovery procedures, and operational standards.
What We’re Looking For
Required
- 2+ years of hands-on DevOps, SRE, Platform Engineering, or Infrastructure Engineering experience.
- Strong production experience with Kubernetes.
- Strong experience with Infrastructure as Code, ideally Terraform/OpenTofu.
- Strong Linux and networking fundamentals.
- Strong understanding of containers, Docker, Kubernetes networking, storage, ingress, scheduling, and resource management.
- Experience designing CI/CD and GitOps workflows.
- Experience operating production workloads in at least two major cloud providers.
- Ability to understand a cloud-native service and design a reasonable self-hosted or open-source equivalent when required.
- Experience with monitoring, logging, tracing, and alerting.
- Solid understanding of IAM, secrets management, TLS, network security, RBAC, and infrastructure hardening.
- Experience operating databases and stateful workloads in Kubernetes or private infrastructure.
- Strong troubleshooting ability across the entire infrastructure stack.
- Ability to design for reliability rather than simply “make the deployment work.”
- Comfortable working directly with enterprise customer infrastructure teams.
- Strong written technical documentation skills.
Strong PlusExperience with:
- AWS, Azure, GCP and/or OCI.
- Managed and vanilla Kubernetes.
- Argo CD and GitOps.
- Helm.
- Terraform / OpenTofu.
- Prometheus.
- Grafana.
- OpenTelemetry.
- Loki or similar logging infrastructure.
- PostgreSQL.
- Redis / Valkey.
- OpenSearch / Elasticsearch.
- Object storage such as S3-compatible systems.
- HashiCorp Vault or equivalent secrets-management systems.
- Air-gapped or restricted-network deployments.
- Private registries and offline container/image distribution.
- High-availability infrastructure.
- Backup and disaster recovery.
- Enterprise SSO infrastructure such as SAML/OIDC.
- Security frameworks and regulated enterprise environments.
AI Infrastructure Experience Is a Major Plus
Gaia operates AI workloads as part of its infrastructure, so experience with any of the following is highly valuable:
- NVIDIA GPUs.
- GPU scheduling in Kubernetes.
- CUDA and NVIDIA Container Toolkit.
- NVIDIA GPU Operator.
- GPU monitoring.
- Model serving infrastructure.
- vLLM or similar inference servers.
- Embedding and reranking workloads.
- GPU memory and inference optimization.
- Running multiple AI workloads on shared GPU infrastructure.
- Capacity planning for inference workloads.
You do not need to be an ML engineer, but you should be comfortable treating AI infrastructure as a production workload rather than a black box.