About Nehme AI Labs:
Nehme AI Labs brings rigorous software engineering to AI. We audit enterprise AI architectures, cut inference costs by 40–60%, and build tools like FlashCheck that make AI systems cheaper and more reliable. We're a small, fully remote, async-first team. We hire based on proof, not pedigree.
About the Role
We're hiring a Senior Software Engineer to build and own the infrastructure behind our AI products — inference services, APIs, and internal tooling. You'll work directly with research to productionize models, make architecture calls, and set the engineering bar. This is a hands-on role: you ship, you own your work end-to-end, and you solve problems without being told how.
What You'll Do
-
Design, build, and operate backend services for model inference, batching, and caching
-
Own features end-to-end: architecture, implementation, deployment, and monitoring
-
Develop APIs and SDKs used by enterprise customers
-
Optimize inference and data pipelines for latency, throughput, and cost
-
Lead technical decisions on architecture and tooling, and explain the tradeoffs in writing
-
Mentor other engineers and raise the code quality bar through review and standards
-
Work with research to productionize new models
-
Keep production systems running and write docs that don't suck
What We're Looking For
-
3+ years of professional experience building and operating production software
-
Strong proficiency in Python, plus experience with at least one systems language (Rust, Go, C++)
-
You've shipped services with real users and kept them running
-
Solid cloud fundamentals (AWS or GCP) and containers (Docker, Kubernetes)
-
Sound technical judgment — you can defend a tradeoff, including the one to not build something
-
A track record of shipping: projects, open-source work, or products you can show us
-
Good async communication skills (we don't do meetings for things that should be docs)
Nice to Have
-
Experience with ML serving frameworks (Triton, vLLM, TGI) or LLM inference at scale
-
Evaluation or benchmark tooling experience
-
Developer tools with real adoption
-
Familiarity with WebAssembly or browser-based ML inference