We are building an enterprise AI platform in which AI agents operate business applications through their user interface, the way a person does. It runs inside customers' own environments, including on-premises and air-gapped sites. We are a small, senior team working in two-week sprints towards a first production release in December 2026.
You will lead the design and build of our agent harness: the framework around the model that turns open-weight LLMs into agents that plan, act, recover, and can be trusted in production. You will also build the agent that operates web applications through the browser. This is the hardest technical problem we have, and you will set the technical direction for the AI engineering team.
What you will do
- Design and build the agent harness:
- the agent loop
- tool registry and tool execution
- context management and compaction
- memory
- permission and approval gates
- sandboxed execution
- retries and recovery
- Build the agent that operates web applications, using vision-language models and browser automation (Playwright, Chrome DevTools Protocol) in isolated browser sandboxes.
- Coordinate specialised agents that work together on a task. Decide where the model should be in control and where code should be (behavioural vs procedural).
- Make every agent run traceable and replayable: structured logs, traces (e.g. Langfuse, OpenTelemetry), and step-level replay.
- Integrate self-hosted open-weight models through an OpenAI-compatible gateway, and expose agent capabilities as MCP tools.
- Work with the Backend, Full-stack and DevOps engineers to run agents reliably in production.
- Set engineering standards for the AI team, review designs and code, and mentor the L4 AI engineers.
What we are looking for
Must have
- 10+ years of software engineering, including 2+ years building LLM-powered applications.
- Has designed and built an agent harness or a substantial agent framework that ran in production or a serious pilot.
- Strong Python and/or TypeScript, with solid systems design: event loops, state machines, concurrency, failure handling.
- Hands-on experience with browser automation (Playwright, Puppeteer, or Chrome DevTools Protocol).
- Deep knowledge of the core AI topics below.
- A track record of technical leadership: owning architecture decisions, mentoring engineers, and raising engineering standards across a team.
Nice to have
- Contributions to an open-source agent framework or harness.
- Experience with vision-language models for UI agents (e.g. Qwen-VL, UI-TARS) or agent-oriented browser frameworks (e.g. Stagehand, browser-use).
- State-machine or workflow engines (e.g. XState, Temporal, Trigger.dev).
- Self-hosting open-weight models (vLLM, SGLang) on Kubernetes and GPUs.
- Building AI systems for on-premises, air-gapped, or regulated environments.
Core AI knowledge
We expect every AI Engineer on our team to be able to discuss these topics with confidence. For this role we expect depth in all of them.
- How LLMs work:
- the Transformer architecture (attention, positional encodings, feed-forward layers), tokenisation, autoregressive decoding and sampling, and KV caching
- the training pipeline: pre-training, SFT, preference optimisation such as RLHF and DPO, and RL with verifiable rewards
- reasoning models and test-time compute; Mixture-of-Experts, quantisation, and serving
- failure modes: hallucination, prompt injection, context rot, and reward hacking
- Agentic AI:
- agent patterns: ReAct, plan-and-execute, reflection, orchestrator–worker, and evaluator–optimiser
- tool design and context engineering: retrieval, memory, compaction, and sub-agents
- the trade-offs of multi-agent systems
- safety controls: permissions, sandboxing, human-in-the-loop, and prompt-injection defence
- Protocols:
- Model Context Protocol (MCP): hosts, clients and servers; tools, resources and prompts; transports; and OAuth
- Agent2Agent (A2A)
- awareness of AG-UI and the OpenTelemetry GenAI conventions
- Agent harnesses:
- what sits around the model: the loop, tool registry, permission layers, sandboxes, context management, hooks, and skills
- hands-on experience with at least one harness or SDK (e.g. Claude Agent SDK, OpenAI Agents SDK, LangGraph, Google ADK, Pydantic AI, or a custom one)
- why the choice of harness changes benchmark scores
- Behavioural vs procedural approaches:
- when to let the model decide the steps (behavioural) and when to define them in code or workflows (procedural)
- how to combine the two, and how the balance shifts as models improve
- this applies both to system design and to how instructions are written
- Benchmarks:
- what current benchmarks measure and what they miss:
- computer use and web: OSWorld, WebArena, Online-Mind2Web, ScreenSpot, BrowseComp
- agents and tools: SWE-bench Verified / Pro, Terminal-Bench, τ²-bench, GAIA
- reasoning: Humanity's Last Exam, ARC-AGI, GPQA Diamond
- how to read results critically: contamination, harness effects, and cost
What we will assess
- System design: design a harness for an agent that operates a web application with no API, covering the loop, tools, context, permissions, recovery, and observability.
- Hands-on exercise: extend a small agent harness with a new tool, an approval gate, and recovery from a failed action.
- Fundamentals: LLM architecture, training, and the behavioural vs procedural trade-off. We want explanations, not memorised definitions.
- Leadership: how you set technical direction and grow other engineers.
Why join