We are building an enterprise AI platform in which AI agents operate business applications through their user interface, the way a person does. It runs inside customers' own environments, including on-premises and air-gapped sites. We are a small, senior team working in two-week sprints towards a first production release in December 2026.
You will own two things that decide whether our agents can be trusted:
- How agents learn a new application. They learn from human demonstrations and from exploring the application themselves, and they keep that knowledge up to date when the application changes.
- How we prove they work. You will build the evaluation framework that measures agent reliability and catches regressions before they reach users.
What you will do
- Build ways for agents to learn an application from human demonstrations and from self-exploration, and turn what they learn into reusable, testable procedures (playbooks).
- Represent what agents know about an application as structured knowledge, for example a knowledge graph of screens, actions, and data.
- Build simulated and recorded environments where agents can practise and be tested safely before acting on live systems.
- Make agents self-correcting: detect when an application changes or behaves unexpectedly, and update their knowledge and playbooks without breaking production tasks.
- Build the evaluation framework:
- task suites and success metrics
- LLM-as-judge with calibration against human labels
- regression testing whenever a prompt, model, or tool changes
- online monitoring of agent reliability
- Improve agent performance through prompt and program optimisation (e.g. DSPy, TextGrad, Ax), and through fine-tuning or reinforcement learning where it pays off.
- Work closely with the AI Engineer – Agent Harness, and mentor the L4 AI engineers on evaluation practice.
What we are looking for
Must have
- 10+ years of software engineering, including 2+ years building LLM-powered applications.
- Has designed and run rigorous evaluations for LLM or agent systems, and used the results to make decisions.
- Experience with at least one of: learning from demonstration, agent memory, or prompt or program optimisation.
- Strong Python. Comfortable with statistics and experiment design.
- Deep knowledge of the core AI topics below.
- A track record of technical leadership: owning architecture decisions, mentoring engineers, and raising engineering standards across a team.
Nice to have
- Reinforcement learning or fine-tuning for LLMs (LoRA / PEFT, SFT, preference tuning).
- Graph databases (e.g. Memgraph, Neo4j) used as agent memory.
- LLM observability and evaluation tools (e.g. Langfuse, Arize Phoenix).
- Working with recorded browser sessions (HAR / WARC, rrweb) as test environments.
- Published research, writing, or open-source work on agent evaluation or learning.
Core AI knowledge
We expect every AI Engineer on our team to be able to discuss these topics with confidence. For this role we expect depth in all of them, especially evaluation and benchmarks.
- How LLMs work:
- the Transformer architecture (attention, positional encodings, feed-forward layers), tokenisation, autoregressive decoding and sampling, and KV caching
- the training pipeline: pre-training, SFT, preference optimisation such as RLHF and DPO, and RL with verifiable rewards
- reasoning models and test-time compute; Mixture-of-Experts, quantisation, and serving
- failure modes: hallucination, prompt injection, context rot, and reward hacking
- Agentic AI:
- agent patterns: ReAct, plan-and-execute, reflection, orchestrator–worker, and evaluator–optimiser
- tool design and context engineering: retrieval, memory, compaction, and sub-agents
- the trade-offs of multi-agent systems
- safety controls: permissions, sandboxing, human-in-the-loop, and prompt-injection defence
- Protocols:
- Model Context Protocol (MCP): hosts, clients and servers; tools, resources and prompts; transports; and OAuth
- Agent2Agent (A2A)
- awareness of AG-UI and the OpenTelemetry GenAI conventions
- Agent harnesses:
- what sits around the model: the loop, tool registry, permission layers, sandboxes, context management, hooks, and skills
- hands-on experience with at least one harness or SDK (e.g. Claude Agent SDK, OpenAI Agents SDK, LangGraph, Google ADK, Pydantic AI, or a custom one)
- why the choice of harness changes benchmark scores
- Behavioural vs procedural approaches:
- when to let the model decide the steps (behavioural) and when to define them in code or workflows (procedural)
- how to combine the two, and how the balance shifts as models improve
- this applies both to system design and to how instructions are written
- Benchmarks:
- what current benchmarks measure and what they miss:
- computer use and web: OSWorld, WebArena, Online-Mind2Web, ScreenSpot, BrowseComp
- agents and tools: SWE-bench Verified / Pro, Terminal-Bench, τ²-bench, GAIA
- reasoning: Humanity's Last Exam, ARC-AGI, GPQA Diamond
- how to read results critically: contamination, harness effects, and cost
What we will assess
- Evaluation design: design an evaluation suite for an agent that operates a web application, including how you would validate an LLM-as-judge.
- Critical thinking: interpret a set of benchmark and evaluation results, and make a model or approach recommendation.
- System design: how an agent should detect that an application has changed and update what it knows, safely.
- Fundamentals: LLM training, evaluation, and the behavioural vs procedural trade-off.
Why join
- Own the part of the platform that makes agents trustworthy: learning, self-correction, and proof that they work.
- Turn current research on agent learning and evaluation into production engineering.
- Shape the AI engineering team and its evaluation standards from day one.