We are building an enterprise AI platform in which AI agents operate business applications through their user interface, the way a person does. It runs inside customers' own environments, including on-premises and air-gapped sites. We are a small team working in two-week sprints towards a first production release in December 2026.
You will put our agents to work on real web portals. Using the team's agent framework, you will teach agents to navigate a wide range of real-world websites, keep them working as those sites change, and make sure the business users who supervise them can rely on the results.
What you will do
- Teach agents to operate new web portals: navigation, search, filtering, pagination, and downloading documents.
- Handle the messy reality of real websites:
- logins and session handling
- steps that need a person (for example multi-factor login), handed over cleanly to a human
- broken search, inconsistent categories, and unexpected pop-ups
- Monitor agent runs, diagnose failures from traces and replays, and fix them.
- Detect when a portal changes and work with the AI Engineer – Agent Learning & Evaluation so agents adapt without breaking.
- Build tests against recorded sessions so changes can be checked before they reach production.
- Work with the business users who supervise the agents to understand what "done" looks like and turn their feedback into improvements.
What we are looking for
Must have
- 6+ years of software engineering, including 1+ year building LLM agents or browser automation at scale.
- Strong experience with browser automation: Playwright, Puppeteer, or Selenium. A good understanding of the DOM, the Chrome DevTools Protocol, network traffic, and authentication flows.
- Strong Python and/or TypeScript.
- Excellent debugging skills: you enjoy working out why something failed on a site you don't control.
- Working knowledge of the core AI topics below.
Nice to have
- Vision-language models for UI agents (e.g. Qwen-VL, UI-TARS) and grounding actions to screen elements.
- Agent-oriented browser frameworks (e.g. Stagehand, browser-use).
- Experience making large-scale web automation reliable (retries, rate limits, change detection).
- Experience with Arabic-language websites.
Core AI knowledge
We expect every AI Engineer on our team to be able to discuss these topics with confidence. For this role we expect depth in agentic AI and computer use, and working knowledge of the rest.
- How LLMs work: the Transformer architecture, tokenisation, decoding and sampling, the training pipeline (pre-training, SFT, RLHF / DPO, RL with verifiable rewards), and common failure modes such as hallucination and prompt injection.
- Agentic AI: agent patterns (ReAct, plan-and-execute, reflection), tool design, context engineering, and safety controls (permissions, sandboxing, human-in-the-loop).
- Computer use: how agents perceive and act on user interfaces: screenshots and vision-language models, DOM and accessibility trees, grounding actions to UI elements, and checking the result after each action.
- Protocols: Model Context Protocol (MCP) and Agent2Agent (A2A).
- Agent harnesses: what sits around the model (the loop, tools, permissions, sandboxes, context management), and hands-on experience with at least one agent framework or SDK.
- Behavioural vs procedural approaches: when to let the model decide the steps and when to define them in code, and how to combine the two.
- Benchmarks: what computer-use and web benchmarks (OSWorld, WebArena, Online-Mind2Web, ScreenSpot) measure, and why real-world reliability differs from benchmark scores.
What we will assess
- Hands-on exercise: build an agent that completes a multi-step task on a test web application, then make it recover from a change to that application.
- Debugging: diagnose a failed agent run from its trace and replay.
- Fundamentals: agentic AI concepts and how UI agents perceive and act on a page.
Why join
- Work on agents that operate real, unpredictable websites in daily use, not curated demos.
- Learn from senior AI engineers building an agent platform from the ground up.
- See your work used directly by the business users who rely on it.