About Skeptical
Most AI in production automates very little, whether the model is a custom classifier or a frontier LLM. A model makes a suggestion, a person checks it, and the queue is as long as it was before, because nobody can state when and how often the prediction is wrong. Skeptical addresses this. A team sets the error rate it is willing to accept. Skeptical's software then clears every decision that falls within it, automatically and in real time, and routes the remainder to human review. The error rate among cleared decisions, measured against the reviewers' decisions, stays at or below the chosen level, with a statistical guarantee. The approach requires no retraining and no changes to the customer's pipeline.
We are a team out of KTH Royal Institute of Technology, and we are growing fast. The method comes from our own peer-reviewed, award-winning research. The hard part is making it hold in production, on real data, on live traffic, for years.
The role
Help build the product and deploy it with our first customers. The role is split between infrastructure and time spent with design partners, understanding how their review queues work and getting them cleared.
What you will do
You will help build the core of the product: a service that takes a customer's AI predictions and error budget and returns, for each prediction, whether it clears or goes to review. Around it, the data pipelines that bring scores and review outcomes in, the recalibration and drift monitoring that keep the guarantee valid over time, and the interfaces customers work in.
You will then deploy it inside design partner environments, in domains such as insurance claims, credit decisions, and aviation safety, and stay close while they use it. You will take ownership of the deployment and resolve the problems that appear in their systems together with the customer, and what you learn there feeds back into the product. You will work in a small team, with real ownership of what you build and how. Our aim is to build something we are proud of, and that requires people who want to be responsible for their part of it.
What we are looking for
You have run a production system that other people depended on, and you know what that requires. You have deployed machine learning models in practice, whether classifiers and scoring pipelines or model-driven workflows with several steps, and you know what it takes to keep them running. You can read a paper on uncertainty quantification and turn it into working code, and you are comfortable deciding what to do next without being told.
You are fluent in Python and SQL, and you have shipped production code in at least one statically typed language. You are at home with containers and at least one cloud, and you have built data pipelines that other services depended on.
If you have wondered why models keep getting smarter while the pipelines that matter still require a person to check every output, you are probably the person we are looking for. Our answer is that AI has gotten smarter, but it has not become more trustworthy. We are building the layer that changes that, and we would like to work with people who think the question was worth asking.
What we offer
A founding title and the equity that comes with it. Ownership of a whole product surface rather than a component of one. A seat in every customer conversation that matters, and a real say in what the product becomes. Founders who have built production systems that the industry runs on, and who will work alongside you. A desk in central Stockholm and colleagues in the same room most days.
How to apply
Tell us what you would do here and why you are the right person to help us solve this problem. Apply at skeptical.io/careers