Job DescriptionThe problemEvery week, a tomato grower tells their buyers how many kilos they will deliver three weeks from now. Contracts, trucks and prices are set on that number. Source builds AI software that greenhouse growers in Europe and North America use for decisions like this one, and our Harvest Forecast for tomato, pepper and cucumber is one of the inputs to that weekly number.
Getting an average week right is not the hard part. The value is in the weeks when the crop does something unusual: a heat wave, a sharp drop after a peak, the start and the end of the season. Those are the weeks when growers need a forecast most, and when forecasts, ours included, are trusted least.
Why This Is Hard
- * Slow feedback. You can replay past seasons with any model. But a new model in production shows its first real result only three to four weeks later, and the weeks that matter come only a few times a year. You cannot A/B test a crop.
- Thin, noisy data. A few plants, measured by hand once or twice a week, stand in for a whole greenhouse. Sensors drift. Every site, variety and grower strategy is a little different.
- Not all of it is biology. Part of every swing comes from the plant, and part from decisions the grower makes after the forecast was made. Telling the two apart is part of the job.
- Strong science, messy greenhouses. Decades of crop and climate science exist. Data alone will not rediscover it, and the science alone does not fit a real greenhouse.
- People act on the number. A forecast that is confidently wrong costs a grower real money, so the uncertainty has to be honest.
We are looking for a Staff Data Scientist who has done this before: taken a forecast of a physical system that people did not trust, and made it a number a business commits money on.
The difficulty is not one clever model. It is the whole system: data quality, how plants are sampled, growers changing their harvest plans, and then the model. There may be ten things to fix before the forecast is good, and only some of them are in the model. You look at the whole system end to end, find the places that matter, do the hardest work yourself, and carry it all the way to something growers use.
This is a senior individual contributor role, technical all the way, with no people management. You report to our Director of Engineering and join our staff engineers, the senior group that works across teams with the CTO and sets the technical bar. Day to day you work with our data scientists, crop scientists and directly with growers. Harvest forecasting is the first problem. The same kind of problem shows up across our other crop models, and we expect your work to reach there too.
The role is based in Amsterdam, because you cannot see the whole system from a distance. You need to be in the room with the data scientists, the crop scientists and the people who talk to growers every week.
What You'll Do
- * Own the accuracy of our harvest forecasts for tomato, pepper and cucumber, especially in the unusual weeks, from diagnosis to a fix growers use
- Decide which problem is worth solving first, including when the fix is in the input data, the plant sampling or the grower's plan rather than the model
- Design and build hybrid models that combine crop physiology and greenhouse climate with statistical and machine learning methods
- Build the evaluation backbone: backtests on past seasons, strong baselines, uncertainty calibration and production monitoring, so the next ten fixes get cheaper
- Ship your models to production with our engineers, and stay responsible for how they behave there
- Raise the bar of the data science team through review, challenge and mentoring
Job RequirementsWe care most about what you have done. Today you might be called a data scientist, forecasting scientist, ML scientist, quant or research engineer. That does not matter to us.
- Around 10 years of hands-on modelling or forecasting, with models of physical or real-world systems that ran in production for more than one season. Energy, weather, process industry, plant breeding, pharma, batteries or trading: the field matters less than the fact that it was real
- You have taken a forecast people did not trust and made it one a business commits money on
- You have shipped models that broke when the world changed (a new season, a new site, a new customer) and fixed them, more than once. You can tell us what broke each time
- Deep expertise in time series and probabilistic forecasting: Bayesian and hierarchical models, gradient boosting, deep learning for sequences, and uncertainty quantification on small, irregular datasets
- Validation that holds up: backtests against strong baselines, error analysis tied to business impact, drift monitoring, and evaluation tools other people kept using after you built them
- Production-grade Python (pandas or polars, scikit-learn, PyTorch or JAX, PyMC or similar) and solid engineering habits (Git, testing, CI/CD, Databricks, AWS), so your models run in production without you
- An MSc or PhD background and the habit of reasoning from first principles
- You work AI-native: coding agents and LLM tools are how you work every day, and you know where they fail
- You have worked next to people with a higher standard than yours, and can say what they taught you
- Clear communication with growers, product and engineering alike
Nice to have
- * A degree in engineering, physics, applied mathematics, statistics or econometrics
- Physics-informed ML, mechanistic or process-based models, data assimilation (for example Kalman filtering) or system identification
- Publications, talks or open-source work on forecasting or scientific ML
You do not need a background in plant science or horticulture. We teach the domain.
Is this for you?
You Will Probably Like It Here If You
- * Want a hard problem that is still open, not a plan to execute
- Would rather find the real cause than tune a metric
- Hold a view under pressure and drop it when the evidence says so
- Want to meet the people who use your work and hear it directly when it is wrong
You Will Probably Not Like It Here If You
- * Need a clear spec, an A/B test or a large, clean dataset to know you are right
- Judge your work by an offline benchmark score
- Want to lead a team rather than solve the problem yourself
- Want to work on the model only, and leave the data, the measurements, the growers and production to someone else
What We Offer
- * Hybrid work: we work together in the Amsterdam office on Mondays and Thursdays
- Meals: lunch provided on office days
- Pension: a contribution of 4.5%
- Well-being: mental health support through OpenUp
- Equipment: MacBook Pro 16"
- Commuting: travel allowance for your commute to the office
- Learning: an annual learning budget of €1,000
- Home setup: a work-from-home budget of €550
- Connectivity: €50 per month for WiFi and phone
- Time off: flexible holiday policy; we encourage at least 25 days per year
- Community: quarterly company events, including dinner and team activities
How We Hire
- * Intro call with the hiring manager or our People & Culture team
- One story: a time a hard project changed direction because of what you did. We go deep until we understand what you did yourself
- Workshop on a real problem. We send you a write-up of how a greenhouse works, how plants are measured and where our forecast struggles, then work through it together, with your own tools including AI. Agreeing with our diagnosis earns nothing. A good next move does
- Final conversation and references
At Source, we value diversity of background and perspective. People assess their own qualifications differently, so please apply even if you do not tick every box. We are looking for the right person, not a perfect CV.
- Note to recruitment agencies: we do not accept unsolicited outreach about our open roles. Please do not contact us about this vacancy.
- Note on AI: during your interviews we use an AI note-taker, so we can be fully present with you rather than typing. There is no video recording. If you'd rather we didn't, just let us know during the interview.