Staff Data Engineer, Robotics
Germany | Full-time, on-site | Competitive six-figure base
We're working with a fast-growing deep tech company in Germany that is developing AI for machines operating in the physical world. As their models scale, so does the volume and complexity of their real-world data, and they're hiring a senior engineer to take ownership of how that data is collected, processed and turned into better models.
This is a hands-on, high-impact role sitting between software and machine learning, with genuine technical authority and a direct line to senior leadership.
The role
- Design and run large-scale distributed processing for video, sensor and time-series data
- Build the infrastructure that gets data efficiently from collection into large GPU training runs
- Work closely with ML researchers to understand where models fall short, and fix it through better data
- Define how multimodal data is structured, stored, versioned and quality-checked
- Develop Python tooling for validation, cleaning and dataset management
- Help shape data collection, annotation processes, and secure handling of sensitive data
What you'll brin
- Significant experience in data engineering, ideally with large volumes of real-world data from robotics, autonomous driving, drones or video AI
- Strong hands-on experience with distributed processing frameworks such as Ray, Spark, Dask or Beam
- Expert, production-grade Python, including asyncio, multiprocessing and performance optimisation
- Strong working knowledge of PyTorch and NumPy, and of how training workloads consume data
- A proven track record building pipelines that feed multi-node GPU training on AWS or GCP
- Experience with high-throughput storage and data loading for training (e.g. S3/GCS, Parquet/Arrow, WebDataset or similar)
- Hands-on experience with video encoding (H.264/H.265, FFmpeg) and robotics or sensor log formats (e.g. ROS bags, MCAP or similar)
- Dataset versioning and lineage at scale using tools such as DVC, LakeFS, Delta Lake or Iceberg
- A solid grasp of machine learning fundamentals, enough to judge which data actually improves a model
Nice to have
- Building training datasets for multimodal or foundation models (e.g. VLA, VLM)
- Time synchronisation across high-frequency sensor streams and 3D coordinate transforms
- Active learning, data-centric ML or failure-case mining workflows
- Workflow orchestration (Airflow, Dagster, Prefect) and containerised infrastructure (Docker, Kubernetes)
- Experience working with annotation vendors or data collection operations
Why it's worth a conversation
- A senior hands-on role with real ownership and influence
- Work on cutting-edge AI with immediate, visible impact
- A strong engineering culture in a company with serious backing