Objective
Your role as Senior Data Engineer is to lead data integration and deliver scalable, reliable pipelines across all companies under the Holding. You will provide trusted data for reporting, analytics, machine learning and generative AI, drive architecture improvements, mentor engineers, and ensure strong data modeling and governance. The role owns the data foundation and works with AI engineers and data scientists on model and application requirements.
Main Tasks
- Lead the design and implementation of scalable batch and real-time ETL/ELT pipelines; use Apache Airflow to orchestrate scheduled workflows.
- Integrate databases, APIs, files, event streams and documents; manage incremental loads, change data capture and schema changes.
- Ensure data integrity and platform scalability; work with DBAs on database performance and availability.
- Maintain data models, source mappings, lineage, technical documentation and automated pipeline tests in version control.
- Collaborate with analysts, AI engineers and data scientists to deliver curated datasets and consistent business definitions.
- Monitor data quality, freshness and pipeline failures; implement alerts, retries, replay and root-cause fixes.
- Optimize queries, transformations, storage and processing costs for high-volume workloads.
- Develop reusable connectors and automation scripts. Use AI-assisted development tools to support coding, testing, debugging, and documentation while reviewing generated outputs for accuracy, security, quality, and maintainability.
- Evaluate data technologies and architectures against business needs, security, maintainability and cost.
- Mentor junior and mid-level engineers; review code and promote engineering and delivery standards.
- Build reproducible datasets and feature pipelines for model training, evaluation and inference with data scientists.
- Build data pipelines for retrieval-augmented generation (RAG): document extraction, chunking, metadata, embeddings and index refresh or deletion.
- Provide governed data access for AI applications and agents, enforcing source permissions, sensitive-data controls and traceability.
Qualifications
- Bachelor’s degree in computer science, Computer Engineering, or related field.
- 4+ years of experience in data engineering/integration, including production pipelines, data warehousing, and SQL. Hands-on Apache Airflow experience, with exposure to analytics/ML data delivery; RAG or vector search experience is a plus.
- Professional proficiency in Local Language and English required.
- Strong SQL and Python for integration, transformation, automation and troubleshooting.
- Production experience with relational and analytical databases, such as PostgreSQL, SQL Server or ClickHouse; familiarity with NoSQL.
- Batch and streaming integration using Kafka or equivalent, including deduplication, retries and incremental processing.
- Strong Apache Airflow skills: DAG development, scheduling, dependencies, retries, monitoring and troubleshooting; familiarity with dbt and automated data quality checks.
- Data modeling for warehouses and lakes, including dimensional models, partitioning and formats such as Parquet.
- Git, code review, CI/CD, Linux and container fundamentals; experience with cloud or on-premises data platforms.
- Understanding of ML dataset preparation, reproducibility, data leakage, and data requirements for training and inference.
- Excellent analytical, communication and mentoring skills; ownership of production issues and delivery quality.
- Practical RAG ingestion, embeddings, vector databases or vector search, and metadata-based access filtering.
- Experience with Spark or Flink, object storage, lakehouse architectures, or cloud data services on AWS, Azure or GCP.
- Experience with telecom datasets, CDRs, network events or TM Forum SID; familiarity with feature stores and data versioning.