Job DescriptionRole Purpose:
To ensure the organization has reliable, well-governed, and analytics-ready data by building and maintaining robust data pipelines and models. The role exists to make trusted data available securely and cost-effectively to support enterprise reporting, BI, and data-driven decision-making.
Key Responsibilities:
1- Enterprise Data Ingestion and Data Engineering;
- Build reusable, parameterized mappings and task flows in Informatica IDMC to standardize data ingestion.
- Implement change data capture (CDC), idempotent loads, schema evolution, and data-quality gates.
- Optimize Big Query loads using partitioning, clustering, and the appropriate load-versus-stream approach.
- Set up version control and CI/CD pipelines for data engineering assets.
2- Data Mapping & Transformation Design;
- Profile source systems and define field-level mappings and transformation rules.
- Specify business logic, including joins, lookups, and derivations.
- Define data-quality rules, exception handling, and reject criteria.
3- Orchestration, Automation & Reliability Engineering;
- Parameterize task flows and configure schedules and dependencies.
- Implement retries, backoff, and checkpointing to ensure reliable processing.
- Integrate monitoring and alerting through the Ops console and ChatOps.
4- Data Roles and Privacy management;
- Define least-privilege IAM roles and service accounts for data access.
- Apply dataset, table, row, and column-level security and data masking.
- Enable audit logging and retention policies, and classify PHI/PII data.
5- Modern Data Architecture & Data Modelling;
- Implement Medallion Architecture across the Bronze, Silver, and Gold layers.
- Develop scalable, curated data models for analytical and reporting needs.
- Define data quality, lineage, and governance standards for curated data.
- Collaborate with business and analytics teams to create trusted datasets.
6- Workload Monitoring, Performance & Cost Optimisation;
- Track data freshness, job health, volumes, and anomalies.
- Monitor SLAs across job duration, errors, cost per TB, and slot usage.
- Tune BigQuery performance through query optimisation and resource management.
- Optimize cost through storage lifecycle management, query tuning, and caching.
Skills
- * Data Engineering Tools: Informatica IDMC, Airflow, Apache Spark.
- Programming Languages: SQL, Python.
- Cloud Technologies: Google Cloud Platform (BigQuery, IAM), Azure/AWS (optional).
- Visualization & BI Platforms: Power BI, Looker, Looker Studio, Tableau.
- Strong in BigQuery (SQL, modelling, tuning).
- Proficient with Informatica IDMC/IICS ETL/ELT.
- Experience with end-to-end enterprise data pipelines using GCP and Informatica IDMC.
- Knowledge of Medallion Architecture.
- Experience building reporting-ready data marts and BI datasets.
- Experience supporting enterprise BI dashboards and self-service analytics is preferred.
- Knowledge of healthcare data (claims, providers, members).
- Understanding of cloud security (IAM, service accounts, private endpoints).
Education
Bachelor’s degree in Computer Science, Information Technology, Information Systems, or a related