COMPANY OVERVIEW
At ARRAY, we're not just another software services company—we're a team of dreamers, innovators, and trailblazers! From startup grit to big-tech aspirations, we're on a mission to redefine technology, put Bahrain on the global tech map, and grow into a powerhouse that inspires. If you're ready to be part of an exciting journey, we want you on our team!
KEY RESPONSIBILITIES
- Data Modelling: Design schemas and semantic layers (medallion bronze/silver/gold) that power AI-driven applications and analytics.
- ETL & Pipeline Architecture: Build and orchestrate ETL/ELT pipelines (batch and streaming) using Airflow and dbt, feeding cloud data warehouses and lakes.
- Lakehouse & Storage Architecture: Design and maintain data lake and lakehouse solutions using open table formats (Apache Iceberg, Delta Lake, or Hudi) over object storage, balancing cost, performance, and query flexibility.
- Database Engineering: Manage relational and NoSQL databases (PostgreSQL, MySQL, Oracle) supporting both transactional and analytical workloads, including performance tuning, replication, and migration between systems.
- Data Governance: Implement role-based access control, permission-aware retrieval, reconciliation, and data quality monitoring across data platforms.
- Cloud & Platform Integration: Deploy and operate data workloads on our Kubernetes-native platform (ArgoCD, Terraform) alongside core AWS services.
MUST-HAVE SKILLS
- Bachelor's degree in Computer Science or a STEM-based subject.
- 5+ years of experience in data or software engineering.
- Strong SQL and Python (or Java) skills.
- Hands-on experience building ETL/ELT pipelines using orchestration tools such as Airflow.
- Experience with cloud data warehouses and lakehouse architectures (e.g., Snowflake, Redshift, BigQuery, Apache Iceberg, or Delta Lake).
- Strong database fundamentals — schema design, indexing, query optimization — across relational (PostgreSQL, MySQL, Oracle) and NoSQL systems.
- Experience with unstructured data stores (Vector DBs, Document DBs) supporting AI/ML retrieval use cases.
- Working knowledge of containerized environments (Docker, Kubernetes) for deploying data workloads.
- CI/CD pipeline experience (GitHub Actions, GitLab CI, Jenkins, or similar).
- Working knowledge of Infrastructure as Code (Terraform or similar).
- Proficiency with version control (Git).
- Strong written and verbal English communication skills.
NICE-TO-HAVE SKILLS
- Apache Spark for large-scale batch and distributed data processing.
- Exposure to streaming/CDC tools (Kafka, Debezium, or Kinesis).
- Experience with data catalog and metadata management tools (Glue Data Catalog, DataHub, or Amundsen).
- Familiarity with Power BI dashboarding and reporting.
- AWS, Azure, or GCP certification.
- Experience migrating legacy on-prem databases to cloud-native platforms.
- Observability stack — Prometheus, Grafana.