Role Overview This is an amazing opportunity for the right person. You will lead the data engineering for a petabyte scale industrial digital twin, at the center of a national initiative delivering the world's largest coral restoration programme, working with imagery, geospatial data and sensor telemetry that few data engineers see in one place. Beyond conventional pipeline work you will build the curated datasets that feed production AI and MLOps pipelines, and own the data governance framework from first principles, in a small senior team where your decisions reach production quickly. This is a hands-on delivery-first role combined with advisory responsibility pertaining to data. You will work with the Databricks Lakehouse (bronze/silver/gold), build and run Azure Data Factory orchestration, manage the real-time ingestion chain from Event Hubs through Function Apps and Batch Accounts, and migrate operational data from source systems. You will also own the continued functional enhancements and optimization of the live application data architecture. This can include schema changes as well as performance tuning. Alongside deep hands-on engineering, you will be the person the team and the client look to for direction on data architecture and data governance: shaping the digital twin platform medallion and application architecture, and standing up and operating the governance framework the engagement commits to the data domain map, ownership register, data quality rules and issue routing, classification-driven access, metadata and lifecycle.
Key Responsibilities Databricks Lakehouse & Pipelines• Build and maintain bronze/silver/gold pipelines in Databricks, processing CSV and other source formats into curated, analysis-ready datasets.• Develop Databricks notebooks/jobs (PySpark/SQL), manage job scheduling for recurring and incremental loads, and tune performance.• Register curated datasets in Unity Catalog with correct metadata, ownership and access tags, per the architecture team's cataloguing standards.• Apply data modeling and data warehousing best practices (dimensional modeling, star schema, slowly changing dimensions) when designing silver/gold layer schemas.• Use Databricks Lakebase — the managed, Postgres-compatible operational database built on the Lakehouse — where a workload needs transactional/OLTP-style access alongside the Delta Lake data. Orchestration & Ingestion• Build and maintain Azure Data Factory pipelines and triggers that move data from source systems into Databricks — this movement is trigger-driven, not automatic.• Consume streaming data from Azure Event Hubs across multiple concurrent sources, and configure triggers so Databricks pulls new/changed data as it arrives.• Develop Azure Function Apps that fire on incoming Event Hub data and initiate downstream processing via Batch Accounts.• Manage Azure Batch Accounts that receive data from third-party tools, IoT/sensor feeds and other external sources, and move it into Databricks once triggered.• Keep the ingestion chain (Event Hub → Function App → Batch Account → Databricks) resilient, monitored, and recoverable from failure. Migration, Security & Storage• Migrate data from the different sources/platforms into Databricks, preserving data integrity and minimizing disruption to source systems.• Support migration of large imagery/scientific-archive data into the platform, including metadata preservation and integrity validation.• Manage credentials and secrets in Azure Key Vault, and implement access controls (RBAC, managed identities) per architecture guidance.• Maintain Azure Data Lake as the platform's storage layer, including file and metadata organisation across Databricks and Batch Account outputs. Downstream Enablement• Prepare curated, governed datasets that feed Power BI/Tableau reporting and Digital Twin applications.• Prepare feature-ready datasets for AI/ML consumption, and support MLflow-tracked training and inference data needs.• Build and maintain the feature store and feature-engineering pipelines that support the programme's ML cadence, working with the ML Lead so models consume governed, versioned datasets rather than ad-hoc extracts.• Provide the data foundation for MLOps — reproducible training and inference datasets, dataset and feature versioning, lineage back to the gold layer, and drift/freshness monitoring that supports model retraining.• Build the data pipelines behind Gen AI and RAG use cases: preparing and chunking documents and scientific/operational text, generating and maintaining embeddings, managing vector indexes, and keeping retrieval sources refreshed and access-controlled so retrieval respects the same classification rules as the rest of the platform.• Enable advanced analytics capabilities on the platform, including graph and knowledge-graph data structures that support simulation, prediction and decision support alongside the Digital Twin.• Work with geospatial (Post GIS) and time-series/sensor data feeds as part of the ingestion and curation pipeline. Data Architecture• Own the Lakehouse data architecture in practice — medallion layering, domain boundaries, curated data products and serving patterns — documented in the programme's digital blueprint, and set the integration and data standards any contributor, internal or third-party, must meet before data enters the governed layers.• Own application level data architecture across all current and future applications.• Advise on the open target-state decisions and work with the Platform Lead and ML Lead so ingestion, curation, reporting and ML consumption are architected as one estate
Data Governance Leadership• Run the initial data assessment and define the governance framework from it — the data domain map, the ownership register with a named owner per domain, and the standards covering data quality, classification-driven access, privacy, metadata and lifecycle.• Implement that framework in the platform rather than on paper: data quality rules with live issue routing to the accountable domain owner, the QA/QC split with the business, and retention, lifecycle and deletion processes operating for governed data.• Own the implementation of Data Quality rules onto relevant applications. Direct application development team on what data quality controls need to be implemented in application and any future applications.• Bring each newly onboarded domain under the framework and the metadata standard as it lands, evidence that governance is actually operating,Core Technology Skills• Azure Databricks, Apache Spark/PySpark, Delta Lake, Unity Catalog, Medallion (bronze/silver/gold) architecture.• Databricks Lakebase (managed, serverless Postgres-compatible OLTP database sharing the Lakehouse's storage layer) for operational/transactional workloads.• Data modeling and data warehousing fundamentals — dimensional modeling, star schema, slowly changing dimensions — applied to curated Lakehouse layers.• Azure Data Factory, ADLS Gen2, pipeline orchestration, and batch + streaming integration patterns.• Azure Event Hubs, Azure Function Apps and Azure Batch Accounts for event-driven and trigger-based ingestion.• Advanced SQL and strong Python/PySpark for data transformation and pipeline development.• Postgre SQL, including Post GIS for geospatial data and exposure to time-series stores (e.g. Timescale DB) for sensor/IoT data.• Azure Key Vault, managed identities and RBAC for secure credential and access management.• Data quality and observability practices: validation, profiling, freshness/completeness checks, pipeline monitoring.• Experience preparing data for Power BI/Tableau consumption and for AI/ML pipelines (feature tables, MLflow-tracked datasets).• Git-based development and CI/CD for data pipelines (Azure Dev Ops or equivalent).• Working familiarity with geospatial/GIS data (Arc GIS-adjacent) and IoT/telemetry data formats.• Data governance frameworks in practice — ownership and stewardship models, data quality rule design, classification-driven access, retention and lifecycle, metadata standards and catalogue management (Unity Catalog, Purview or equivalent).• Lakehouse and data-product architecture — medallion design, domain boundaries, serving patterns, and the ability to document and defend architecture decisions to a client audience.• Data lineage, cataloguing and observability tooling, and the ability to evidence that controls are operating rather than merely defined.• Cost and performance optimisation of Azure data workloads (reserved compute/storage, cluster sizing, endpoint and logging optimisation).• AI/ML data engineering — feature stores, feature engineering at scale, training/inference dataset preparation, MLflow-tracked datasets, and the data side of MLOps (versioning, lineage, drift and freshness monitoring).• Gen AI and RAG data patterns — document and text preparation and chunking, embedding generation, vector stores/indexes (Databricks Vector Search, pgvector or equivalent), retrieval evaluation, and applying data classification and access control to retrieval sources.
نظرة عامة على الدور: هذه فرصة مذهلة للشخص المناسب. ستقود هندسة البيانات لتوأم رقمي صناعي على نطاق بيتابايت، في قلب مبادرة وطنية تقدم أكبر برنامج لاستعادة الشعاب المرجانية في العالم، حيث ستعمل مع الصور والبيانات الجغرافية المكانية وقياسات المستشعرات التي نادراً ما يراها مهندسو البيانات في مكان واحد. بالإضافة إلى عمل خطوط أنابيب البيانات التقليدية، ستقوم ببناء مجموعات بيانات منظمة تغذي خطوط أنابيب الذكاء الاصطناعي وعمليات تعلم الآلة (MLOps)، وستتولى مسؤولية إطار عمل حوكمة البيانات من المبادئ الأولى، ضمن فريق صغير من كبار الخبراء حيث تصل قراراتك إلى مرحلة التنفيذ بسرعة. هذا دور عملي يركز على التسليم المباشر بالإضافة إلى مسؤولية استشارية تتعلق بالبيانات. ستعمل باستخدام Databricks Lakehouse (الطبقات البرونزية/الفضية/الذهبية)، وستقوم ببناء وتشغيل تنسيق Azure Data Factory، وإدارة سلسلة استيعاب البيانات في الوقت الفعلي من مراكز الأحداث (Event Hubs) عبر تطبيقات الوظائف (Function Apps) وحسابات الدفعات (Batch Accounts)، وترحيل البيانات التشغيلية من الأنظمة المصدرية. كما ستتولى مسؤولية التحسينات الوظيفية المستمرة وتحسين بنية بيانات التطبيقات الحية، بما في ذلك تغييرات المخطط وضبط الأداء. بجانب الهندسة العملية العميقة، ستكون الشخص الذي يتطلع إليه الفريق والعميل للحصول على التوجيه بشأن بنية البيانات وحوكمتها: تشكيل بنية منصة التوأم الرقمي، وإنشاء وتشغيل إطار عمل الحوكمة الذي يلتزم بخريطة نطاق البيانات، وسجل الملكية، وقواعد جودة البيانات، وتوجيه المشكلات، والوصول القائم على التصنيف، والبيانات الوصفية، ودورة الحياة.
المسؤوليات الرئيسية: بحيرة بيانات Databricks وخطوط الأنابيب • بناء وصيانة خطوط الأنابيب (البرونزية/الفضية/الذهبية) في Databricks، ومعالجة ملفات CSV والتنسيقات المصدرية الأخرى إلى مجموعات بيانات منظمة وجاهزة للتحليل. • تطوير دفاتر ملاحظات/وظائف Databricks (PySpark/SQL)، وإدارة جدولة المهام للأحمال المتكررة والتزايدية، وضبط الأداء. • تسجيل مجموعات البيانات المنظمة في Unity Catalog مع البيانات الوصفية والملكية وعلامات الوصول الصحيحة، وفقاً لمعايير الفهرسة الخاصة بفريق الهندسة المعمارية. • تطبيق أفضل الممارسات لنمذجة البيانات وتخزينها (النمذجة الأبعاد، مخطط النجمة، الأبعاد المتغيرة ببطء) عند تصميم مخططات الطبقات الفضية/الذهبية. • استخدام Databricks Lakebase - قاعدة البيانات التشغيلية المُدارة والمتوافقة مع Postgres المبنية على Lakehouse - عند احتياج سير العمل إلى وصول معاملات/OLTP إلى جانب بيانات Delta Lake. التنسيق والاستيعاب • بناء وصيانة خطوط أنابيب Azure Data Factory والمشغلات التي تنقل البيانات من الأنظمة المصدرية إلى Databricks. • استهلاك البيانات المتدفقة من Azure Event Hubs عبر مصادر متزامنة متعددة، وتكوين المشغلات لتقوم Databricks بسحب البيانات الجديدة/المغيرة عند وصولها. • تطوير تطبيقات Azure Function التي تعمل عند وصول بيانات Event Hub وبدء المعالجة اللاحقة عبر حسابات الدفعات. • إدارة حسابات Azure Batch التي تتلقى بيانات من أدوات الطرف الثالث، وخلاصات IoT/المستشعرات ومصادر خارجية أخرى، ونقلها إلى Databricks بمجرد تفعيلها. • الحفاظ على سلسلة الاستيعاب مرنة ومراقبة وقابلة للاسترداد من الفشل. الترحيل والأمان والتخزين • ترحيل البيانات من مصادر/منصات مختلفة إلى Databricks، مع الحفاظ على سلامة البيانات وتقليل التعطيل للأنظمة المصدرية. • دعم ترحيل بيانات الصور الكبيرة/الأرشيفات العلمية إلى المنصة، بما في ذلك الحفاظ على البيانات الوصفية والتحقق من السلامة. • إدارة بيانات الاعتماد والأسرار في Azure Key Vault، وتطبيق عناصر التحكم في الوصول (RBAC، الهويات المُدارة) وفقاً لتوجيهات الهندسة المعمارية. • صيانة Azure Data Lake كطبقة تخزين للمنصة، بما في ذلك تنظيم الملفات والبيانات الوصفية عبر مخرجات Databricks وحساب الدفعات. تمكين المعالجة اللاحقة • إعداد مجموعات بيانات منظمة ومحكومة تغذي تقارير Power BI/Tableau وتطبيقات التوأم الرقمي. • إعداد مجموعات بيانات جاهزة للميزات لاستهلاك الذكاء الاصطناعي/تعلم الآلة، ودعم احتياجات بيانات التدريب والاستدلال التي يتتبعها MLflow. • بناء وصيانة مخزن الميزات وخطوط أنابيب هندسة الميزات التي تدعم وتيرة تعلم الآلة للبرنامج. • توفير أساس البيانات لـ MLOps - مجموعات بيانات التدريب والاستدلال القابلة للتكرار، وإصدارات مجموعات البيانات والميزات، ومراقبة الانجراف/الحداثة التي تدعم إعادة تدريب النماذج. • بناء خطوط أنابيب البيانات خلف حالات استخدام الذكاء الاصطناعي التوليدي (Gen AI) وRAG. الحوكمة • إجراء التقييم الأولي للبيانات وتحديد إطار عمل الحوكمة منه - خريطة نطاق البيانات، سجل الملكية، والمعايير التي تغطي جودة البيانات، الوصول القائم على التصنيف، الخصوصية، والبيانات الوصفية ودورة الحياة. • تنفيذ هذا الإطار في المنصة فعلياً: قواعد جودة البيانات مع توجيه المشكلات الحية إلى مالك النطاق المسؤول، وعمليات الاستبقاء ودورة الحياة والحذف للبيانات المحكومة. • امتلاك تنفيذ قواعد جودة البيانات على التطبيقات ذات الصلة. توجيه فريق تطوير التطبيقات بشأن عناصر التحكم في جودة البيانات التي يجب تنفيذها في التطبيق.