hirq
← All jobs

cloudbox

Senior Data & AI Engineer

Cape Town, Western Cape · On-site · Full-Time · Tech

Apply well, not just fast

Create a free account and upload your resume to get a match score, keyword gaps, a tailored resume, a cover letter and interview prep for this job.

About the role

AzureDatabricksCI/CDGitObservabilityLLMsRAGPower BI
About the role Cloudbox is building a unified multi-tenant data and reporting platform on Microsoft Fabric, aggregating 10+ vendor APIs (Autotask PSA, Microsoft 365/Graph, Azure Monitor, Microsoft Defender, Huntress, Proofpoint, Cato SASE, BambooHR, Xero, and internal systems) into a medallion lakehouse for hedge fund and financial-services clients. You'll own the data platform end-to-end, from raw vendor ingestion through to a query-ready gold layer and an AI-assisted reporting layer. This is a greenfield build: you will design the medallion architecture, build it, and then run it as its long-term owner. What you'll own - Metadata-driven OneLake medallion lakehouse architecture (bronze/silver/gold) - Azure Functions-based vendor ingestion pipelines across the full vendor set, fanned out per tenant from a central metadata registry, starting with an Autotask-first vertical slice - Dimensional modelling and semantic model design - the semantic layer serving dashboards and reporting, with row-level security (RLS) for client-level access - The AI proof-of-concept - grounded, provenance-backed natural-language querying over platform data, with anti-fabrication controls - Data quality, reconciliation, and observability for the pipeline layer - structured logging, SLOs, and alerting, with reconciliation proven against synthetic vendor data before production - Data governance with Microsoft Purview - lineage, classification, and sensitivity labelling across the medallion layers - Documentation and knowledge transfer - runbooks, architecture decision records, and pairing with the wider team, so the platform never carries a bus factor of one What we're looking for - Senior-level experience with Microsoft Fabric, OneLake, or equivalent modern lakehouse platforms (Databricks, Synapse acceptable if willing to specialise into Fabric) - Strong dimensional modelling and semantic model design background (star schemas, conformed dimensions, Fabric/Power BI semantic models) - Experience building production data pipelines from third-party/vendor APIs - ideally within a metadata-driven ingestion framework - Hands-on with the Azure services surrounding Fabric - Function Apps, Key Vault, Blob Storage, and Static Web Apps - Source control and CI/CD experience with both Azure DevOps and GitHub, including Fabric Git integration; Bicep or equivalent infrastructure-as-code a plus - Security-first engineering habits - managed identities and Key Vault for all secrets; no credentials in code, ever - Working knowledge of Microsoft Purview (or equivalent) for governance, lineage, and sensitivity labelling - this platform serves regulated financial-services clients - Familiarity with grounding techniques for LLM-based querying over structured data (RAG, citation/provenance patterns) is a plus - the core screen for this role is lakehouse and modelling depth, and the AI layer can be developed in-role with specialist support - Comfortable being the technical anchor for the platform's most critical dependency chain Nice to have: Exposure to Azure AI Foundry (model deployment, evaluation, prompt flow), Fabric IQ and its ontology layer, Fabric data agents, MCP-style agent tool-calling, Azure Key Vault / Managed HSM and customer-managed key (CMK) patterns, and Fabric capacity planning (F-SKU sizing) - these support later platform tiers and none is a screening requirement