hirq
← All jobs

Saudi AZM

AI Engineer

Riyadh, Riyadh Province, Saudi Arabia · On-site · fulltime_permanent · AZM ICT - PMO

Apply well, not just fast

Create a free account and upload your resume to get a match score, keyword gaps, a tailored resume, a cover letter and interview prep for this job.

About the role

AzureFastAPIGCPKubernetesCI/CDObservabilityLLMsRAGMLOpsIncident ResponseTechnical Writing
AI systems - Build, fine-tune, and evaluate LLM systems for domain-specific tasks (QLoRA / PEFT on open-weight models such as Llama-3 and Mistral). - Design reproducible evaluation harnesses and A/B test frameworks with tracked metrics: task success rate, safety rate, and latency distributions (p50/p95). - Architect multi-agent and RAG systems (LangGraph, FastAPI, vector databases) from prototype through production. - Implement safety guardrails — input/output validation, allowlist/denylist policies, and controls that reduce invalid or high-risk model actions. - Translate business use cases into deployable prototypes with measurable acceptance criteria, and demo them to stakeholders. Platform & infrastructure - Design and operate cloud infrastructure and MLOps workspaces (Azure, OCI, or GCP) for AI workloads on Kubernetes and containerized runtimes. - Build CI/CD pipelines and GitOps-based release promotion (Argo CD) across development, test, and production environments. - Implement end-to-end observability (Azure Monitor, Application Insights, ELK) with defined detection and response targets. - Apply network and perimeter security baselines (FW/WAF), automated code quality and SCA scanning (SonarQube, Black Duck), and gated pipelines. - Own disaster recovery design — automated backups, failover, and documented RTO/RPO commitments. Engineering leadership - Lead and mentor a cloud/AI operations team; define monitoring, incident response, and release governance practices with clear uptime and MTTR targets. - Standardize SDLC practices — branching strategy, PR governance, release management, delivery reporting — to improve lead time and deployment frequency. - Consolidate engineering tooling and workflows; drive migrations and platform standardization where fragmentation slows delivery. - Produce handover documentation and runbooks that make systems auditable and operationally transferable. - Support vendor and licensing negotiations for cloud enterprise agreements.