Senior technical leadership role requiring strong hands-on experience across modern data and AI platforms, DevOps/DataOps, production AI, observability and Service Management integration.
Key Success Measures- Data & AI ProductOps practice established, adopted and scaled.
- Clear and reusable standards established across DataOps, MLOps, LLMOps and AgentOps.
- ProductOps capabilities successfully implemented within clients’ environments.
- Clear production ownership, operational readiness and L0/L1/L2/L3 support models established.
- Improved product health and SLO compliance.
- Reduced incident detection and restoration time.
- Reduced recurring production issues and unnecessary engineering escalations.
- Increased automation, proactive monitoring and self-service.
- Strong integration between Data & AI ProductOps and Enterprise Service Management.
- Increased reuse of standards, patterns, playbooks and delivery accelerators.
- Strong internal capability development and effective client knowledge transfer.
Responsibilities
Practice Development
- Establish and lead the Data & AI ProductOps practice, including the operating model, standards, methods, governance, reusable patterns, playbooks and implementation approach.
- Define the practice across DataOps, MLOps, LLMOps and AgentOps, with clear standards for operating Data, BI, ML, GenAI and Agent-based products in production.
- Develop reusable ProductOps assets, including operational readiness standards, support models, SLO frameworks, monitoring patterns, runbooks, implementation templates and delivery accelerators.
Product Operations
- Define and implement operational readiness requirements covering product ownership, criticality, support levels, SLAs/SLOs, monitoring, alerting, runbooks, escalation, recovery, dependencies and rollback.
- Establish clear L0/L1/L2/L3 support models, with L0 focused on automation and self-service, L1 on Service Desk support and initial triage, L2 on ProductOps-led operational support and product-level troubleshooting, and L3 on complex issues requiring Engineering expertise.
- Establish product health and observability across availability, performance, data freshness and quality, pipeline and integration health, model and AI performance, usage, cost and other product-specific operational measures.
- Lead product-level incident investigation, restoration and root-cause analysis, coordinating across Engineering, Platform, Service Management, Security and vendors.
- Integrate ProductOps with enterprise Incident, Problem, Request, Change, Release, Knowledge, Service Level and Major Incident Management processes.
- Ensure recurring incidents and operational issues are converted into permanent corrective actions, automation opportunities and product improvement backlog items.
- Establish effective change and release practices supported by automated testing, CI/CD, versioning, controlled deployment, and rollback.
- Establish regular Product Operations reviews covering product health, SLOs, incidents, recurring problems, operational trends, technical debt, improvements and releases.
Client Delivery
- Assess clients Data & AI ProductOps maturity and define target operating models, support structures, service levels, observability, automation and implementation roadmaps.
- Design and implement ProductOps capabilities for Data, BI, ML, GenAI and Agent-based products within clients’ environments.
- Lead clients solutioning, architecture workshops and implementation engagements, defining roles, responsibilities, support models, operational processes, tooling and Service Management integration.
- Provide architecture and implementation oversight for complex Data and AI production environments, engaging senior stakeholders across Product, Engineering, Platform, Governance, Security and Service Management.
- Support proposals, technical solutioning and advisory engagements across Data & AI ProductOps, DataOps, MLOps, LLMOps and AgentOps.
Leadership Expectations
- Build and scale a new Data & AI ProductOps practice.
- Provide strong technical leadership while remaining hands-on where required.
- Lead multidisciplinary teams across Data, AI, Engineering, Platform, Product and Service Management.
- Provide architecture and implementation oversight across complex enterprise environments.
- Engage confidently with senior clients and internal stakeholders.
- Translate emerging Data and AI technologies into practical enterprise operating practices.
- Develop internal talent and build sustainable technical capability.
Qualifications
- Bachelor’s degree in Computer Science, Engineering, Data/AI, Information Systems, or a related field.
- Relevant certifications across cloud, Data & AI platforms, DevOps/MLOps, Service Management, architecture, or data management are an advantage.
- 10+ years of experience across Data Engineering, AI/ML, DataOps, DevOps, SRE, Product Operations, Platform Engineering or related disciplines.
- Strong hands-on experience with modern enterprise data platforms and production Data/AI solutions.
- Demonstrated experience establishing or leading production operating capabilities for Data and AI products.
- Strong practical experience across DataOps, MLOps, LLMOps and AgentOps.
- Strong experience with operational readiness, observability, SLAs/SLOs, monitoring, support models, incident and problem management, CI/CD, automated testing, release management, recovery and rollback.
- Experience operating BI and data products, data pipelines, APIs, integrations, ML models, GenAI applications and AI Agents in production.
- Practical experience with model lifecycle management, model serving, drift monitoring, RAG, vector search, AI evaluation, AI observability, Agent tracing, tool execution and guardrails.
- Strong understanding of enterprise ITSM and Service Management, with experience integrating Data and AI product teams into established support, incident, problem, change and release processes.
- Strong incident leadership, troubleshooting, root-cause analysis and production problem-solving capabilities.
- Demonstrated experience developing technical standards, operating models, reusable patterns and implementation methods.
- Strong client-facing consulting, architecture, solutioning and stakeholder-management capabilities.
- Strong experience working with government agencies or large enterprises in Qatar or the GCC is highly preferred.
Preferred Technology Experience
Experience across a combination of the following technology areas:
- Data & Analytics Platforms: Databricks, Informatica IDMC, Microsoft Fabric, Power BI, Azure Data Services, Spark and Delta Lake.
- Data Integration & Streaming: Informatica, Kafka, APIs, batch integration, streaming and CDC patterns.
- AI, ML & GenAI: MLflow, Azure AI/ML, model serving and monitoring, RAG, vector search, AI evaluation and AI observability.
- GenAI & Agent Frameworks: LangChain, LangGraph, Semantic Kernel or equivalent GenAI and Agent orchestration frameworks.
- DevOps & Platform Engineering: Azure DevOps, GitHub, CI/CD, Docker, Kubernetes, Terraform, Python and SQL.
- Observability & Monitoring: OpenTelemetry, Azure Monitor, Application Insights or equivalent enterprise monitoring platforms.
- Service Management: ServiceNow or equivalent enterprise ITSM platforms.