Engineering

Ai Platform Engineer At International Rescue Committee

International Rescue Committee·Nairobi, kenya·Full Time·Internship
EngineeringFull TimeInternship
-Closes Oct 16, 2026

Major Responsibilities

AI Systems Administration & Operations (40%)

  • Serve as primary technical administrator across IRC enterprise AI environments, currently including Anthropic (Claude) and OpenAI platform deployments
  • Manage user access, API key governance, workspace configurations, and environment-level settings across AI platforms
  • Monitor system health, usage patterns, and API performance across AI tools; triage and resolve operational issues as they arise
  • Maintain and improve observability across AI system

Tracking uptime, error rates, token consumption, and integration reliability

  • Oversee and document configuration changes, environment updates, and deployment procedures across managed platforms
  • Support responsible use by flagging anomalous usage patterns and coordinating with InfoSec on policy adherence and access controls

Integrations & Technical Implementation (35%)

  • Coordinate with the DevOps, SW Engineering and Data Engineering team(s) on deployment processes, environment access, and infrastructure dependencies required to build and maintain AI integrations
  • Follow established change management procedures for all configuration changes, environment updates, and integration deployments, including documentation, testing, and appropriate approvals before pushing to production
  • Develop lightweight scripts, connectors, and automations to support AI-assisted workflows across teams, primarily in Python and/or JavaScript/TypeScript
  • Troubleshoot integration failures, data flow issues, and API connectivity problems across the AI ecosystem
  • Collaborate with the data engineering team on AI/KM pipeline work, including vector store ingestion, retrieval configuration, and source data connections
  • Contribute to technical design discussions with engineering partners, translating operational requirements into implementable solutions
  • Maintain technical documentation for all integrations, including architecture notes, runbooks, and dependency maps

Monitoring, Resource Optimization & InfoSec Liaison (15%)

  • Track and report on AI resource utilization across platforms, identifying opportunities to reduce waste and improve cost efficiency in coordination with the AI
  • Serve as the technical point of contact with the InfoSec team on matters related to AI system security, data handling, access controls, and compliance requirements
  • Support risk assessments and security reviews for new AI tools or integrations by providing accurate technical context on system behavior and data flows
  • Contribute to the development of technical SOPs and best-practice guidelines for AI system use, in coordination with the AI Platform Support Director and relevant stakeholders

Stakeholder Support & Collaboration (10%)

  • Act as a technical resource for program and operations teams adopting AI tools, including answering implementation questions, supporting troubleshooting, and identifying configuration solutions
  • Participate in rollout planning for new AI capabilities, providing grounded input on technical feasibility, integration requirements, and operational readiness
  • Collaborate with the AI Platform Support Director on onboarding documentation and technical guidance materials for end users
  • Contribute to sprint and project planning with accurate estimates on technical effort and dependencies

Required Experience & Skills

AI & Cloud Platforms

  • Hands-on experience administering enterprise AI platforms (Anthropic, OpenAI, Azure OpenAI, or comparable tools), including API management, access controls, and environment configuration
  • Familiarity with LLM application infrastructure: prompt pipelines, Model Context Protocol (MCP), other tool-calling integration frameworks, vector databases, retrieval-augmented generation (RAG) patterns, and embedding workflows
  • Experience working with Databricks or comparable data/ML platforms is a strong plus

Integration & Development

  • Proficiency in Python and/or JavaScript for scripting, automation, and lightweight integration work
  • Experience building and maintaining REST API integrations, including authentication patterns, webhook handling, and error management
  • Comfort reading and working within existing codebases without requiring significant architectural guidance
  • Familiarity with version control (Git) and standard deployment practices for scripts and integrations

Systems Administration & Monitoring

  • Experience monitoring distributed systems or SaaS platforms, including setting up alerting, reviewing logs, and diagnosing performance or availability issues
  • Familiarity with usage/cost monitoring for cloud or API-based services
  • Comfort operating in live production environments where reliability and data integrity are critical

Security & Compliance

  • Working knowledge of information security principles as they apply to SaaS and API-based systems: access controls, credential management, data handling, and audit logging
  • Ability to engage constructively with InfoSec teams, providing clear technical context to support reviews and risk assessments

Collaboration & Communication

  • Ability to communicate technical concepts clearly to non-technical colleagues and program staff
  • Experience contributing to cross-functional teams alongside product, engineering, and operations stakeholders
  • Strong documentation habits: runbooks, SOPs, architecture notes, and internal guides

Key skills

BA/BSc/HND

At a glance

Company

International Rescue Committee

Location

Nairobi, kenya

Employment

Full Time

Experience

Valid until

October 16, 2026

Created

October 2, 2026

More opportunities

Similar roles you might like

Senior Ai Platform Engineer (cloud) - Ke At Absa Bank Limited

Absa Bank Limited

Nairobi, kenya

full-time

Job Summary Absa Group's Chief Data Analytics and Applied AI Office (CDAIO) requires a technically exceptional and commercially grounded AI Platform Engineer (Cloud) to design, build, operate, and continuously optimise the multi-cloud AI infrastructure that powers the bank's enterprise AI capability. The AI capability must enable the CDAIO to fulfil its mandate as steward of the bank's AI capabilities through the end-to-end delivery of the AI platform enablement, governance and acceptable use in service of the bank's strategic and commercial objectives. This role is the engineering backbone of a platform that supports various live AI projects across four business units (CIB, PPB, BB, and AR) and ten countries. This role demands deep technical mastery in cloud AI infrastructure, AI FinOps, zero-trust security architecture, agentic AI infrastructure, and platform observability, combined with the commercial fluency to govern AI compute costs at enterprise scale and communicate trade-offs to senior business and finance stakeholders. The role includes but not limited to applying critical thinking, design thinking, and problem-solving skills in an agile team environment to solve complex platform engineering challenges, delivering high-quality solutions at optimal cost to serve, in full compliance with Absa's Enterprise-Wide Risk Management Framework, Group Architecture standards, and AI Responsible Use Policy. The successful candidate carries full accountability for building high-performing, scalable, enterprise-grade Platform services. As well as build capability in others to do the same. Job Description KEY FOCUS AREAS AI Platform Engineering and Architecture: Design and operation of enterprise-grade, multi-cloud AI platform infrastructure supporting bank-wide AI delivery at scale across the AI platform stack (AWS Bedrock, Databricks AI, Microsoft Azure AI Foundry, Hugging Face, and GPU clusters). AI FinOps and Compute Cost Governance: Full accountability for AI compute cost models, chargeback and showback frameworks, provisioned throughput optimisation, and monthly cost-per-use-case reporting to Group Finance across all four business units. Platform Observability and SLA Engineering: AI-specific service reliability standards, observability tooling, and incident management for production AI workloads serving 43 live projects across ten countries. AI Security Architecture and Zero Trust: Zero-trust security design, OAuth / OIDC integration, prompt injection controls, and data residency compliance protecting Absa's AI platform across the different country jurisdictions. Agentic AI Infrastructure: Design and operation of the infrastructure layer enabling multi-agent AI systems, autonomous workflows, tool-calling architectures, and agent orchestration at enterprise scale. Accountabilities Platform Engineering and Architecture Lead the design, deployment, and continuous optimisation of Absa's multi-cloud AI platform stack: AWS Bedrock, Databricks AI, Microsoft Azure AI Foundry, Hugging Face Model Hub, and on-demand GPU clusters. Architect scalable, resilient, and reusable platform components including AI Gateway configuration, model serving infrastructure, vector database deployments, and data pipeline integration to support bank-wide AI delivery. Define and maintain infrastructure-as-code (IaC) standards (e.g. using Terraform or Pulumi), enabling repeatable, auditable multi-cloud AI deployments across Absa's operating territories (10 countries). Lead the design and operation of agentic AI infrastructure: orchestration runtime environments (e.g. Microsoft Foundry Agent Service, AWS Bedrock Agents), tool-calling schemas, agent memory and state management patterns, and multi-agent communication protocols. Develop and enforce cloud-agnostic model serving patterns to reduce platform lock-in and ensure workload portability across the CDAIO's multi-vendor stack. Identify and select appropriate internal and external technologies to deliver AI platform services; apply excellent judgement in continuously improving platform engineering practices. Take full accountability for end-to-end platform quality, completeness, and user experience across the development, deployment, and operational lifecycle. Positively contribute to the design and evolution of Group Architecture, infrastructure standards, and AI platform governance frameworks AI FinOps and Compute Cost Governance Own the AI compute cost model for the CDAIO, including chargeback and showback frameworks for Databricks DBU consumption, AWS Bedrock token-based pricing, Azure AI Foundry provisioned throughput units, and GPU cluster utilisation across all four business units. Design and maintain FinOps dashboards and cost attribution reports using AWS Cost Explorer, Databricks System Tables cost analytics, and Azure OpenAI utilisation tooling providing monthly cost-per-use-case reporting to Group Finance and the CDAIO COO. Evaluate and manage provisioned throughput versus on-demand consumption trade-offs for production AI workloads, presenting optimisation recommendations to the CDAIO and BU technology leads. Identify and execute AI compute cost optimisation opportunities: workload scheduling, spot instance strategies for training workloads, model distillation to reduce inference cost, and right-sizing of GPU clusters. Create business cases and solution specifications for AI platform investments and governance processes, including CTO and architecture approvals. Collaborate with the FinOps capability within the CDAIO COO to align AI platform costs to agreed budget envelopes and ensure spend anomalies are detected and escalated proactively Platform Observability and SLA Engineering Define, implement, and own AI-specific SLAs and OLAs covering inference latency, platform availability, token throughput, API gateway response times, and model serving reliability, with explicit targets agreed with each business unit technology lead. Implement and maintain AI platform observability tooling (e.g. Prometheus, Grafana, Datadog, Databricks Lakehouse Monitoring, or equivalent) providing real-time visibility of platform health, model drift alerts, and capacity utilisation. Design and operate incident management processes for AI platform failures: on-call runbooks, escalation paths, post-incident reviews, and root-cause remediation, ensuring minimal disruption to live AI projects across Absa's footprint. Lead service improvement initiatives, translating performance data into platform enhancement programmes and continuously reducing mean time to recovery (MTTR) across the platform estate. Own the release and change management process for AI platform components, including change governance, cutover management, and operational readiness sign-off in alignment with Absa's Group Technology change framework. Use production performance monitoring and customer data to inform technical design and implementation decisions; leverage systems and processes to measure, monitor, and manage platform performance bank-wide AI Security Architecture and Zero Trust Design and implement zero-trust security architecture for AI platform APIs and services such as OAuth 2.0 / OIDC integration, JWT/JWE/JWS token management, role-based access control (RBAC), and attribute-based access control (ABAC) for AI workloads. Implement prompt injection prevention, output filtering, and data exfiltration controls at the AI Gateway layer, protecting data confidentiality for all LLM and agentic AI interactions across business units. Design and enforce data residency and sovereignty controls for AI platform deployments across Absa's operating countries, ensuring compliance with country-specific data localisation requirements and cross-border data transfer restrictions. Conduct and maintain AI-specific threat models in collaboration with the Chief Information Security Office, covering third-party AI vendor risks (Databricks, AWS, Microsoft, Hugging Face), model supply chain integrity, and adversarial ML attack vectors. Apply and maintain all Group risk, governance, compliance, and regulatory standards and frameworks; hold accountability for all risk associated with AI platform engineering decision-making. Update, develop, and maintain all platform documentation in accordance with organisational technical standards and risk and governance frameworks. People, Capability and Agile Delivery Lead and develop a team of AI Platform Engineers, establishing clear performance objectives, providing regular coaching and feedback, and building a high-performance, self-directed squad aligned to agile delivery practices. Cascade platform direction across the team; ensure alignment on platform strategy, performance objectives, and delivery priorities. Assume end-to-end accountability for the right people in the right teams to deliver the platform strategy. Leverage coaching techniques across all squad-related activity to drive higher-quality design and deployment of AI platform services. Maintain comprehensive technical documentation, architectural decision records (ADRs), and operational runbooks for all platform components, ensuring service continuity is independent of individual staffing changes and contractor dependencies are actively mitigated. Conduct peer reviews, testing, and problem-solving within and across the broader CDAIO engineering community; identify and develop needed skills in self and others. Support the AI Embedment and Training capability in developing platform onboarding materials and self-service guides to accelerate business unit adoption of AI platform services. Proactively lead agile practices, remove barriers to success, and ensure seamless delivery in a continuously changing environment. Qualifications And Experience Education/ Qualification: Postgraduate degree in a quantitative discipline such as Computer Science, Data Science, Mathematics, Statistics, Engineering, or equivalent ((Masters-essential or PhD-advantageous). Certification in: Cloud - AWS Solutions Architect Professional, AWS Machine Learning Specialty, or Microsoft Azure AI Engineer Associate). FinOps - FinOps Foundation Certified Practitioner (FOCP) or equivalent AI cost governance credential. Security Certification - Certified Cloud Security Professional (CCSP) or AWS Security Specialty. IaC Certification - HashiCorp Terraform Associate or equivalent infrastructure-as-code credential. Work Experience 5-8 years of progressive leadership experience in Cloud AI Platform Engineering, with production experience managing multi-cloud AI platform stacks across at least two of: AWS Bedrock/SageMaker, Databricks AI, Microsoft Azure AI Foundry, or Hugging Face enterprise deployments. 2 - 3-year experience in the following: AI FinOps and Cost Governance: Demonstrated ownership of AI compute cost models and FinOps reporting in a multi-BU or multi-cloud environment, with evidence of cost optimisation outcomes. AI Security Architecture: Designing and implementing zero-trust AI security (OAuth/OIDC, JWT), prompt injection controls, data residency compliance in a regulated environment. Agentic AI Infrastructure: Production design of agent orchestration infrastructure (such as LangGraph, AutoGen, Foundry Agent Service, Bedrock Agents), tool-calling APIs, and agent state management. Platform Observability: Operating AI-specific observability tooling for inference latency, drift alerting, and capacity management (such as Prometheus, Grafana, Datadog, or Lakehouse Monitoring). Infrastructure-as-Code: Terraform, Pulumi, or equivalent for multi-cloud, multi-region AI infrastructure deployments; CI/CD pipeline design for platform components. Regulated Industry: AI platform engineering in financial services or a similarly regulated sector with model risk governance and change management obligations. Regulated Industry: AI platform engineering in financial services or a similarly regulated sector with model risk governance and change management obligations Advantageous: People leadership: Leading or mentoring a team of platform or infrastructure engineers in an agile delivery environment. Pan-African Deployments: Delivering AI platform services across multiple African jurisdictions with awareness of data localisation and cross-border data transfer requirements. Knowledge And Skills Multi-Cloud AI Platform Architecture: Expert design and operation of AWS Bedrock, Databricks AI, Azure AI Foundry, and Hugging Face in enterprise production environments across multiple business units and geographies. Agentic AI Infrastructure: Practical production knowledge of agent orchestration frameworks (LangGraph, AutoGen, Foundry Agent Service, Bedrock Agents), tool-calling API design, agent memory architecture, and multi-agent coordination patterns. AI FinOps and Cost Management: Chargeback and showback model design; DBU and token cost attribution; provisioned throughput versus on-demand optimisation; GPU cluster cost management; spend anomaly detection and FinOps dashboarding. AI Security and Zero Trust: OAuth 2.0, OIDC, JWT/JWE/JWS; RBAC and ABAC for AI workloads; prompt injection prevention; data exfiltration controls at the Gateway layer; AI threat modelling and data residency compliance. Infrastructure-as-Code: Terraform, Pulumi, or AWS CDK for multi-cloud AI infrastructure; CI/CD pipeline design for platform components; container orchestration using Docker, Kubernetes, and Helm. Platform Observability: Prometheus, Grafana, Datadog, OpenTelemetry, and Databricks Lakehouse Monitoring; custom metric design for AI workload health including inference latency, token throughput, and model drift. Cloud-Agnostic Model Serving: ONNX, BentoML, Triton Inference Server; containerised model deployment patterns for portability across AWS, Azure, and Databricks environments. MLOps Tooling: Working knowledge of MLflow, Kubeflow, Airflow, and CI/CD for ML, sufficient to collaborate effectively with AI Solution Engineers on model deployment and lifecycle management GPU and HPC Architecture: On-demand GPU cluster management; spot instance strategies; high-performance compute cost optimisation for large-scale model training and fine-tuning workloads. Enterprise Risk and Governance: Absa Enterprise Wide Risk Management Framework; Group Architecture standards; AI Responsible Use Policy; POPIA; country-specific data localisation requirements across Absa's ten operating countries. Agile Delivery: Sprint planning, backlog management, and continuous delivery practices in a self-directed squad environment; experience removing delivery barriers in a fast-moving, multi-stakeholder context. Education Bachelor's Degree: Information Technology

2 months ago

Ai Platform Engineer (cloud) - Ke At Absa Bank Limited

Absa Bank Limited

Nairobi, kenya

full-time

Job Summary Absa Group's Chief Data Analytics and Applied AI Office (CDAIO) requires an experienced and technically capable AI Platform Engineer (Cloud) to support the design, deployment, operation, and continuous improvement of the multi-cloud infrastructure powering the bank's enterprise AI capability. The role will contribute to the delivery of secure, scalable, reliable, and cost-effective AI platform services across multiple business units and countries. The platform supports AI use cases across Corporate and Investment Banking (CIB), Personal and Private Banking (PPB), Business Banking (BB), and Absa Regional Operations (AR). The successful candidate will work across technologies such as AWS Bedrock, Databricks AI, Microsoft Azure AI Foundry, Hugging Face, Kubernetes, and GPU-based infrastructure. The role requires practical experience in cloud platform engineering, infrastructure-as-code, AI workload deployment, platform observability, cloud cost optimisation, security controls, and agentic AI infrastructure. The role includes applying critical thinking, design thinking, and problem-solving skills within an agile engineering environment to address complex platform challenges. The AI Platform Engineer will work closely with senior engineers, architects, security teams, FinOps specialists, and AI Solution Engineers to deliver high-quality platform services in line with Absa's architecture, risk, security, and responsible AI requirements. The successful candidate will take accountability for assigned platform components and services while contributing to the broader performance, resilience, and user experience of the enterprise AI platform. Job Description Key Focus Areas AI Platform Engineering and Architecture - Support the design, deployment, and operation of enterprise-grade, multi-cloud AI infrastructure across AWS Bedrock, Databricks AI, Microsoft Azure AI Foundry, Hugging Face, and GPU environments. AI FinOps and Compute Cost Optimisation - Monitor AI infrastructure consumption, support cost allocation and reporting, and identify opportunities to optimise token usage, Databricks consumption, provisioned throughput, and GPU utilisation. Platform Observability and Reliability - Implement and maintain monitoring, alerting, dashboards, and operational processes to ensure the availability, performance, and reliability of production AI platform services. AI Security and Zero-Trust Controls - Implement security controls for AI platform APIs, model endpoints, data pipelines, and agentic AI services in line with Absa's security architecture and regulatory requirements. Agentic AI Infrastructure - Support the deployment and operation of infrastructure enabling AI agents, tool-calling services, autonomous workflows, agent memory, and orchestration frameworks. Agile Engineering and Collaboration - Deliver platform enhancements through agile practices while collaborating with engineers, architects, business units, security teams, risk stakeholders, and third-party technology providers. Accountabilities Platform Engineering and Architecture Support the design, deployment, configuration, and operation of Absa's multi-cloud AI platform stack, including AWS Bedrock, Databricks AI, Microsoft Azure AI Foundry, Hugging Face, and GPU clusters. Build and maintain reusable platform components such as AI Gateway configurations, model serving environments, vector databases, API integrations, data pipelines, and containerised workloads. Develop and maintain infrastructure-as-code using technologies such as Terraform, Pulumi, AWS CDK, or equivalent tools. Contribute to repeatable and auditable infrastructure deployments across multiple cloud environments, regions, and operating countries. Configure and support agentic AI infrastructure, including orchestration environments, tool-calling APIs, agent memory, state management, and integration with enterprise systems. Implement cloud-agnostic model serving patterns that improve workload portability across AWS, Azure, Databricks, and Kubernetes-based environments. Support Kubernetes-based AI workloads using Docker, Kubernetes, and Helm. Assist with the evaluation and implementation of new platform technologies, services, and engineering patterns. Participate in architectural reviews, technical design sessions, peer reviews, and platform improvement initiatives. Create and maintain architectural diagrams, configuration documentation, operational procedures, and technical standards. Take accountability for the quality, performance, and operational readiness of assigned platform components. Escalate complex architectural, security, capacity, and operational risks to senior engineers and platform leadership. AI FinOps and Compute Cost Optimisation Monitor and analyse AI platform consumption across Databricks, AWS, Azure, GPU infrastructure, and third-party services. Support the development and maintenance of chargeback and showback frameworks for business units and individual AI use cases. Assist with cost attribution for Databricks DBU consumption, AWS Bedrock token usage, Azure AI Foundry provisioned throughput, and GPU workloads. Develop and maintain FinOps dashboards and cost reports using tools such as AWS Cost Explorer, Databricks System Tables, Azure Cost Management, and cloud-native monitoring services. Contribute to monthly cost-per-use-case reporting for Finance, platform leadership, and business unit stakeholders. Identify opportunities to optimise AI compute costs through workload scheduling, infrastructure right-sizing, token usage controls, caching, spot instances, and efficient model selection. Support assessments of provisioned throughput versus on-demand consumption for production AI workloads. Monitor spend anomalies and escalate unexpected usage, capacity, or budget risks. Provide technical input into business cases and investment proposals for AI platform services. Work closely with FinOps specialists and senior platform engineers to ensure infrastructure consumption remains within agreed budget parameters. Platform Observability and SLA Engineering Implement and maintain observability tooling for AI platform infrastructure and production AI services. Build dashboards and alerts covering: Inference latency Platform availability Token throughput API gateway response times Model endpoint health GPU and compute utilisation Databricks workload performance Vector database performance Capacity utilisation Model drift indicators Use tools such as Prometheus, Grafana, Datadog, OpenTelemetry, Databricks Lakehouse Monitoring, or equivalent technologies. Support the implementation and monitoring of AI-specific service-level agreements and operational-level agreements. Participate in incident response, troubleshooting, root-cause analysis, and post-incident reviews for AI platform failures. Develop and maintain operational runbooks, support procedures, escalation paths, and recovery documentation. Investigate platform performance issues and implement corrective or preventative actions. Support release, change, and configuration management processes for AI platform components. Conduct technical validation and operational readiness checks before platform changes are released into production. Use performance and usage data to recommend improvements to platform scalability, resilience, reliability, and cost efficiency. Contribute to initiatives focused on reducing incident volumes and mean time to recovery. AI Security Architecture and Zero Trust Implement zero-trust security controls for AI platform APIs, services, model endpoints, and agentic AI workloads. Configure and maintain authentication and authorisation controls using: OAuth 2.0 OpenID Connect JWT, JWE, and JWS Role-based access control Attribute-based access control Managed identities and service principals Support the implementation of prompt injection prevention, output filtering, content controls, and data loss prevention mechanisms at the AI Gateway layer. Implement controls to reduce the risk of unauthorised access, data exfiltration, insecure tool-calling, and excessive agent permissions. Support data residency and sovereignty controls across Absa's operating countries. Work with architecture, security, risk, and legal stakeholders to ensure AI workloads comply with applicable data localisation and cross-border transfer requirements. Contribute to AI-specific threat modelling covering model endpoints, agentic workflows, third-party AI providers, model supply chains, APIs, vector stores, and adversarial machine learning risks. Remediate identified security vulnerabilities and configuration risks within agreed timelines. Maintain platform documentation and evidence required for security reviews, audits, architecture approvals, and risk governance processes. Apply Absa's Enterprise-Wide Risk Management Framework, Group Architecture standards, information security requirements, and AI Responsible Use Policy in all engineering activities. Agentic AI Infrastructure Support the deployment and operation of agent orchestration technologies such as LangGraph, Microsoft Azure AI Foundry Agent Service, Amazon Bedrock Agents, AutoGen, or equivalent frameworks. Configure infrastructure for agent tools, APIs, memory services, vector stores, workflow engines, and enterprise system integrations. Implement secure tool-calling patterns, including identity propagation, permission controls, audit logging, timeout management, and failure handling. Support agent state management, session persistence, memory controls, and multi-agent communication patterns. Implement monitoring and tracing for agent execution paths, tool calls, latency, errors, and resource consumption. Work with AI Solution Engineers to move agentic AI solutions from development into controlled, production-ready environments. Contribute to platform standards for agent testing, deployment, monitoring, rollback, and lifecycle management. Investigate and resolve infrastructure issues affecting the performance, security, or reliability of agentic AI workloads. Agile Delivery and Capability Development Participate actively in sprint planning, backlog refinement, daily stand-ups, technical demonstrations, and retrospectives. Estimate engineering effort and deliver assigned platform features within agreed timelines and quality standards. Collaborate with platform engineers, cloud engineers, AI Solution Engineers, architects, security specialists, data engineers, and business unit technology teams. Participate in code reviews, infrastructure reviews, testing, troubleshooting, and technical problem-solving. Contribute to platform engineering standards, reusable templates, automation libraries, and delivery accelerators. Maintain comprehensive technical documentation, architectural decision records, deployment guides, and operational runbooks. Share technical knowledge and provide guidance to junior engineers and other members of the engineering community. Support the development of platform onboarding materials, self-service documentation, and user guides for business unit technology teams. Proactively identify technical dependencies, delivery risks, and operational barriers and escalate these appropriately. Remain current with developments in cloud AI platforms, agentic AI, MLOps, FinOps, AI security, and platform engineering. Qualifications And Experience Education and Qualifications Bachelor's degree in Computer Science, Information Technology, Data Science, Mathematics, Statistics, Engineering, or a related quantitative discipline is essential. A postgraduate qualification is advantageous. Relevant practical experience may be considered where supported by a strong record of cloud and platform engineering delivery. Advantageous Certifications One or more of the following certifications would be advantageous: Cloud AWS Certified Solutions Architect AWS Certified Machine Learning Engineer Microsoft Certified: Azure AI Engineer Associate Microsoft Certified: Azure Solutions Architect Expert Databricks Certified Data Engineer or Machine Learning certification Infrastructure-as-Code HashiCorp Certified: Terraform Associate Equivalent Terraform, Pulumi, or cloud infrastructure certification FinOps FinOps Certified Practitioner Equivalent cloud cost management or financial operations certification Security Certified Cloud Security Professional AWS Certified Security Microsoft Security, Compliance, and Identity certification Equivalent cloud or cybersecurity certification Work Experience Approximately 4 to 6 years of relevant experience in cloud engineering, platform engineering, DevOps, MLOps, infrastructure engineering, or AI platform engineering. At least 2 years of practical experience supporting cloud-based data, machine learning, generative AI, or AI platform workloads in a production environment. Production experience with at least two of the following: AWS Bedrock or Amazon SageMaker Databricks Microsoft Azure AI Foundry or Azure Machine Learning Hugging Face Kubernetes-based model serving Practical infrastructure-as-code experience using Terraform, Pulumi, AWS CDK, or an equivalent technology. Experience building or supporting CI/CD pipelines for cloud infrastructure, platform components, data services, or machine learning workloads. Experience with Docker, Kubernetes, Helm, APIs, identity integration, and cloud-native platform services. Experience implementing monitoring, dashboards, alerts, and operational support processes for production platforms. Working knowledge of cloud cost management, cost allocation, capacity monitoring, and infrastructure optimisation. Experience applying cloud security controls, identity and access management, secrets management, and secure API integration. Experience working within enterprise risk, architecture, security, and change management processes. Experience in financial services, telecommunications, healthcare, insurance, or another regulated industry is advantageous. Knowledge And Skills Multi-Cloud AI Platform Engineering - Practical knowledge of designing, deploying, and supporting AI services across AWS, Microsoft Azure, Databricks, Hugging Face, or Kubernetes-based environments. Agentic AI Infrastructure - Working knowledge of agent orchestration frameworks, tool-calling API patterns, agent memory, state management, tracing, and multi-agent workflows. AI FinOps and Cost Management - Knowledge of cloud consumption models, token-based pricing, Databricks DBUs, provisioned throughput, GPU utilisation, chargeback and showback reporting, and spend anomaly detection. AI Security and Zero Trust - Working knowledge of OAuth 2.0, OIDC, JWT, RBAC, ABAC, API security, managed identities, secrets management, prompt injection controls, data loss prevention, and secure agent tool access. Infrastructure-as-Code - Strong practical experience with Terraform, Pulumi, AWS CDK, or equivalent infrastructure automation technologies. Containerisation and Orchestration - Experience with Docker, Kubernetes, Helm, container registries, workload scheduling, resource allocation, and production container operations. Platform Observability - Experience with Prometheus, Grafana, Datadog, OpenTelemetry, cloud-native monitoring tools, or Databricks Lakehouse Monitoring. Cloud-Agnostic Model Serving - Working knowledge of containerised model deployment and serving technologies such as ONNX, BentoML, Triton Inference Server, Kubernetes, or equivalent frameworks. MLOps Tooling - Working knowledge of MLflow, Kubeflow, Airflow, model registries, feature stores, automated testing, and CI/CD for machine learning workloads. GPU Infrastructure - Understanding of GPU workload deployment, capacity management, right-sizing, spot instance strategies, and cost optimisation for model training and inference. Enterprise Risk and Governance - Working knowledge of information security, technology risk, architecture governance, responsible AI, privacy, data residency, and change management requirements within a regulated environment. Agile Delivery - Experience working in agile engineering teams using sprint planning, backlog management, iterative delivery, peer review, testing, and continuous improvement practices. Education Bachelor's Degree: Information Technology

2 months ago

Director Of Ai And Digital Workplace Platforms

International Rescue Committee

kenya

full-time

Job Background / Overview The IRC is investing in enterprise AI and modern workplace tools to improve the speed, quality, and scale of its programs and operations. As the organization expands its use of platforms like Microsoft 365, Box, Anthropic Claude, Microsoft Copilot, and Power Platform, it needs a senior leader to set the direction for how these tools are managed, adopted, and governed across a globally distributed workforce. The Director of AI and Digital Workplace Platforms owns this portfolio. The role is responsible for the strategy, operations, and staff enablement across the full digital workplace stack. That includes setting the roadmap for platform adoption and investment, managing vendor operations, building and leading a team that delivers platform administration and user support, and driving measurable adoption of AI tools across the organization. This role works closely with the Head of AI and Program Technology Development, as well as cross-functional partners in IT, People & Culture, and departmental leadership teams. The focus here is on the tools and platforms that staff use every day, and making sure IRC gets real value from them. Major Responsibilities Digital Workplace & AI Platform Strategy (25%) Define and maintain the roadmap for the digital workplace stack: M365, Box, Claude, Copilot, and Power Platform Evaluate new tools and platform capabilities; make recommendations on where to invest, expand, or sunset Build the business case for platform investments, including cost-benefit analysis and projected ROI Align platform strategy with broader IT and organizational priorities Track industry trends in enterprise AI and workplace technology to keep IRC current Vendor & Platform Operations (20%) Manage day-to-day vendor operations with Anthropic, Microsoft, and other digital workplace platform providers, including license and account-level issues, feature roadmap tracking, and technical escalations; coordinate with senior leadership on broader strategic and contractual matters Oversee license management, provisioning, and cost optimization across all platforms Ensure platform configurations, access controls, and security settings meet organizational standards Monitor platform health, usage trends, and license utilization; report to leadership on a regular cadence Track vendor policy changes, feature releases, and pricing shifts; assess impact and coordinate response AI Adoption & Enablement (25%) Drive adoption of AI tools across the digital workplace portfolio, identifying high-impact use cases and working with departmental teams to put them into practice Define success metrics for AI adoption and track progress against them Partner with appropriate change management and training leadership on communications, training priorities, and rollout planning for platform changes and new capabilities Surface training needs and content gaps to the teams responsible for staff learning; provide platform-specific context and user feedback Collect and synthesize user feedback on an ongoing basis; route recurring issues and feature requests to the right teams Coordinate with the Head of AI and Program Technology Development to ensure consistency in how AI tools are positioned and supported Governance & Compliance (15%) Contribute to AI governance processes and policy updates by providing platform-level insight, usage data, and operational perspective to the AI Risk & Governance Committee Ensure platform usage aligns with organizational policies, data privacy requirements, and the IRC’s AI governance framework Provide platform usage data, access logs, and incident reports to support governance reviews Flag usage patterns or behaviors that present compliance or security risks Maintain documentation of platform configurations, access decisions, and policy changes Team Leadership & Operations (15%) Lead and manage the Digital Workplace team, including platform administrators and support staff Set priorities, manage workload, and ensure the team delivers responsive, high-quality support Own the support model for digital workplace platforms: ticketing, triage, escalation, and resolution Maintain the help resource library covering FAQs, how-to guides, and reference materials across all platforms Manage the team budget and track spend against plan Required Experience & Skills Strategy & Leadership 8+ years of experience in IT, enterprise applications, or digital workplace roles, with at least 3 years in a leadership position managing a team Experience setting platform or product strategy in an enterprise environment Track record of building business cases and managing budgets for technology investments Ability to translate organizational needs into a clear platform roadmap Platform Administration & Operations Deep working knowledge of Microsoft 365 administration, including licensing, security, compliance, and governance Experience with at least two of the following: Box, Power Platform, Anthropic Claude, Microsoft Copilot Experience managing vendor operations, including license management, technical escalations, and feature roadmap tracking Comfort working with usage data and producing clear, actionable reports for leadership AI Adoption & Enablement Hands-on experience with enterprise AI tools as a user and/or administrator Experience driving technology adoption across a large, distributed organization Ability to identify high-value use cases and measure adoption outcomes User Support & Communication Experience designing and running a user support function, including ticketing, triage, and escalation Strong written and verbal communication skills; able to explain technology clearly to non-technical audiences Experience creating and maintaining user-facing help content Governance & Compliance Familiarity with AI governance frameworks and responsible AI use guidelines Working knowledge of data privacy principles as they apply to SaaS and AI platforms Preferred Experience in a nonprofit, international development, or humanitarian sector IT environment Experience managing Power Platform environments (Power Apps, Power Automate, Power BI) Background in change management, organizational learning, or staff enablement Organization The following positions report into this role: Associate Director, Applications Support Sr. Application Support Specialist Support Specialist (Open) AI Platform Engineer Working Environment This role is remote, with the possibility of in-person work depending on location. Stakeholders are located across multiple time zones (GMT-8 to GMT+3), which may require occasional early or late meetings. This is a remote position open to internal candidates based in countries where IRC operates who have the right to work in their location. Successful candidates will be hired on a local employment contract and according to local salary scale. This role is open to candidates located and with the right to work in United States of America, United Kingdom, or Kenya. Compensation: (US Pay Rate: $158,492-184,536/yr; UK Pay Rate: 77,499-93,814/yr ). Ranges are based on various factors including the labor market, job type, internal equity, and budget. Exact offers are calibrated by work location, individual candidate experience and skills relative to the defined job requirements. Equal Opportunity Employer: IRC is an Equal Opportunity Employer. IRC considers all applicants on the basis of merit without regard to race, sex, color, national origin, religion, sexual orientation, age, marital status, veteran status, disability or any other characteristic protected by applicable law. Professional Standards: All International Rescue Committee workers must adhere to the core values and principles outlined in IRC Way - Standards for Professional Conduct. Our Standards are Integrity, Service, Equality and Accountability. In accordance with these values, the IRC operates and enforces policies on Safeguarding, Conflicts of Interest, Fiscal Integrity, and Reporting Wrongdoing and Protection from Retaliation. IRC is committed to take all necessary preventive measures and create an environment where people feel safe, and to take all necessary actions and corrective measures when harm occurs. IRC builds teams of professionals who promote critical reflection, power sharing, debate, and objectivity to deliver the best possible services to our clients. US Benefits: We offer a comprehensive and highly competitive set of benefits. In the US, these include: 10 sick days, 10 US holidays, 20-25 paid time off days depending on role and tenure, medical insurance starting at $163 per month, dental starting at $6.50 per month, and vision starting at $5 per month, FSA for healthcare and commuter costs, a 403b retirement savings plans with immediately vested matching, disability & life insurance, and an Employee Assistance Program which is available to our staff and their families to support counseling and care in times of crisis and mental health struggles.

3 months ago