AI Platform Engineer (Cloud)
ABSA BANK LIMITED
Sandton, Gauteng
Job description
Empowering Africa’s tomorrow, together…one story at a time.
With over 100 years of rich history and strongly positioned as a local bank with regional and international expertise, a career with our family offers the opportunity to be part of this exciting growth journey, to reset our future and shape our destiny as a proudly African group.
Job Summary
Absa Group’s Chief Data Analytics and Applied AI Office (CDAIO) requires an experienced and technically capable AI Platform Engineer (Cloud) to support the design, deployment, operation, and continuous improvement of the multi-cloud infrastructure powering the bank’s enterprise AI capability.
The role will contribute to the delivery of secure, scalable, reliable, and cost-effective AI platform services across multiple business units and countries. The platform supports AI use cases across Corporate and Investment Banking (CIB), Personal and Private Banking (PPB), Business Banking (BB), and Absa Regional Operations (AR).
The successful candidate will work across technologies such as AWS Bedrock, Databricks AI, Microsoft Azure AI Foundry, Hugging Face, Kubernetes, and GPU-based infrastructure. The role requires practical experience in cloud platform engineering, infrastructure-as-code, AI workload deployment, platform observability, cloud cost optimisation, security controls, and agentic AI infrastructure.
The role includes applying critical thinking, design thinking, and problem-solving skills within an agile engineering environment to address complex platform challenges. The AI Platform Engineer will work closely with senior engineers, architects, security teams, FinOps specialists, and AI Solution Engineers to deliver high-quality platform services in line with Absa’s architecture, risk, security, and responsible AI requirements.
The successful candidate will take accountability for assigned platform components and services while contributing to the broader performance, resilience, and user experience of the enterprise AI platform.Job Description
Key Focus Areas
- AI Platform Engineering and Architecture - Support the design, deployment, and operation of enterprise-grade, multi-cloud AI infrastructure across AWS Bedrock, Databricks AI, Microsoft Azure AI Foundry, Hugging Face, and GPU environments.
- AI FinOps and Compute Cost Optimisation - Monitor AI infrastructure consumption, support cost allocation and reporting, and identify opportunities to optimise token usage, Databricks consumption, provisioned throughput, and GPU utilisation.
- Platform Observability and Reliability - Implement and maintain monitoring, alerting, dashboards, and operational processes to ensure the availability, performance, and reliability of production AI platform services.
- AI Security and Zero-Trust Controls - Implement security controls for AI platform APIs, model endpoints, data pipelines, and agentic AI services in line with Absa’s security architecture and regulatory requirements.
- Agentic AI Infrastructure - Support the deployment and operation of infrastructure enabling AI agents, tool-calling services, autonomous workflows, agent memory, and orchestration frameworks.
- Agile Engineering and Collaboration - Deliver platform enhancements through agile practices while collaborating with engineers, architects, business units, security teams, risk stakeholders, and third-party technology providers.
Accountabilities
Platform Engineering and Architecture
- Support the design, deployment, configuration, and operation of Absa’s multi-cloud AI platform stack, including AWS Bedrock, Databricks AI, Microsoft Azure AI Foundry, Hugging Face, and GPU clusters.
- Build and maintain reusable platform components such as AI Gateway configurations, model serving environments, vector databases, API integrations, data pipelines, and containerised workloads.
- Develop and maintain infrastructure-as-code using technologies such as Terraform, Pulumi, AWS CDK, or equivalent tools.
- Contribute to repeatable and auditable infrastructure deployments across multiple cloud environments, regions, and operating countries.
- Configure and support agentic AI infrastructure, including orchestration environments, tool-calling APIs, agent memory, state management, and integration with enterprise systems.
- Implement cloud-agnostic model serving patterns that improve workload portability across AWS, Azure, Databricks, and Kubernetes-based environments.
- Support Kubernetes-based AI workloads using Docker, Kubernetes, and Helm.
- Assist with the evaluation and implementation of new platform technologies, services, and engineering patterns.
- Participate in architectural reviews, technical design sessions, peer reviews, and platform improvement initiatives.
- Create and maintain architectural diagrams, configuration documentation, operational procedures, and technical standards.
- Take accountability for the quality, performance, and operational readiness of assigned platform components.
- Escalate complex architectural, security, capacity, and operational risks to senior engineers and platform leadership.
AI FinOps and Compute Cost Optimisation
- Monitor and analyse AI platform consumption across Databricks, AWS, Azure, GPU infrastructure, and third-party services.
- Support the development and maintenance of chargeback and showback frameworks for business units and individual AI use cases.
- Assist with cost attribution for Databricks DBU consumption, AWS Bedrock token usage, Azure AI Foundry provisioned throughput, and GPU workloads.
- Develop and maintain FinOps dashboards and cost reports using tools such as AWS Cost Explorer, Databricks System Tables, Azure Cost Management, and cloud-native monitoring services.
- Contribute to monthly cost-per-use-case reporting for Finance, platform leadership, and business unit stakeholders.
- Identify opportunities to optimise AI compute costs through workload scheduling, infrastructure right-sizing, token usage controls, caching, spot instances, and efficient model selection.
- Support assessments of provisioned throughput versus on-demand consumption for production AI workloads.
- Monitor spend anomalies and escalate unexpected usage, capacity, or budget risks.
- Provide technical input into business cases and investment proposals for AI platform services.
- Work closely with FinOps specialists and senior platform engineers to ensure infrastructure consumption remains within agreed budget parameters.
Platform Observability and SLA Engineering
- Implement and maintain observability tooling for AI platform infrastructure and production AI services.
- Build dashboards and alerts covering:
- Inference latency
- Platform availability
- Token throughput
- API gateway response times
- Model endpoint health
- GPU and compute utilisation
- Databricks workload performance
- Vector database performance
- Capacity utilisation
- Model drift indicators
- Use tools such as Prometheus, Grafana, Datadog, OpenTelemetry, Databricks Lakehouse Monitoring, or equivalent technologies.
- Support the implementation and monitoring of AI-specific service-level agreements and operational-level agreements.
- Participate in incident response, troubleshooting, root-cause analysis, and post-incident reviews for AI platform failures.
- Develop and maintain operational runbooks, support procedures, escalation paths, and recovery documentation.
- Investigate platform performance issues and implement corrective or preventative actions.
- Support release, change, and configuration management processes for AI platform components.
- Conduct technical validation and operational readiness checks before platform changes are released into production.
- Use performance and usage data to recommend improvements to platform scalability, resilience, reliability, and cost efficiency.
- Contribute to initiatives f
Good to know
How do I apply for this job?
Tap "Apply on Indeed" to open the original listing, where you can read the full description and apply directly. JobsZA never charges you to apply, and you should never pay money to get a job.
Found on Indeed · Posted 4 days ago
Similar jobs in Sandton
Hospitality Placements
R20K - R20K/mo