On-site Full Time
--
Avensys Consulting

Job Details

Avensys is a reputed global IT professional services company headquartered in Singapore. Our service spectrum includes enterprise solution consulting, business intelligence, business process automation, and managed services. Given our decade of success, we have evolved to become one of the top trusted providers in Singapore and service a client base across banking and financial services, insurance, information technology, healthcare, retail and supply chain. We are hiring for "AI Infrastructure and Platform Operations Resident - strong exp in NVAIE, NIM, Triton endpoint support and troubleshooting ,Kubernetes administration and cluster operations, GPU orchestration, Power Edge, Power Scale F210/One FS, AI networking, RoCE/RDMA, and storage performance" for Riyadh KSA
Experience - 8 + Years
Job Summary:BCM and server lifecycle management Kubernetes administration and cluster operations Run:ai resource pools, quotas, priorities, fair-share, reservations, and queue management GPU health, utilization, capacity, and workload-placement monitoring Power Edge, Power Scale F210/One FS, AI networking, RoCE/RDMA, and storage performance NVAIE, NIM, Triton endpoint support and troubleshooting Linux, incident management, change control, patching, upgrades, backup, and recovery Perform recurring health reviews across servers, GPUs, management and infrastructure nodes, Power Scale, switching, storage paths, Kubernetes, BCM, Run:ai, and critical AI services. Maintain BCM provisioning, node baselines, configuration consistency, lifecycle coordination, and recovery procedures. Administer Kubernetes nodes, namespaces, workloads, GPU resources, and platform add-ons within the MOJ operating model. Operate H200 inference, H200 fine-tuning, and L40S RAG resource pools, including quotas, priorities, fair-share, reservations, fractional-GPU allocation, and approved preemption. Monitor utilization, queue time, idle capacity, throughput, contention, endpoint availability, and workload placement. Support Power Scale capacity and protection reviews and investigate AI-network, storage, GPU-fabric, and endpoint issues. Coordinate firmware, driver, operating-system, Kubernetes, BCM, Run:ai, and NVAIE maintenance planning with MOJ change processes. Provide first-line incident triage, impact assessment, escalation, recovery tracking, and post-incident improvement actions. Maintain runbooks for cluster health, provisioning, workload submission, GPU allocation, endpoint troubleshooting, patching, upgrades, incident response, and recovery.
WHAT’S ON OFFERYou will be remunerated with an excellent base salary and entitled to attractive company benefits. Additionally, you will get the opportunity to enjoy a fun and collaborative work environment, alongside a strong career progression. To submit your application, please apply online or email your UPDATED CV in Microsoft Word format to [Click to show email] Your interest will be treated with strict confidentiality. CONSULTANT DETAILSConsultant Name: Saranya MAvensys Consulting Pte Ltd EA Licence 12C5759Privacy Statement: Data collected will be used for recruitment purposes only. Personal data provided will be used strictly in accordance with the relevant data protection law and Avensys' personal information and privacy policy.

Similar Jobs

About Avensys Consulting
Saudi, Jeddah
Information Technology and Services