L3 System / Cloud Infrastructure / Nutanix / Automation Engineer (Ref 26504)
Job Description
Responsibilities • Define and drive the organization's cloud and platform infrastructure strategy, architecture, standards, and multi-year roadmap, ensuring scalable, secure, resilient, and cost-efficient solutions. • Serve as the L3/L4 technical authority and final escalation point for complex infrastructure, platform, and cross-domain incidents, leading root cause analysis (RCA) and implementing permanent resolutions. • Design, develop, and maintain Infrastructure as Code (IaC), platform automation, and GitOps practices using technologies such as Terraform and Python/Go. • Architect, implement, and continuously improve platform resilience, high availability, disaster recovery (DR), fault tolerance, and service reliability, including defining and maintaining SLOs and SLIs. • Design, optimize, secure, and manage enterprise Kubernetes environments, including networking, security, lifecycle management, and platform operations. • Establish, govern, and enforce platform engineering, security, infrastructure, and CI/CD standards, while mentoring engineers and promoting engineering best practices. • Independently manage and resolve complex incidents, service requests, problems, and changes within agreed SLAs, ensuring accurate documentation, timely ticket updates, stakeholder communication, and appropriate escalations. • Proactively identify, investigate, analyse, and resolve platform issues, leveraging advanced troubleshooting techniques, operational diagnostics, and root cause analysis to prevent recurrence. • Collaborate with clients, stakeholders, cross-functional teams, and automation teams to deliver platform improvements, optimise operational efficiency, and automate routine tasks. • Produce and maintain technical documentation, share knowledge, coach L1–L3 engineers, and contribute to quality assurance, operational excellence, and continuous service improvement. • Lead or contribute to infrastructure projects, platform enhancements, disaster recovery implementation and testing, and other technology initiatives as required. • Perform other related duties as assigned.
Requirements • Bachelor's degree in Information Technology, Computer Science, or a related discipline (or equivalent practical experience). • Strong expertise in virtualization, cloud infrastructure, storage, automation, and enterprise platform technologies. • Hands-on experience with VMware, Azure, AWS, Ansible, GitLab, DevOps/SRE practices, containers, and enterprise tools such as Veeam, Rubrik, Splunk, CyberArk, Opswat, NVIDIA AI, and storage platforms. • Relevant industry certifications are highly desirable, including VMware, Microsoft Azure, AWS, Veeam, and Rubrik certifications. • Excellent understanding of IT change management, with experience planning, assessing risks, executing changes, and documenting mitigation plans. • Strong communication and collaboration skills with cross-functional, multicultural teams and stakeholders. • Proven ability to work effectively in a fast-paced, high-pressure environment while managing multiple priorities. • Strong client-focused mindset with a commitment to delivering exceptional service and customer experience. • Excellent planning, organizational, problem-solving, and active listening skills, with the flexibility to adapt to changing business needs.
Licence no: 12C6060