Sourcing Criteria and Requirements for Senior Engineer - Infrastructure / DevOps

Must Haves

  • Extensive experience (8+ years) in Infrastructure, Platform Engineering, or SRE

  • Proven track record building and managing production infrastructure from the ground up

  • Deep expertise in Kubernetes in production at scale (beyond managed cloud services)

  • Strong cloud infrastructure knowledge (AWS, Azure, or GCP)

  • Infrastructure as Code proficiency (Terraform, Ansible, Helm, or similar)

  • Experience with GPU cluster management (NVIDIA H100 or equivalent) in production

  • Experience with air-gapped, on-premises, or sovereign deployment environments

  • Strong automation mindset with a focus on scalability, security, and operational resilience

  • Experience designing secure, scalable production environments for enterprise or government clients

  • Strong incident response discipline: SLOs, runbooks, post-mortems

  • Platform reliability and operational resilience expertise

  • Ability to work independently in a growing technology business

Nice to Haves

  • Background in cybersecurity, enterprise SaaS, defence technology, or other highly regulated environments

  • Experience with InfiniBand networking and NVIDIA NVLink topology

  • Prior leadership or mentoring experience of infrastructure/platform teams

  • Familiarity with FIPS-compliant and HSM-protected environments

Location and Working Pattern

  • Initially remote with travel to Abu Dhabi as required

  • Future relocation welcomed but not essential

  • Preference for candidates working close to Gulf Standard Time for collaboration

Other Preferences

  • Candidates should demonstrate an engineering mindset prioritising reliability, scalability, and automation rather than traditional sysadmin roles

  • Search should include varied titles such as Senior DevOps Engineer, Platform Engineer, Site Reliability Engineer, Infrastructure Architect

  • Experience supporting software engineering teams is important

Business Context

  • Role supports sovereign, intelligence-led AI systems in complex, high-security environments

  • Infrastructure stability, security and scalability are fundamental to Elile's products

  • Responsible for infrastructure underpinning AI and intelligence platforms rather than AI model development itself