TixelJobs
E
Eplusincvia Greenhouse

Sr Systems Engineer – AI Infrastructure (Req#1248)

REMOTEPosted 6d ago
MLOpsSeniorFull-time

Not sure if you're a good fit?

Upload your resume and TixelJobs AI will compare it against Sr Systems Engineer – AI Infrastructure (Req#1248) at Eplusinc. Get a match score, missing keywords, and improvement tips before you apply.

Free preview · Your resume stays private

About the Role


Overview


We are looking for an Sr Systems engineer with strong hands-on experience in enterprise infrastructure, including GPU-accelerated platforms, VMware, and Kubernetes. This role will support and scale infrastructure for AI/ML workloads across hybrid environments that include NVIDIA DGX, HGX, and Nutanix.


Your Impact


 

The essential functions of this position include: 

  • Support deployment and maintenance of NVIDIA DGX, HGX, and GPU-accelerated systems
  • Administer VMware environments, including ESXi, vCenter, and virtual networking/storage.
  • Deploy and support Kubernetes clusters across various environments and distros (e.g., RKE, OpenShift, AKS, EKS, GKE)
  • Perform day-to-day system administration across compute, storage, and networking layers
  • Automate infrastructure tasks using Shell scripts, Ansible, or similar tools
  • Collaborate with DevOps, data science, and engineering teams to ensure scalable, resilient infrastructure for AI/ML workloads
  • Monitor infrastructure health and performance; participate in troubleshooting and root cause analysis

Qualifications


 

  • Min 6 years of experience in systems engineering or enterprise infrastructure roles
  • Experience with NVIDIA DGX, HGX, or other GPU-based compute platforms
  • Strong working knowledge of VMware virtualization technologies
  • Proficient with Kubernetes and experience managing multiple distributions (e.g., RKE, OpenShift, AKS, EKS)
  • Understanding of enterprise storage, networking, and system monitoring tools
  • Working knowledge of Mellanox (NVIDIA) networking and InfiniBand environments
  • Understanding of RDMA and high-speed interconnects is a plus
  • Experience with High Availability (HA), clustering, and failover technologies
  • Scripting and automation experience (e.g., Bash, Python, Ansible)
  • Strong communication, documentation, and troubleshooting skills
  • Comfortable working independently in a remote environment

 



Who We Are

At ePlus, we believe technology is a people business. Our team is passionate, skilled, and driven to deliver solutions that make a real difference. Join us and be part of a culture that values collaboration, innovation, and extraordinary results.

Corporate Values

  • Share