E
Eplusincvia Greenhouse
Sr Systems Engineer – AI Infrastructure (Req#1248)
REMOTEPosted 6d ago
MLOpsSeniorFull-time
Not sure if you're a good fit?
Upload your resume and TixelJobs AI will compare it against Sr Systems Engineer – AI Infrastructure (Req#1248) at Eplusinc. Get a match score, missing keywords, and improvement tips before you apply.
Free preview · Your resume stays private
About the Role
Overview
We are looking for an Sr Systems engineer with strong hands-on experience in enterprise infrastructure, including GPU-accelerated platforms, VMware, and Kubernetes. This role will support and scale infrastructure for AI/ML workloads across hybrid environments that include NVIDIA DGX, HGX, and Nutanix.
Your Impact
The essential functions of this position include:
- Support deployment and maintenance of NVIDIA DGX, HGX, and GPU-accelerated systems
- Administer VMware environments, including ESXi, vCenter, and virtual networking/storage.
- Deploy and support Kubernetes clusters across various environments and distros (e.g., RKE, OpenShift, AKS, EKS, GKE)
- Perform day-to-day system administration across compute, storage, and networking layers
- Automate infrastructure tasks using Shell scripts, Ansible, or similar tools
- Collaborate with DevOps, data science, and engineering teams to ensure scalable, resilient infrastructure for AI/ML workloads
- Monitor infrastructure health and performance; participate in troubleshooting and root cause analysis
Qualifications
- Min 6 years of experience in systems engineering or enterprise infrastructure roles
- Experience with NVIDIA DGX, HGX, or other GPU-based compute platforms
- Strong working knowledge of VMware virtualization technologies
- Proficient with Kubernetes and experience managing multiple distributions (e.g., RKE, OpenShift, AKS, EKS)
- Understanding of enterprise storage, networking, and system monitoring tools
- Working knowledge of Mellanox (NVIDIA) networking and InfiniBand environments
- Understanding of RDMA and high-speed interconnects is a plus
- Experience with High Availability (HA), clustering, and failover technologies
- Scripting and automation experience (e.g., Bash, Python, Ansible)
- Strong communication, documentation, and troubleshooting skills
- Comfortable working independently in a remote environment
Ready to apply?
This job is active. Apply now to get in early.