Senior Storage Systems Engineer - remote in the US
Full-time DirectorJob Overview
Overview
The role is to deploy, integrate, and operate high-performance storage for GPU-accelerated compute and AI platforms. You will own the storage layer where Kubernetes meets bare metal — standing up NFS-based high-performance storage, wiring it into clusters via CSI, and tuning it to keep data flowing to GPU workloads at scale. Work spans hybrid, edge, and air-gapped deployments built on the Mirantis K0rdent stack.
About the Role
We are looking for a senior systems engineer who treats storage as infrastructure to be automated, observed, and tuned — not hand-managed. The right candidate is fluent in Kubernetes storage, deeply versed in Linux storage and networking fundamentals down to the kernel and NFS-client layer, and knows how to make high-performance NAS actually perform under demanding workloads. You should reach for infrastructure-as-code and GitOps by default, be self-directed in diagnosing performance and reliability issues end to end, set operational standards for others to follow, and communicate clearly across teams. Bare-metal hardware experience is a strong plus, but deep Linux storage knowledge is essential.
Responsibilities
1. Storage Integration & Operation
- Integrate NFS-based high-performance storage (e.g., VAST, Dell PowerScale) into Kubernetes clusters via CSI, storage classes, and persistent volumes.
- Tune the NFS data path — mount options, nconnect/RDMA, Linux client, and network settings — for high-throughput, low-latency GPU/AI workloads.
- Deploy and operate storage services and operators; manage capacity, quotas, snapshots, and lifecycle.
2. Linux Platform & System Integration
- Configure and optimize Linux systems for storage workloads, including driver setup, file system layout, network tuning, and kernel parameter optimization.
- Deliver storage integration for k0s-based Kubernetes via Cluster API (CAPI) and K0rdent management/child cluster topologies.
- Operate storage in fully disconnected (air-gapped) environments, including local artifact/mirror connectivity (Harbor) and PKI/TLS considerations.
3. Automation & Observability
- Automate storage provisioning and configuration with infrastructure-as-code (Terraform/OpenTofu) and GitOps pipelines (ArgoCD or Flux).
- Build monitoring, alerting, and observability for storage performance, capacity, and health.
- Diagnose and resolve performance, reliability, and scaling issues across the storage stack.
Make Your Resume Now