Python Inference Engineer
Full-time Mid-Senior LevelJob Overview
What You’ll Do
- Build and improve the inference layer of the Gcore Inference platform.
- Integrate and operate inference frameworks such as vLLM, SGLang, NVIDIA Dynamo, and TensorRT-LLM.
- Bring new language and multimodal models into production.
- Improve inference latency, throughput, memory use, GPU utilization, and cost efficiency.
- Debug performance and reliability issues across model code, inference frameworks, GPU execution, networking, and Kubernetes.
- Work with platform, infrastructure, product, and customer-facing teams to turn inference improvements into reliable product features.
- Contribute improvements to open-source inference projects when appropriate.
Make Your Resume Now