Staff Software Engineer, Inference
Core
Design and operate CoreWeave's Kubernetes-native inference platform for low-latency, high-throughput AI workloads at massive scale.
Role type
Staff Software Engineer (IC5), hands-on technical leader
Builds
Kubernetes-native inference platform, request routing, scheduling, GPU resource management
Domain
Cloud infrastructure, distributed systems, AI inference (via careerplan.io/jobs/4670593006-staff-software-engineer-inference-at-coreweave)
Deliverable
production ML models
Required skills
Go, Python, C++, Kubernetes orchestration, distributed systems, networking, performance optimization, low-latency system design, inference batching, memory optimization
Preferred skills
vLLM, Triton, TensorRT-LLM, Ray Serve, TorchServe, CUDA, NCCL, RDMA, NUMA, GPU interconnects, large-scale AI/ML infrastructure
Technologies
Kubernetes, Go, Python, C++, vLLM, Triton, TensorRT-LLM, Ray Serve, TorchServe, CUDA, NCCL, RDMA
Responsibilities
Lead cross-team design initiatives for the inference platform, optimize inference performance (latency, throughput, GPU utilization), improve system reliability at scale, work on scheduling and batching strategies, implement memory optimization techniques
Seniority
Staff, hands-on IC with cross-team influence