Senior ML Software Engineer, Data Plane
Core
Design and implement the inference data plane for large language models running on custom hardware, covering model execution, memory management, and data movement.
Role type
Senior IC machine-learning software engineer (custom hardware inference)
Builds
High-performance inference software for large distributed models on custom accelerators
Domain
AI infrastructure / Custom hardware acceleration / LLM serving
Deliverable
production ML models
Required skills
C/C++, Linux systems, computer architecture, parallel computing, compute kernel development, model validation, CI/CD pipeline ownership
Preferred skills
PyTorch, vLLM, CUDA, distributed systems (RDMA), hardware simulation, speculative decoding, KV cache optimization
Technologies
PyTorch, vLLM, SGLang, Dynamo, TorchXLA, TensorRT, CUDA
Responsibilities
Develop and optimize compute kernels for custom ML accelerators; Implement and validate LLM architectures end-to-end; Integrate custom accelerator backends into open-source ML serving frameworks; Build and maintain test infrastructure for model correctness; Profile and optimize inference workloads; Mentor engineers and drive design reviews
Seniority
Senior, hands-on IC