ML Software Engineer, Data Plane
Core
Design and implement the inference data plane for large language models running on custom hardware, covering model execution, memory management, and data movement.
Role type
Senior IC software engineer specializing in ML inference systems and custom hardware integration
Builds
High-performance inference software stack for large distributed models on custom accelerators
Domain
Artificial Intelligence / Machine Learning Infrastructure / Custom Hardware
Deliverable
production ML models
Required skills
low-level code optimization, custom hardware architecture, parallel computing, distributed systems, model validation, profiling, CI/CD pipeline integration
Preferred skills
LLM fundamentals (transformer, MoE), ML frameworks (PyTorch, JAX, vLLM, TensorRT), GPU/TPU/Neuron deployment
Technologies
PyTorch, vLLM, JAX, SGLang, Dynamo, TorchXLA, TensorRT, CUDA
Responsibilities
Develop and optimize compute kernels for custom ML accelerators; Implement and validate LLM architectures end-to-end; Integrate custom accelerator backends into open-source serving frameworks; Build and maintain test infrastructure for model correctness; Profile and optimize inference workloads for latency and throughput; Own features from design through production integration
Seniority
Senior, hands-on IC