Senior AI Infrastructure Engineer
Core
Build and operate the infrastructure for production AI models, GPU clusters, and distributed training systems to support clinical AI products serving millions of patients.
Role type
Senior IC AI Infrastructure Engineer
Builds
Production model serving pipelines, GPU cluster management, and distributed training infrastructure for healthcare AI
Domain
Healthcare + AI Infrastructure
Deliverable
production ML models
Required skills
LLM deployment and inference serving, GPU cluster orchestration, performance profiling and debugging, distributed systems programming, Python, Linux, containers
Preferred skills
Distributed training frameworks (PyTorch FSDP, Megatron), inference engine tuning (MoE, quantization, speculative decoding), GPU interconnects (NCCL, RDMA), CUDA/Triton kernel development
Technologies
vLLM, SGLang, TensorRT-LLM, Triton Inference Server, Kubernetes, Slurm, Python, Go, C++, Rust
Responsibilities
Build model-serving infrastructure with autoscaling and rollback capabilities; manage GPU cluster scheduling and resource allocation; optimize inference latency and throughput; support distributed training job management; create observability dashboards and incident traceability; own production reliability and recovery procedures; optimize compute costs and capacity planning
Seniority
Senior, hands-on IC