Software Engineer - ML Infrastructure
Core
Build ML infrastructure and data systems enabling research teams to train and deploy state-of-the-art medical imaging models.
Role type
Senior IC ML Infrastructure Engineer
Builds
Distributed training stacks, data pipelines, and inference systems for medical imaging foundation models
Domain
Healthcare / Medical Imaging / AI Infrastructure
Deliverable
production ML models
Required skills
Python, PyTorch/JAX, distributed training (FSDP/DeepSpeed/Megatron), data pipeline engineering (Spark/Airflow/BigQuery), cloud infrastructure (AWS/GCP), containerization (Docker/Kubernetes)
Preferred skills
Reinforcement learning training infrastructure, high-performance inference (vLLM/SGLang/TensorRT), MLOps, DICOM/healthcare data standards, privacy-preserving systems
Technologies
PyTorch, JAX, Spark, Airflow, BigQuery, Snowflake, Databricks, Chalk, AWS, GCP, Docker, Kubernetes, vLLM, SGLang, TensorRT, Triton, DICOM
Responsibilities
Build and optimize distributed training infrastructure for foundation models on large-scale medical imaging; Build reinforcement learning training stacks for online, multi-reward RL at scale; Design high-throughput data loading and preprocessing for volumetric and multimodal datasets; Implement robust data pipelines to collect and store large-scale medical imaging data; Partner with researchers to prototype ideas and deliver production-ready code; Contribute to production serving, deployment, and monitoring pipelines
Seniority
Senior, hands-on IC