Senior Software Engineer, Machine Learning Infrastructure
Core
Build and operate shared infrastructure for production ML and AI systems, including data pipelines, model serving, and evaluation harnesses for frontier AI labs.
Role type
Senior IC machine learning infrastructure engineer
Builds
Scalable ML platforms, LLM orchestration systems, evaluation benchmarks, and post-training workflows for AI training and inference
Domain
Generative AI, Machine Learning Infrastructure, Cloud Infrastructure
Deliverable
production ML models | infrastructure
Required skills
Python, Go, TypeScript, Kubernetes, Docker, Terraform, CI/CD, Cloud infrastructure (AWS/GCP), ML infrastructure (model serving, feature stores, embeddings), Modern data platforms (BigQuery, Airflow, Spark, Beam/Dataflow), LLM orchestration and optimization
Preferred skills
Ray, Anyscale, KubeRay, Ray Serve, vLLM, Triton, PyTorch, GPU-backed inference, LLM evaluation frameworks, Vertex AI, Bigtable, Redis, Post-training techniques (fine-tuning, RLHF, reward modeling), Agentic systems, MCP integrations
Technologies
Kubernetes, Docker, Terraform, AWS, GCP, BigQuery, Airflow, Spark, Beam, Dataflow, Ray, vLLM, Triton, PyTorch, Vertex AI, Bigtable, Redis
Responsibilities
Build and operate shared infrastructure for production ML and AI; Develop and scale LLM platform with provider integrations and observability; Build evaluation infrastructure including LLM eval harnesses and benchmarks; Support post-training workflows including fine-tuning and RL pipelines; Optimize inference infrastructure for GPU serving and autoscaling; Partner with teams to productionize models and establish ML best practices; Improve reliability, scalability, and developer experience of the ML platform
Seniority
Senior, hands-on IC