Software Engineer - Machine Learning Platform
Core
Design, build, and maintain an end-to-end ML training platform on AWS, focusing on productionizing GPU workloads, optimizing performance/cost, and enabling Gen AI/LLM fine-tuning workflows.
Role type
Senior IC Machine Learning Platform Engineer
Builds
Scalable, reliable ML training systems and pipelines on AWS
Domain
Cloud Infrastructure / Machine Learning / Generative AI
Deliverable
production ML models | infrastructure
Required skills
Python, Deep Learning frameworks (PyTorch/TensorFlow), Kubernetes (EKS), AWS services (EC2, EKS, S3, IAM, VPC, CloudWatch), Distributed training (DDP, FSDP, DeepSpeed), CI/CD pipelines, GPU profiling and optimization, Responsible AI practices
Preferred skills
Multi-cloud training experience, Cloud-native networking/storage patterns, Large dataset pipeline optimization (Parquet, WebDataset), Distributed compute frameworks (Spark, Ray, Dask), Workflow orchestration (Airflow), Training cost/performance optimization, Observability (metrics, logs, traces, GPU telemetry)
Technologies
AWS, Kubernetes, PyTorch, TensorFlow, Python, Spark, Ray, Dask, Airflow, NodeJS
Responsibilities
Design and maintain end-to-end ML training platform; Run and tune single-node and distributed GPU training workloads; Build and operate training infrastructure on Kubernetes; Enable Gen AI and LLM training and fine-tuning workflows; Implement observability for training systems; Partner with data engineering and platform teams on security and cost guardrails; Improve developer experience via standardized containers and CI/CD; Validate AI-assisted code outputs using peer review and automated testing.
Seniority
Senior, hands-on IC
