SDE II, ML Infra Services, Annapurna Labs
Core
Lead the design and implementation of an ML infrastructure platform to run, optimize, and analyze machine learning workloads on AWS ML accelerators (Inferentia, Trainium, Neuron).
Role type
Senior IC software engineer (ML infrastructure platform)
Builds
ML infrastructure platform for capacity management, workload scheduling, and fleet orchestration across ML accelerators
Domain
Cloud-scale machine learning infrastructure, custom silicon, heterogeneous compute
Deliverable
production ML models | product features
Required skills
ML infrastructure design, workload scheduling, fleet orchestration, capacity management, system architecture, performance profiling, optimization, resource management, distributed systems design, observability, telemetry, Go/Java/Python, Kubernetes
Preferred skills
Large-scale ML infrastructure for recommendation/search/ads, application and kernel-level performance profiling, integrated software/hardware performance analysis, debugging complex distributed systems
Technologies
Kubernetes, Go, Java, Python, Javascript/TypeScript
Responsibilities
Design and code solutions for software architecture efficiency, create metrics and implement automation, resolve root cause of software defects, build high-impact solutions for large customer base, participate in design discussions and code reviews
Seniority
Senior, hands-on IC