Staff Machine Learning Engineer
Core
Design and operate multi-region model-serving fleets, production ML infrastructure (queueing, orchestration, deployment), and CI/CD tooling to ensure reliability and cost efficiency.
Role type
Staff Machine Learning Engineer (Infrastructure & Platform)
Builds
Production ML infrastructure, model-serving fleets, customer-specific pipelines, and deployment tooling.
Domain
Cloud Infrastructure & Machine Learning Operations
Deliverable
infrastructure
Required skills
Distributed systems operations, ML workload management, Kubernetes, cloud infrastructure, CI/CD, Infrastructure as Code, Incident response, Technical leadership, Mentoring.
Preferred skills
Celery, RabbitMQ, Kafka, Ray, MLflow, model registries, feature stores, ML observability, Python, Ruby on Rails, MongoDB, Redis.
Technologies
Kubernetes, Celery, RabbitMQ, Kafka, Ray, MLflow, MongoDB, Redis
Responsibilities
Drive technical vision for model training, deployment, serving, and observability; Lead incident response for ML systems; Improve engineering quality through design/code reviews; Connect technical decisions to business outcomes.
Seniority
Staff, hands-on IC with strategic leadership
