Senior Software Engineer, ML Ops
Core
Design, develop, and scale machine learning infrastructure and automation to support enterprise AI systems, bridging the gap between research and production.
Role type
Senior IC MLOps Engineer
Builds
ML application development and deployment infrastructure on AWS and on-premises
Domain
Healthcare AI / Machine Learning Operations
Deliverable
production ML models
Required skills
Kubernetes, AWS, Python, Infrastructure-as-Code (Terraform, Helm), Observability (Prometheus, Grafana, Datadog), CI/CD for ML, Scalable backend architecture
Preferred skills
PyTorch, Scikit-learn, Data workflow orchestration (Airflow, Kubeflow), Streaming data processing (Kafka, Flink, Spark), Security and compliance in ML systems
Technologies
AWS, Kubernetes, Terraform, Helm, Prometheus, Grafana, Datadog, Airflow, Kubeflow, Kafka, Flink, Spark Streaming
Responsibilities
Architect and build infrastructure and automation for ML application development and deployment; Drive system design and lead architectural discussions for the MLOps suite; Research, evaluate, and implement new MLOps tools and best practices; Optimize ML workflows for efficient and reproducible deployment and monitoring; Automate ML operations including CI/CD, feature engineering pipelines, and deployment strategies; Mentor junior engineers and enforce high coding standards
Seniority
Senior, hands-on IC