Site Reliability Engineer - Applied Machine Learning Engine (Singapore)
Core
Build and operate massively distributed recommendation systems globally, ensuring high availability of centric machine learning services through system engineering and automation.
Role type
Senior IC Site Reliability Engineer (Applied Machine Learning)
Builds
Highly automated systems and pipelines for distributed recommendation services
Domain
Internet / Recommendation Systems / Distributed Systems
Deliverable
production ML models
Required skills
distributed systems troubleshooting, large-scale system design, Python or C/C++ programming, hardware/software integration, performance analysis, software testing and validation
Preferred skills
code optimization, routine task automation, machine learning frameworks (TensorFlow, Pytorch, MXNet, PaddlePaddle), algorithms and data structures
Technologies
TensorFlow, PyTorch, MXNet, PaddlePaddle
Responsibilities
Research, design, and develop computer and network software; Analyze user needs and develop software solutions; Update software and enhance existing capabilities; Work with hardware engineers to integrate systems and develop performance requirements
Seniority
Mid-Senior, hands-on IC