SRE运维开发工程师(J73254)
Core
Design and develop online services and automation tools for financial joint modeling products to ensure reliability, stability, and data quality.
Role type
Senior Site Reliability Engineer (SRE)
Builds
Financial joint modeling online services, model deployment architectures, and reliability automation systems.
Domain
Financial services + Cloud-native infrastructure
Deliverable
production ML models
Required skills
Linux, Kubernetes, Docker, Prometheus, Grafana, Java, Python, Shell, capacity management, fault analysis, performance tuning
Preferred skills
Large internet company experience, SaaS service stability experience, disaster recovery planning
Technologies
Kubernetes, Docker, Chart, Prometheus, Grafana, Linux, Java, Python, Shell
Responsibilities
Design stability solutions for online services including prevention, loss mitigation, degradation, and capacity management; Lead implementation of reliability automation systems; Design deployment architectures for model products; Implement monitoring systems and disaster recovery plans; Optimize large-scale ML model online prediction systems.
Seniority
Senior, hands-on IC