SRE高级软件工程师(云原生与AI Agent方向)-Data
Core
Design and maintain cloud-native infrastructure and business systems, building AI-driven SRE agents to automate operations and ensure high availability.
Role type
Senior Site Reliability Engineer (SRE) specializing in Cloud-Native and AI Agents
Builds
Cloud-native infrastructure, AI-powered SRE agents (LLM/RAG-based), and automated stability platforms
Domain
Cloud-native computing, AIOps, Large Language Models (LLM), RAG
Deliverable
production ML models
Required skills
Distributed systems architecture, Linux kernel internals, Go, Kubernetes, Microservices, System-level debugging, LLM integration, RAG implementation, FinOps
Preferred skills
Service Mesh, OpenTelemetry, Prometheus, Capacity planning, Natural language interaction design
Responsibilities
Lead service lifecycle management and capacity planning; Architect and develop AI agents for fault diagnosis and self-healing; Build platform tools to automate toil and emergency responses; Analyze performance bottlenecks in high-concurrency systems; Optimize compute and storage costs via elastic scaling and hybrid deployment.
