Senior Software Engineer - AI Infrastructure Performance Insights & Observability
Core
Designing and optimizing high-performance AI infrastructure systems and observability solutions to support large-scale compute workloads.
Role type
Senior Software Engineer (AI Infrastructure)
Builds
Scalable AI compute platforms and real-time performance monitoring tools
Domain
Cloud Infrastructure / AI Systems / Observability
Deliverable
infrastructure
Required skills
Distributed systems, performance optimization, observability stack design, cloud architecture, system reliability, debugging complex infrastructure issues, API design, container orchestration
Preferred skills
Experience with AI/ML workload patterns, real-time data streaming, multi-tenant system design
Technologies
Kubernetes, Prometheus, Grafana, Go, Python, gRPC, Istio
Responsibilities
Architecting high-throughput data pipelines for infrastructure metrics, developing tools to visualize and analyze system performance bottlenecks, implementing automated alerting and remediation workflows, collaborating with AI teams to optimize resource allocation, maintaining system health and uptime for critical compute clusters, documenting infrastructure standards and best practices