Staff Software Engineer - AI Research Infrastructure
Core
Design and build services to schedule, orchestrate, and observe large-scale training and inference experiment workloads across thousands of GPUs for Databricks AI Research.
Role type
Staff Software Engineer (AI Research Infrastructure)
Builds
Research stack, job submission/scheduling systems, monitoring tools, and CI/testing infrastructure for research code.
Domain
AI Research Infrastructure / Distributed Systems / Cloud Computing
Deliverable
infrastructure
Required skills
distributed systems design, large-scale backend services, systems programming (C++, Rust, Go, Java, Scala), cluster schedulers, resource managers, job orchestration (Kubernetes, Slurm, Ray), ML training/inference workflows, operational excellence, technical mentorship
Preferred skills
experience with HPC clusters, GPU fleets, cloud providers, translating research needs to infra realities
Technologies
Kubernetes, Slurm, Ray, C++, Rust, Go, Java, Scala
Responsibilities
Design and implement infrastructure for large-scale experiments and model training; Enable rapid iteration from idea to experiment via powerful abstractions; Create tooling to improve research developer productivity; Influence the long-term roadmap for research computation; Serve as a technical mentor for other engineers.
Seniority
Staff, hands-on IC with mentorship