Staff Software Engineer - AI Research Infrastructure
Core
Designing and building services to schedule, orchestrate, and observe large-scale training and inference experiment workloads across thousands of GPUs for Databricks AI Research.
Role type
Staff Software Engineer (AI Research Infrastructure)
Builds
Research stack, job submission/scheduling systems, monitoring tools, and CI/testing infrastructure for research code.
Domain
AI Research Infrastructure / Distributed Systems / Cloud Computing
Deliverable
infrastructure
Required skills
distributed systems design, large-scale backend services, systems programming (C++, Rust, Go, Java, Scala), cluster schedulers, resource managers, job orchestration, ML training/inference workflows, operational excellence, technical mentorship
Preferred skills
GPU cluster management, cloud provider expertise, experiment management systems
Technologies
Kubernetes, Slurm, Ray, HPC clusters, GPU fleets
Responsibilities
Design and implement infrastructure for large-scale experiments and model training; Enable researchers to run experiments quickly via powerful abstractions; Create tooling to improve research developer productivity; Influence the long-term roadmap for research computation; Serve as a technical mentor for other engineers.
Seniority
Staff, hands-on IC with mentorship