Senior Deep Learning Systems Engineer, Datacenters
Core
Analyze performance and power consumption of deep learning applications on datacenter-class hardware to influence datacenter design and optimization.
Role type
Senior Deep Learning Systems Engineer
Builds
High-performance datacenter architectures and software infrastructure for AI workloads
Domain
Datacenter hardware, System Software, Deep Learning
Deliverable
production ML models
Required skills
System software development, GPU kernel programming (CUDA), Deep Learning frameworks (PyTorch, TensorFlow), C/C++ programming, Python programming, Computer system architecture analysis, Performance modeling and profiling
Preferred skills
Silicon architecture knowledge, Containerization platforms (Docker), Datacenter workload managers (Slurm), Performance monitoring tools (perf, gprof, nvidia-smi, dcgm)
Technologies
CUDA, PyTorch, TensorFlow, Linux, C++, Python, Docker, Slurm, perf, gprof, nvidia-smi, dcgm
Responsibilities
Develop software infrastructure to characterize and analyze Deep Learning applications, Evolve cost-efficient datacenter architectures for Large Language Models, Work with experts to develop analysis and profiling tools, Analyze system and software characteristics of DL applications, Develop methodologies to measure key performance metrics and estimate efficiency improvement
Seniority
Senior, hands-on IC