Senior ML Infrastructure Engineer, Inference Platform
Core
Design and build cloud-agnostic, reliable, and cost-efficient platforms for serving state-of-the-art machine learning models for autonomous vehicles and AI-driven products.
Role type
Senior ML Infrastructure Engineer (Inference Platform)
Builds
Robust ML inference service supporting real-time, batch, and experimental inference needs for GM's AI teams.
Domain
Automotive AI / Distributed Systems / High-Performance Computing
Deliverable
production ML models
Required skills
Distributed systems design, Python or C++, ML inference frameworks (Triton, RayServe, vLLM), GPU optimization, system architecture, observability, technical leadership
Preferred skills
Zero-to-one platform building, API design for ML workflows, Ray framework, large-scale data processing, telemetry integration, hardware acceleration
Technologies
Triton, RayServe, vLLM, Ray, GPUs (B200, H100, A100)
Responsibilities
Design and implement core platform backend software components; Lead technical decision-making on model serving strategies and auto-scaling; Drive development of monitoring and observability; Proactively research and integrate state-of-the-art serving frameworks; Contribute to open source projects.
Seniority
Senior, hands-on IC with technical leadership