CareerPlanGet AI match score →

Application Software Engineer, Inference

Palo Alto - 1530💼 Full-time💰 $135,000–$135,000🗓 2026-07-09 → 2026-07-31

Core

Design and optimize large-scale AI inference platforms to serve mission-critical models for SpaceX's launch vehicles and Starlink systems.

Role type

Senior IC application software engineer (LLM inference systems)

Builds

High-throughput, low-latency distributed inference infrastructure for internal SpaceX AI applications

Domain

Aerospace / Large Language Model Inference

Deliverable

production ML models

Required skills

Rust, C++, distributed systems design, low-level GPU optimization, model serving frameworks, system observability, CI/CD

Preferred skills

LLM inference engines (SGLang, vLLM, TensorRT-LLM), speculative decoding, agent SDKs, Kubernetes, gRPC, Python/Go

Technologies

SGLang, vLLM, TensorRT-LLM, Triton, PostgreSQL, ClickHouse, MongoDB, Docker, Kubernetes, gRPC

Responsibilities

Architect scalable distributed infrastructure for model serving; Optimize latency and throughput via GPU kernels and quantization; Build high-concurrency serving systems with 100% uptime; Own end-to-end components like request routing and SDK development; Benchmark and accelerate inference engines; Develop custom tracing and debugging tools; Create CI/CD infrastructure for deployment

Seniority

Senior, hands-on IC

Sourced via greenhouse · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Greenhouse ↗