CareerPlanGet AI match score →

Machine Learning Engineer

💼 Full-time🗓 2026-07-26

Core

Design and build the high-level architecture of a next-generation MLOps platform for deploying, scaling, and monitoring GenAI models (LLMs, speech, vision, diffusion) across cloud and on-prem environments.

Role type

Senior IC Machine Learning Engineer (MLOps Infrastructure)

Builds

Scalable, reliable, GPU-accelerated ML workflows and MLOps infrastructure for a GenAI inference platform.

Domain

Generative AI, Cloud Infrastructure, Distributed Systems

Deliverable

infrastructure

Required skills

System design, Distributed systems, GPU-based ML workloads, Software engineering fundamentals, Infrastructure-as-code, Cloud platforms, ML fundamentals, Multi-step pipeline orchestration, Linux internals, Networking, Performance tuning

Preferred skills

Modern inference stacks (TensorRT, Triton, vLLM/TGI, SGLang), Model optimization, CUDA concepts, LLM/VLM/ASR pipeline experience, CI/CD, Docker, High-availability system design

Technologies

Terraform, Ansible, AWS, GCP, Azure, Linux, Docker, GitHub workflows

Responsibilities

Design and implement core architecture for GPU-accelerated workloads at scale, Formalize heterogeneous ML workloads and build orchestration abstractions, Build internal systems for continuous deployment across multi-cloud environments, Create frameworks for reliability and observability, Develop internal tooling for benchmarking and model deployment, Troubleshoot complex systems and distributed workloads

Seniority

Senior, hands-on IC

Rewrite
## About the company Simplismart is a GenAI inference platform to deploy, scale, and monitor any GenAI model (LLMs, speech, vision, or diffusion) across cloud or on-prem. Built for strict SLAs, enterprise-grade security, and full observability. Its modular design lets you optimize for cost or latency or auto-select the best topology per workload. ## About the role In this role, you will design and build the high-level architecture of Simplismart's MLOps platform from the ground up, enabling scalable, reliable, and GPU-accelerated ML workflows across the product ecosystem. ## Responsibilities - Design and implement the core architecture of a next-generation MLOps platform capable of running diverse GPU-accelerated workloads at scale. - Formalise and standardize heterogeneous ML workloads—including LLM/VLM/ASR/diffusion pipelines and build orchestration abstractions for them. - Build internal systems for continuous deployment of services, modules, and model pipelines across multi-cloud and hybrid environments. - Create frameworks for high reliability, observability, and fault-tolerance for mission-critical inference, training, and data pipelines. - Collaborate closely with Applied ML and Core ML teams to improve system reliability, latency, and cost efficiency. - Develop internal tooling to benchmark, evaluate, and deploy models quickly and consistently. - Ship production-grade code and infrastructure using strong engineering fundamentals and test-driven development (TDD), aligning with Simplismart's engineering culture. - Troubleshoot complex systems, performance bottlenecks, GPU behavior, and distributed workloads. ## Requirements - Deep technical expertise in system design, distributed systems, and GPU-based ML workloads. - Strong software engineering fundamentals (data structures, APIs, testing, debugging). - Experience with infrastructure-as-code (Terraform, Ansible) and cloud platforms (AWS/GCP/Azure). - Strong knowledge of ML fundamentals, model architectures (Transformers, CNNs), and inference behavior. - Ability to build, maintain, and reason about multi-step pipelines (ETL → model → evaluation → deploy). - Strong systems knowledge: Linux internals, networking, performance tuning, GPU memory behavior. - Ability to work independently, own large ambiguous problems, and collaborate across teams. - Excellent communication skills—able to articulate design decisions, tradeoffs, and system impacts clearly. ## Nice to have - Experience with modern inference stacks such as TensorRT, Triton, vLLM/TGI, SGLang. - Exposure to quantization, model optimization, or CUDA concepts. - Hands-on experience with Llama/Mistral, Whisper, or Stable Diffusion pipelines. - Familiarity with CI/CD, Docker, GitHub workflows, and IaC-driven deployments. - Experience designing high-availability or fault-tolerant production systems. ## What we offer - Opportunity to define and lead the brand identity of a fast-growing GenAI company. - Work closely with leadership on high-impact initiatives from global event campaigns to overall storytelling. - Be part of a team that values design as a strategic lever, not just execution. - Competitive compensation and growth opportunities in a high-energy startup environment.
Sourced via wellfound · Listed on CareerPlan, which tracks 70,000+ jobs from 20+ sources.
Apply on Wellfound ↗