Software Development Engineer.5
Core
Design and build enterprise AI platforms to automate operations, manage infrastructure, and support AI-powered decision-making at scale.
Role type
Senior AI Reliability Engineer (AI Platform Engineering & MLOps)
Builds
AI platforms, AI Agents, MCP servers, secure AI gateways, RAG pipelines, AI copilots, and production-grade inference systems.
Domain
AI Engineering, MLOps, Cloud Infrastructure, Platform Engineering
Deliverable
production ML models | infrastructure
Required skills
Python, Golang, LLMs, AI Agents, MLOps, Kubernetes, AWS, Terraform, Distributed Systems, API Design, Observability
Preferred skills
LangGraph, LangChain, MCP, Vector Databases, Agentic AI, Prompt Engineering, AI Evaluation Frameworks, Kubeflow, Ray, KServe, BentoML, Triton Inference Server, NVIDIA ecosystem, Prometheus, Grafana, OpenTelemetry, Jaeger, New Relic, Elastic
Technologies
Kubernetes, AWS, Terraform, Docker, Helm, ArgoCD, GitOps, Prometheus, Grafana, OpenTelemetry, Jaeger, New Relic, Elastic, LangGraph, LangChain, MLflow, Kubeflow, Ray, KServe, BentoML, Triton Inference Server
Responsibilities
Design and build enterprise AI platforms for deploying and scaling LLM-powered applications; Develop AI Agents and MCP servers to automate engineering workflows; Build secure AI gateways and RAG pipelines; Create AI copilots for incident response and operational workflows; Implement MLOps pipelines for model lifecycle management and inference; Build observability platforms for AI systems including latency, cost, and hallucination detection; Develop autonomous operations systems for incident investigation and remediation.
Seniority
Senior, hands-on IC