AI/ML/LLM Systems Engineer - Enterprise AI Platform Engineer
Core
Design and maintain infrastructure for deploying and optimizing large language models (LLMs) and vision models on enterprise-scale AI platforms.
Role type
Senior IC AI/ML Systems Engineer (LLM Infrastructure)
Builds
Scalable, secure, high-performance AI/ML/LLM inference pipelines and platforms for enterprise operations.
Domain
Energy industry + Cloud-native AI/ML infrastructure
Deliverable
production ML models
Required skills
LLM deployment and optimization, Kubernetes (K8s), Docker, OpenShift, Python, SQL, NVIDIA GPU cluster management, inference scaling, distributed computing, SLA/SLO planning, Elasticsearch, PostgreSQL, Haystack, CI/CD pipelines (Git, Bitbucket, Jenkins, ArgoCD), monitoring and dashboarding
Preferred skills
Experience with NVIDIA SuperPods, vector databases, workflow orchestration
Technologies
NVIDIA SuperPods, Kubernetes, Docker, OpenShift, Elasticsearch, PostgreSQL, Haystack, Git, Bitbucket, Jenkins, ArgoCD
Responsibilities
Deploy and manage LLMs/vision models on NVIDIA SuperPods/Cloud; Build scalable inference pipelines using K8s/Docker/OpenShift; Optimize inference performance; Benchmark and evaluate LLMs; Implement LLMOps frameworks with observability; Integrate vector and relational databases; Maintain CI/CD pipelines; Ensure high availability of AI workflows; Develop monitoring and alerting systems.
Seniority
Senior, hands-on IC