Machine Learning Infrastructure Engineer, GenAI Technology
Core
Design and operate high-performance distributed systems and infrastructure to support large-scale generative AI and machine learning workloads, enabling faster model iteration and reliable end-to-end ML workflows.
Role type
Senior Machine Learning Infrastructure Engineer (GenAI)
Builds
Production-ready GenAI infrastructure, distributed training/inference systems, and automated deployment pipelines for ML models.
Domain
Financial technology / Generative AI / Cloud Infrastructure
Deliverable
infrastructure
Required skills
Distributed systems design, Container orchestration (Kubernetes), Public cloud platforms (AWS/GCP/Azure), ML operations tools (MLflow, Ray, Airflow, Kubeflow, Terraform), Reinforcement learning concepts, Python, Systems-level programming (Go/C++/Rust), Performance profiling and optimization, GPU/accelerator compute management
Preferred skills
None stated
Technologies
Kubernetes, AWS, Google Cloud Platform, Azure, MLflow, Ray, Airflow, Kubeflow, Terraform
Responsibilities
Design and implement high-performance infrastructure for GenAI/ML workloads; Design and operate distributed systems for model training, hyperparameter tuning, inference, and data preprocessing; Collaborate with ML researchers to optimize compute utilization and inference latency; Develop and automate deployment, orchestration, and CI/CD pipelines; Implement observability, monitoring, and cost-management strategies for GPU environments; Evaluate and integrate emerging hardware/software technologies; Drive security, compliance, and operational runbooks for GenAI infrastructure; Troubleshoot and optimize performance across GPU and CPU compute stacks; Document architecture and mentor engineers.
Seniority
Senior, hands-on IC