Principal Engineer, Machine Learning, SMAI
Core
Architecting and executing large-scale custom model training and fine-tuning jobs, and designing autonomous AI Agents to automate complex manufacturing workflows for Micron's memory solutions market.
Role type
Principal Machine Learning Engineer (Agentic AI & GenAI)
Builds
Scalable AI/ML solutions, custom GenAI models, and autonomous AI Agents for manufacturing processes.
Domain
Semiconductor manufacturing / Memory solutions / Generative AI
Deliverable
production ML models
Required skills
GPU architecture management, distributed training strategies (FSDP, DeepSpeed, Megatron-LM), fine-tuning LLMs (PEFT, LoRA, QLoRA), inference engine optimization (vLLM, TensorRT-LLM), AI Agent framework development (LangChain, LangGraph, CrewAI), multi-step reasoning and planning, CI/CD pipeline creation for ML, data structure optimization in cloud data warehouses.
Preferred skills
Computer Science or Statistics background, experience with Snowflake and Google Cloud platforms, mixed-precision techniques (FP16/BF16), profiling and debugging GPU performance bottlenecks.
Technologies
PyTorch, Nsight Systems, Snowflake, Google Cloud, LangChain, LangGraph, CrewAI, vLLM, TensorRT-LLM, DeepSpeed, Megatron-LM, FSDP, NVLink.
Responsibilities
Architect and execute large-scale custom model training and fine-tuning jobs on multi-node, multi-GPU clusters; Optimize training throughput and memory efficiency using distributed training strategies; Design and develop autonomous AI Agents capable of multi-step reasoning and tool execution; Implement Agentic frameworks to orchestrate LLM interactions with internal APIs and databases; Profile and debug GPU performance bottlenecks to maximize hardware utilization; Build and maintain data/solution pipelines that feed machine learning models and GenAI applications; Design and optimize data structures in data management systems to enable AI/ML solutions.
Seniority
Principal, hands-on IC