Research Engineer, AI Models
Core
Research and implement state-of-the-art techniques to accelerate AI inference (quantization, sparsity, distillation, speculative decoding) and build fine-tuning/post-training pipelines for large language models and multimodal systems.
Role type
Research Engineer (AI Models & Hardware Co-Design)
Builds
Production-quality AI inference pipelines and benchmarking frameworks optimized for in-memory computing silicon.
Domain
Artificial Intelligence / Hardware Systems / Inference Optimization
Deliverable
production ML models
Required skills
Quantization, Sparsity, Distillation, Speculative Decoding, Caching Strategies, Model Fine-tuning, Post-training Optimization, Hardware-aware Optimization, Benchmarking Frameworks, System Architecture Analysis
Preferred skills
Compiler optimization, Silicon architecture knowledge, Multimodal system design (via careerplan.io/jobs/4252539009-research-engineer-ai-models-at-enchargeai36)
Technologies
PyTorch, TensorFlow, CUDA, C++, Python, MLIR (implied by hardware/compiler context)
Responsibilities
Research and implement techniques to accelerate AI inference; Partner with hardware/compiler teams to optimize models for silicon; Develop rigorous benchmarking frameworks to characterize tradeoffs between quality, latency, and power.
Seniority
Mid-Senior, hands-on IC
