Lead Data Engineer (GenAI / LLM Applications)
Core
Design, build, and maintain scalable data platforms and LLM-powered solutions (RAG pipelines, intelligent agents) supporting clinical trial operations and business intelligence.
Role type
Lead Data Engineer (GenAI / LLM Applications)
Builds
Scalable data pipelines, LLM-powered applications, and data architectures for clinical trial operations.
Domain
Healthcare / Clinical Trials / Data Engineering / Generative AI
Deliverable
production ML models | infrastructure
Required skills
Python (Flask, Django, pandas, NumPy), SQL (Oracle, MS SQL Server, PostgreSQL, Snowflake), Cloud platforms (AWS), Workflow orchestration (Apache Airflow), Git/CI-CD, Data modeling, LLM frameworks (LangChain), AI-assisted development tools (GitHub Copilot)
Preferred skills
Clinical trial lifecycle knowledge, Data visualization (Plotly, Power BI), Front-end technologies (HTML5, CSS3, JavaScript), Data analysis/cleansing
Technologies
AWS (S3, EC2, Secrets Manager, Bedrock, Lambda), Snowflake, Apache Airflow, LangChain, GitHub Copilot, Flask, Django, Plotly, Power BI, Oracle, MS SQL Server, PostgreSQL
Responsibilities
Design and develop scalable software architectures and data pipelines; Write clean, reusable Python code; Leverage AI-assisted tools to build LLM-powered solutions; Develop and optimize complex SQL; Implement ETL pipelines; Establish data quality frameworks and governance; Deploy and manage solutions on AWS; Troubleshoot production issues; Partner with stakeholders to translate business needs into technical solutions.
Seniority
Senior, hands-on IC with leadership responsibilities