← Browse all jobs

F

Lead Architect

Fractal · Gurugram, India

FULL TIME

Job Description

<p>It's fun to work in a company where people truly BELIEVE in what they are doing! </p><p><b>Role overview: </b> </p><p>We’re building a next-gen LLMOps team at Fractal to industrialize GenAI implementation and shape the future of GenAI engineering. This is a hands-on technical leadership role for AI engineers with strong ML and DevOps skills — ideal for those who love building scalable systems from the ground up. You will be designing, deploying, and scaling GenAI and Agentic AI applications with robust lifecycle automation and observability. </p><p><b>Required Qualifications: </b> </p><li>10 - 14 years of experience in working on ML projects that includes product building mindset, strong hands on skills, technical leadership, leading development teams </li><li>Model development, training, deployment at scale, monitoring performance for production use cases </li><li>Strong knowledge on Python, Data Engineering, FastAPI, NLP </li><li>Knowledge on Langchain, Llamaindex, Langtrace, Langfuse, LLM evaluation, MLFlow, BentoML </li><li>Should have worked on proprietary and open-source LLMs </li><li>Experience on LLM fine tuning including PEFT/CPT </li><li>Experience in creating Agentic AI workflows using frameworks like CrewAI, Langraph, AutoGen, Symantec Kernel </li><li>Experience in performance optimization, RAG, guardrails, AI governance, prompt engineering, evaluation, and observability </li><li>Experience in GenAI application deployment on cloud and on-premises at scale for production using DevOps practices </li><li>Experience in DevOps and MLOps </li><li>Good working knowledge on Kubernetes and Terraform </li><li>Experience in minimum one cloud: AWS / GCP / Azure to deploy AI services </li><li>Team player with excellent communication and presentation skills </li><p><b>Must have skills: </b> </p><li>Product thinking that includes ideation, prototyping, and scale internal accelerators for LLMOps </li><li>Architect and build scalable LLMOps platforms for enterprise-grade GenAI systems </li><li>Design and manage end-to-end LLM pipelines from data ingestion and embedding to evaluation and inference </li><li>Drive LLM-specific infrastructure<b>: </b> memory management, token control, prompt chaining, and context optimization </li><li>Lead scalable deployment frameworks for LLMs using Kubernetes and GPU-aware scaling </li><li>Build agentic AI operations capabilities including agent evaluation, observability, orchestration and reflection loops </li><li>Guardrails &amp; Observability: Implement output filtering, context-aware routing, evaluation harnesses, metrics logging, and incident response </li><li>Platform Automation for LLMOps: Drive end-to-end automation with Docker, Kubernetes, GitOps, DevOps, Terraform, etc. </li><p><b>Product Thinking </b>: Ideate, prototype, and scale internal accelerators and reusable components for LLMOps </p><p><b>GenAI Engineering </b>: Productionize LLM-powered applications with modular, reusable, and secure patterns </p><p><b>Pipeline Architecture </b>: Create evaluation pipelines — including prompt orchestration, feedback loops, and fine-tuning workflows </p><p><b>Prompt &amp; Model Management </b>: Design systems for versioning, AI governance, automated testing, and prompt quality scoring </p><p><b>Scalable Deployment </b>: Architect cloud-native and hybrid deployment strategies for large-scale inference </p><p><b>Guardrails &amp; Observability </b>: Implement output filtering, context-aware routing, evaluation harnesses, metrics logging, and incident response </p><p><b>DevOps &amp; Platform Automation </b>: Drive end-to-end automation with Docker, Kubernetes, GitOps, Terraform, etc. </p><p><b>Must-Have Technical Skills </b> </p><li><b>LLMOps frameworks </b>: LangChain, MLflow, BentoML, Ray, Truss, FastAPI </li><li><b>Prompt evaluation and scoring systems </b>: OpenAI evals, Ragas, Rebuff, Outlines </li><li><b>Cloud-native deployment </b>: Kubernetes, Helm, Terraform, Docker, GitOps </li><li><b>ML pipeline </b>: Airflow, Prefect, Feast, Feature Store </li><li><b>Data stack </b>: Spark/Flink, Parquet/Delta, Lakehouse patterns </li><li><b>Cloud </b>: Azure ML, GCP Vertex AI, AWS Bedrock/SageMaker </li><li><b>Languages </b>: Python (must), Bash, YAML, Terraform HCL (preferred) </li><p>If you like wild growth and working with happy, enthusiastic over-achievers, you'll enjoy your career with us! </p><p></p><p><b><b>Hiring Related Queries </b> </b> </p> <p></p><p>This inbox does not process resume submissions. All applications must be made through posted job openings </p><p>Not the right fit? Let us know you're interested in a future opportunity by clicking in the top-right corner of the page or create an account to set up email alerts as new job postings become available that meet your interest! </p>

Details

CompanyFractal
LocationGurugram, India
TypeFULL TIME
Nichegeneral

Similar Jobs

m

SAP ABAP Consultant.

msg global solutions

N

Sr. Area Process Management Professional

Novo Nordisk

M

Wedding Specialist

Marriott International

I

Epic OpTime & Anesthesia Analyst

Interscripts, Inc.

2

Carpenter

2coms