About the role
structured by ORIOnsite: Gurgaon | 6 Days Working Experience: 2+ years Important: We’re specifically looking for someone who has spent their experience building and shipping real AI agents, agentic systems, or RAG systems used by real users at scale. This is not a general software engineering role .
What you will do
- Design and build RAG pipelines — chunking, embeddings, vector search, re-ranking
- Build agents and orchestration logic from scratch (no framework crutch) — tool-use, multi-step reasoning, state management
- Train and fine-tune models where off-the-shelf LLMs fall short — dataset curation, fine-tuning, LoRA, evaluation of trained models
- Set up eval harnesses and benchmarks to catch regressions before users do
- Implement guardrails, hallucination detection, and prompt-injection defenses
What they are looking for
- 2+ years of hands-on experience building and shipping production AI/LLM systems
- Strong experience building AI agents, agentic workflows, or RAG systems used by real users at scale
- Strong backend fundamentals — API design, DB modeling, Python
- End-to-end RAG fluency — embeddings, vector DBs, re-ranking
- Ability to design and build agent/orchestration systems from first principles, without relying on frameworks like LangChain/LangGraph
- Hands-on experience with model training/fine-tuning — LoRA, dataset curation, evaluation
- Docker, Kubernetes, and MLOps practices — CI/CD for models, versioning, monitoring
- Redis or similar for caching/session state
Before you apply
- Must have 2+ years of experience predominantly hands-on building and shipping production AI agents, agentic systems, or RAG systems.
- Position requires onsite working in Gurgaon, 6 days a week.
Full posting text
Onsite: Gurgaon | 6 Days Working
Experience: 2+ years
Important: We’re specifically looking for someone who has spent their experience building and shipping real AI agents, agentic systems, or RAG systems used by real users at scale. This is not a general software engineering role . If your experience is primarily in backend/software engineering and you’ve only recently started working with LLMs or AI, please don’t apply .
We’re looking for someone whose ~2+ years of experience is predominantly hands-on AI engineering - building production agents, RAG pipelines, LLM systems, orchestration, evaluation, and related infrastructure.
Tech Stack: Python, LLM APIs, RAG, Vector DBs, Model Training/Fine-tuning, Docker, Kubernetes, Redis, MLOps, FastAPI, Postgres, AWS
About the Role Build and ship production LLM systems - not prototypes. You'll own the full lifecycle: retrieval architecture, agent/orchestration design, model training/fine-tuning, evaluation, and deployment for features with real users and real failure consequences.
Responsibilities Design and build RAG pipelines — chunking, embeddings, vector search, re-ranking
Build agents and orchestration logic from scratch (no framework crutch) — tool-use, multi-step reasoning, state management
Train and fine-tune models where off-the-shelf LLMs fall short — dataset curation, fine-tuning, LoRA, evaluation of trained models
Set up eval harnesses and benchmarks to catch regressions before users do
Implement guardrails, hallucination detection, and prompt-injection defenses
Optimize for cost, latency, and context-window efficiency — caching, streaming, Redis
Design backend APIs and data models independent of the AI layer
Containerize and deploy services with Docker/Kubernetes; build MLOps pipelines for model versioning, monitoring, and rollout
Run A/B tests and iterate on prompt/model performance
Maintain observability and tracing across LLM pipelines
Required Skills 2+ years of hands-on experience building and shipping production AI/LLM systems
Strong experience building AI agents, agentic workflows, or RAG systems used by real users at scale
Strong backend fundamentals — API design, DB modeling, Python
End-to-end RAG fluency — embeddings, vector DBs, re-ranking
Ability to design and build agent/orchestration systems from first principles, without relying on frameworks like LangChain/LangGraph
Hands-on experience with model training/fine-tuning — LoRA, dataset curation, evaluation
Docker, Kubernetes, and MLOps practices — CI/CD for models, versioning, monitoring
Redis or similar for caching/session state
Seniority level: Entry level
Employment type: Full-time
Job function: Engineering and Information Technology
Industries: Legal Services