Open to AI Engineer roles

Pankaj Sanger

> AI Engineer · LLM Systems · Agentic AI · RAG pipelines

I build production-ready AI systems from development to deployment.

Two years of professional AI engineering, built on five years of analytics. Designing and deploying LLM, ML, and deep learning applications with Docker and AWS.

01About Me

I build end-to-end AI systems that solve real business problems, from LLM-powered applications and intelligent retrieval systems to machine learning and deep learning solutions. My work spans data ingestion, model development, orchestration, deployment, and production infrastructure.

One example is the Amazon Intelligence Assistant, where I built the complete system: product scraping, retrieval pipeline, LLM orchestration, guardrails, React frontend, Dockerized deployment with a headless browser, and hosting on AWS EC2. Beyond LLM applications, I've also developed machine learning and deep learning models for real-world use cases, focusing on scalable, production-ready solutions.

Professionally, at Mobilnxt, I work on AI and social intelligence solutions for Dabur India, building automated pipelines for transcript analysis, review mining, sentiment analysis, consumer insight generation, and large-scale market intelligence across multiple brands. These systems process live data continuously and deliver actionable insights for business and product teams.

My foundation in data analytics complements my AI engineering work by ensuring every system I build is measurable, reliable, and focused on delivering business impact, not just technical capability. The details of my professional experience and projects are covered in the sections below.

Currently
AI & Social Intelligence at Mobilnxt (on the Dabur India account)
Looking for
A full-time AI Engineering role across Generative AI, LLMs, Machine Learning, Deep Learning, and AI applications, building production-ready systems that solve real-world problems.
Based in
NCR, India
Pankaj Sanger, AI Engineer

02Selected work

Three systems I designed, built and debugged end to end. The write-ups go into the decisions I had to make and what I got wrong first, rather than listing the stack again.

012026

Amazon Product & Review Intelligence

Scrapes Amazon listings and reviews, then answers questions about them in plain English. A self-query retriever pulls the constraint out of your sentence and applies it as a real filter. Built end to end, guardrailed, containerised and running on an EC2 box.

docker image · aws ec2embedembedif allowedanswerSession Cookiesauto-loginScrapersURL · Excel · keywordSQLiteupsert on rescrapeChromaproduct_detailsSelf-QueryNL → filtersChromacustomer_reviewsQuestionInput Guardon-topic · jailbreakLLM Answergrounded + citedOutput Guardredact PII · fix toxic
  • Three ingestion modes: single URL, bulk Excel upload, live keyword search
  • SQLite upserts keep repeat scrapes idempotent
  • Dual Chroma collections with independently tuned self-query retrievers
  • Every answer cites the products it drew from
  • Separate Ask Products and Ask Reviews surfaces, one per collection
  • Migrated from a Streamlit prototype to a React + FastAPI front end
  • Input guards reject off-topic and prompt-injection attempts before retrieval
  • Output guards redact reviewer PII and strip toxic sentences from answers
  • Multi-stage Docker build: Node frontend stage, Python backend, headless Chromium
  • Deployed on AWS EC2 via Docker Compose, with bind-mounted state and log rotation
RAGSelf-Query RetrievalChromaDBGuardrailsDockerAWS EC2FastAPIReact
022025 — 2026

YouTube Transcript Intelligence Studio

Searches YouTube, pulls transcripts through a two-source fallback, translates everything into English, and makes the whole library searchable with FAISS.

on missembedretrieveKeyword + FiltersYouTube Data APImetadata + statsTranscript APIApify ActorfallbackNormalizededup → ENFAISS900-tok chunksAnswertop-20 chunks
  • Dual-source transcript retrieval with automatic Apify fallback
  • Korean, Japanese and Hindi transcripts normalized to English before embedding
  • Duration and transcript-availability filters applied at collection time
  • Four tabs — Collect, Review, Library, Ask AI — over one shared dataset
  • Channel-level aggregation across the saved library
  • Persisted FAISS index decoupled from dataset collection
RAGFAISSOpenAI EmbeddingsYouTube Data APIApifyStreamlit
032025 — 2026

Agentic Research Assistant

A LangGraph agent that works out for itself whether a question needs live data, then reaches YouTube and Reddit through an MCP server it calls mid-conversation.

tooldirectresultsstateUser TurnAgent NodeLangGraphConditional Edgetool or answer?FastMCP ServerYouTube ToolReddit ToolResponseSQLite Memorycheckpoint
  • The agent decides per turn whether to answer directly or call a tool
  • YouTube and Reddit search served over the Model Context Protocol
  • Multi-turn chat with session memory checkpointed to SQLite
  • Tools reusable by any MCP-compatible client
LangGraphFastMCPTool CallingAgentsSQLiteStreamlit

03Smaller builds

Machine learning, deep learning and agent orchestration, including a sentiment model trained on review data from a company I was working at.

04Toolkit

Grouped by the layer of the stack they sit in.

05Experience

Where the AI work happens now, and the five years of analysis it was built on top of.

Data Analyst — AI & Social Intelligence

Mobilnxt Pvt Ltd · client: Dabur India

Aug 2024Presentnow

This is where I build LLM systems against live data. I own social listening and competitive intelligence for a 25-brand FMCG portfolio, and I've spent most of the last two years replacing the manual research behind it with pipelines that run themselves.

  • Built LLM pipelines with OpenAI, Claude, Mistral, Hugging Face and Ollama that handle transcript processing, sentiment analysis and review mining. Research that used to take a week of reading now runs unattended against live sources.
  • Wrote the ingestion and monitoring layer in Python and Pandas, tracking sentiment, competitor activity and reputation risk across Instagram, Facebook, YouTube, Reddit, Google Trends and news, with alerting on top.
  • Turn that output into 70+ weekly reports for senior leadership on campaign performance, consumer behaviour, influencer activity and category trends — the part that decides whether any of the engineering mattered.
  • Produced NPD and competitive intelligence covering share of voice, sentiment and emerging category trends for the Digital Marketing and Innovation teams.
PythonPandasOpenAIClaudeMistralHugging FaceOllamaBrandwatchLooker Studio

Data Analyst

Livpure Pvt Ltd

Mar 2019Jul 2023

Field operations and customer analytics for a service business covering all of India. Several of the changes I recommended were rolled out nationally.

  • Got field engineer productivity from 4 service requests a day to 5.5 across a 350-engineer national workforce, by analysing how calls were routed and where capacity was sitting idle. That's a 37.5% gain.
  • Cut rental delivery turnaround by 30% after working through the fulfilment pipeline and finding where it was actually stalling. The fixes went out nationally.
  • Helped the water purifier line reach a 4.8-star Amazon rating, the best in its segment, by reading reviews systematically and getting what they said back to the teams who could act on it.
  • Handled data coming from CRM, Molten and the field teams, and presented recommendations to senior management directly rather than through a layer.
PythonSQLExcelCRMDashboarding

B.Tech, Mechanical Engineering

3rd University Rank at college level

79% aggregate

06 — Contact

Hiring for an AI Engineering role?

I'm looking for a team putting AI in front of actual users. Generative AI, LLM engineering, ML or deep learning: I care less about which layer than about whether it ships. If that sounds like yours, I'd like to hear about it. Email reaches me fastest.