Ingest
Connectors, scrapers and APIs with retries, backoff and schema-drift guards.
- REST
- Async queues
- Redis
Senior AI Engineer · LLM Systems & RAG Specialist. Production RAG, multi-agent orchestration and the backend infrastructure that keeps them measurable, traced, evaluated and costed, not guessed at.
Currently building
ThreeZinc AI Platform
Shipped for teams and brands including












The uncomfortable part
A demo that works once isn't a system. The real work starts when retrieval drifts, APIs fail, schemas change, and costs spiral. That's where I've spent the last three years—building retrieval pipelines, agent workflows, tracing, evaluations, and cost controls that make AI reliable in production.
0%+
RAGAS faithfulness
up from 62% in production
0%
lower inference latency
provider routing + caching
0%
fewer pipeline failures
fault-tolerant async ingestion
0%
LLM cost reduction
per-query, at sustained traffic
Every number here is measured,not estimated — pulled from RAGAS eval runs, LangSmith traces and provider billing on live production traffic.
I own all six. That ownership is the difference between an AI feature that demos well and one that survives a quarter of real traffic.
Connectors, scrapers and APIs with retries, backoff and schema-drift guards.
Semantic chunking, embeddings and metadata that retrieval can actually filter on.
Agent graphs with explicit state, tool use and structured outputs — no prompt soup.
RAGAS harnesses on faithfulness, context precision and answer relevancy, in CI.
Traces on every node — tokens, latency, failure modes — plus guardrails and alerts.
Request routing, caching, prompt compression and quantization where it pays off.
Four things people hire me for, and the stack each one runs on. Not a list of logos — tooling that has carried real production traffic.
We really liked your examples, but there's one more thing — probably the most important:
Semantic chunking, cross-encoder reranking and metadata filtering — tuned against RAGAS, not vibes. Pinecone, ChromaDB, FAISS and pgvector in production at ~5K daily queries.
LangGraph-style agent graphs wired to real connectors — Instagram, LinkedIn, internal APIs — with explicit handling for rate limits, scraping edge cases and schema drift.
Distributed Node.js and FastAPI backends, async job queues for compute-heavy inference, intelligent request routing, response caching and prompt compression.
LangSmith tracing across every agent node, Arize AI model monitoring, RAGAS evaluation harnesses and A/B tests on prompt variants and retrieval strategies.
Models, retrieval, storage, runtime, telemetry — a dozen moving parts, each with its own failure mode. Somebody has to own the place they all meet. On the last two teams, that was me.
Pick a system to see how it actually runs — the diagram animates the same path production traffic takes.
How it runs
Connectors land raw posts; the queue absorbs rate limits and schema drift.
Multi-Agent Orchestration System
Social connectors feed a supervisor that routes work to specialist agents, with every hop traced.
Diagrams animate the real data path — packets follow the same wires the system does. Dashed lines are telemetry, not traffic.
Full-stack products, developer tools and 3D web experiments — the range behind the specialisation.
Shipped across health-tech, real estate, ed-tech, recruitment, export and D2C commerce — handed over and still running.
Generated cast, generated worlds, cut and scored end to end. The same appetite for pipelines, pointed at something that has to hold an audience instead of a p95.
Long-form cuts on the channel.
Posted as @thinkai.io.
Notes on shipping AI features.
Tools have changed, but responsibility hasn't.
On agile practice and what engineering excellence actually costs day to day.
#SoftwareEngineering#Agile

No matter the size or the sector, the common thread is an AI feature that has to work for real users, on real data, with somebody accountable for the numbers.
Gaup Media Pvt Ltd · New Delhi, India
Feb 2026 — Present
Digitally Next · New Delhi, India
Apr 2024 — Feb 2026
Happymonk AI · Bengaluru, India · Promoted from internship
Mar 2023 — Apr 2024
Happymonk AI · Bengaluru, India · Internship
Dec 2022 — Mar 2023
Tap any of these for what I've actually done with it. No proficiency bars — a percentage next to a logo has never told anyone anything.
No tool here is on the list because it looked good on a slide. Each one earned its place solving a specific production problem.
The orchestration layer. Agent graphs with explicit state beat one long prompt every time — they can be traced, tested and reasoned about.
If it isn't traced, it isn't real. Tokens, latency and failure modes go on a dashboard before the feature goes to users.
Retrieval quality is a data problem before it's a model problem. Chunking, metadata and reranking do most of the heavy lifting.
The unglamorous half. Async queues, request routing and honest API contracts are what keep an LLM feature from falling over under load.
The AI has to land somewhere. Shipping the interface myself means the contract between model and UI never gets lost in translation.
Where it all runs. Containers, managed Postgres and object storage — boring, predictable, easy to hand over.
If yours isn't here, ask it directly — I answer anything with a concrete problem attached.
I own the AI layer end to end. That means architecture and retrieval design, prompt engineering, RAG pipeline tuning, deployment, evaluation harnesses, observability and cost control — plus the backend the whole thing sits on. In my last two roles I was the sole owner of that layer, which is a very different job from bolting an API call onto an existing product.
Open to senior AI / LLM engineering projects
Send me the problem — the retrieval that keeps hallucinating, the pipeline that fails silently, the bill nobody can explain. I'll tell you straight whether I'm the right person for it.