Skip to content
All projects
Flagship systemAI/MLRAGEvals2025

Production RAG Chat System

LangChain + GPT + Pinecone

85%+

RAGAS faithfulness

<800ms

p95 retrieval @ 10K docs

$0.003

per query, cached

Overview

A production RAG chatbot at 85%+ RAGAS faithfulness, with a full evaluation harness and quantization experiments.

A production RAG chatbot delivering 85%+ RAGAS faithfulness — up from a 62% baseline — through semantic chunking, cross-encoder reranking and metadata filtering. Benchmarked at under 800ms p95 retrieval latency across 10K documents at roughly $0.003 per query with caching. The repo ships a full RAGAS evaluation framework covering context precision, faithfulness and answer relevancy, with documented eval runs, test cases and a LoRA fine-tuning experiment for domain-specific QA in the evals/ folder. Inference optimisation was explored via GGUF/Q4_K_M quantization on Mistral 7B for cost-sensitive deployments, with methodology written up alongside the results.

Stack

  • LangChain
  • GPT
  • Pinecone
  • RAGAS
  • Mistral 7B
  • LoRA
  • Python
Year
2025
Focus
AI/ML, RAG, Evals
Role
Architecture & delivery

More work

All projects
01
Flagship2026

ThreeZinc AI Platform

Multi-Agent Orchestration System

A multi-agent platform orchestrating Instagram and LinkedIn connectors for automated content generation and campaign analytics.

  • LangGraph
  • LangSmith
  • Node.js
  • +4
02
Flagship2025

Multi-modal Generation Platform

Image, Video & 3D Inference Backend

A unified backend fronting image, video and 3D generative models, with async queues for compute-heavy inference.

  • FastAPI
  • Python
  • Redis
  • +3
CIAO! interface
2024

CIAO!

3D Beverage Brand Experience

A fully responsive 3D brand site for a kombucha label, built with Three.js and GSAP scroll choreography.

  • Vite
  • React
  • Three.js
  • +1