85%+
RAGAS faithfulness
<800ms
p95 retrieval @ 10K docs
$0.003
per query, cached
Overview
A production RAG chatbot at 85%+ RAGAS faithfulness, with a full evaluation harness and quantization experiments.
A production RAG chatbot delivering 85%+ RAGAS faithfulness — up from a 62% baseline — through semantic chunking, cross-encoder reranking and metadata filtering. Benchmarked at under 800ms p95 retrieval latency across 10K documents at roughly $0.003 per query with caching. The repo ships a full RAGAS evaluation framework covering context precision, faithfulness and answer relevancy, with documented eval runs, test cases and a LoRA fine-tuning experiment for domain-specific QA in the evals/ folder. Inference optimisation was explored via GGUF/Q4_K_M quantization on Mistral 7B for cost-sensitive deployments, with methodology written up alongside the results.
Stack
- LangChain
- GPT
- Pinecone
- RAGAS
- Mistral 7B
- LoRA
- Python
- Year
- 2025
- Focus
- AI/ML, RAG, Evals
- Role
- Architecture & delivery
More work
All projectsThreeZinc AI Platform
Multi-Agent Orchestration System
A multi-agent platform orchestrating Instagram and LinkedIn connectors for automated content generation and campaign analytics.
- LangGraph
- LangSmith
- Node.js
- +4
