RAG development services
RAG development: AI answers grounded in your documents, with citations
Retrieval-augmented generation (RAG) lets a language model answer from your own contracts, manuals, tickets or knowledge base instead of from memory. Done well, every answer points back to the passage it came from, so your users can check it in one click.
I am a freelance RAG developer with 18 years of software engineering behind me, from a decade of enterprise Java at HCL to shipping production RAG platforms such as BriefPilot and DocuChat. I build the whole pipeline — ingestion, retrieval, generation, evaluation — and the product around it: auth, billing, multi-tenancy and monitoring.
Problems I solve
Answers that can't be trusted
A chatbot that sounds confident but invents clauses or policies is worse than no chatbot. I design retrieval and prompts so the model answers only from retrieved text and shows the source.
Messy, scanned or structured documents
PDFs with columns, scanned pages, tables and numbered clauses break naive pipelines. I combine text extraction with OCR fallback and chunk on the document's real structure.
Costs that grow with every question
Repeated questions should not cost a fresh model call. Semantic caching and model routing keep inference spend predictable as usage grows.
A demo that never became a product
Many RAG prototypes stall at the notebook stage. I ship them as multi-user products with auth, tenant isolation, quotas, billing and tracing.
What I build
- Document ingestion
- PDF, Word, web pages and scans; text extraction with Tesseract OCR fallback; metadata and versioning so updated documents replace old chunks.
- Structure-aware chunking
- Chunks that follow sections, articles, clauses or headings, so a retrieved chunk is a unit a human would recognise and cite.
- Vector search and retrieval
- Embeddings stored in PostgreSQL with pgvector (or your existing vector database), with filters per tenant, document or permission.
- Cited, streamed answers
- Token-streamed responses over Server-Sent Events, with each claim linked to the source passage and page.
- Model routing and failover
- Claude, GPT-4o and Gemini behind one interface with automatic failover, plus semantic caching for repeat questions.
- Evaluation
- Automated answer-quality and hallucination checks, so you can change prompts, models or chunking and see whether quality went up or down.
Related work

Contract-review RAG platform
BriefPilot
Contract review that cites the exact clause, reconciles amendments and grades every contract against your playbook.
Read the case study →

Multi-tenant enterprise RAG chatbot
DocuChat
A multi-tenant RAG chatbot platform with WhatsApp, Telegram, live human handoff and an 18 KB embeddable widget.
Read the case study →
How we'll work
Step 1
Look at your data first
We start with a sample of real documents and real questions. The data shape decides the chunking and retrieval strategy.
Step 2
Build a measurable baseline
A small evaluation set of questions with expected sources, so quality is a number rather than an impression.
Step 3
Ship the pipeline end to end
Ingestion, retrieval and answers working on a preview URL early, then improved in weekly iterations.
Step 4
Productionise
Auth, multi-tenancy, rate limits, cost tracking, monitoring and documentation your team can own.
Technology
- Python
- FastAPI
- LangChain
- PostgreSQL + pgvector
- OpenAI
- Claude
- Gemini
- Tesseract OCR
- Laravel
- Next.js
- Docker
Frequently asked questions
What is RAG and when do I need it?
RAG retrieves the most relevant passages from your own documents and gives them to the language model as context, so the answer is grounded in your data. You need it when answers must reflect private or frequently changing information that a general model does not know.
Can the system show where each answer came from?
Yes. Citations are a core part of how I build RAG: every answer links to the clause, section or page it was drawn from, so users can verify it.
Which language models and vector databases do you use?
Usually Claude, GPT-4o or Gemini with PostgreSQL and pgvector, because it keeps vectors next to your relational data. I can also work with the model provider or vector store you already use.
How do you reduce hallucinations?
By retrieving well-structured chunks, instructing the model to answer only from that context, returning citations, and running automated evaluations that flag unsupported answers before changes go live.
How long does a RAG project take?
It depends on your documents and integrations. After a discovery call I send a fixed-scope proposal with milestones and dates, and you see working software on a preview URL every week.
Tell me what you're building
A short description of the problem and your data is enough. I'll reply with questions and a time for a call.
Start a project