RAG and LLM development
RAG and LLM application development
Retrieval-augmented generation (RAG) lets a large language model answer from your own documents instead of its general training data. Azimuthforge builds RAG systems and LLM applications that cite their sources, respect your access rules, and are scored against an evaluation set before launch.
What RAG is, in one paragraph
Retrieval-augmented generation (RAG) is a way of building AI systems where the model looks things up before it answers. When someone asks a question, the system searches a trusted set of documents, selects the most relevant passages, and gives them to a large language model (LLM) to write an answer that cites those passages. The result is an assistant that answers from your current information, and that can show its working.
How a RAG system works
- Ingest. Documents are collected from their sources (drives, wikis, ticketing tools, databases) and kept in sync as they change.
- Chunk. Each document is split into passages that keep headings and context together, so a passage makes sense on its own.
- Embed and index. Passages are converted to vector embeddings and stored in a search index, alongside keyword search for exact terms such as product codes.
- Retrieve. A question is matched against the index, filtered by the user's access rights.
- Re-rank. The best candidates are re-ordered by relevance, so the model sees the strongest evidence first.
- Generate with citations. The LLM writes an answer using only the retrieved passages and cites each one.
- Evaluate and monitor. Answers are scored against an evaluation set before launch, and production questions are logged to find gaps.
RAG vs fine-tuning vs prompting
| Prompting only | RAG | Fine-tuning | |
|---|---|---|---|
| Uses your private documents | Only what fits in the prompt | Yes, searched on every question | Only what was in the training data |
| Handles changing information | Manually | Yes, re-indexed as documents change | No, needs retraining |
| Shows sources | No | Yes, with citations | No |
| Good for | Simple, generic tasks | Knowledge-heavy questions and search | Specific formats, styles or task patterns |
| Typical effort | Low | Medium | Medium to high |
In practice, most business assistants start with RAG. Fine-tuning is added when a task needs behaviour that examples and prompting cannot reliably produce.
What we build with RAG and LLMs
- Internal knowledge assistants over HR policies, SOPs, engineering docs and product manuals.
- Customer support answers grounded in help-centre articles and past tickets, with hand-off to an agent.
- Contract and document Q&A that finds clauses, compares versions and flags risks for a reviewer.
- Search across tickets, emails and records that understands meaning rather than exact keywords.
- LLM features inside your product, such as summaries, drafting and classification, built on your data.
- Assistants on WhatsApp and websites, covered in WhatsApp AI chatbot development.
How we measure answer quality
Every RAG system we ship has an evaluation set: real questions with expected answers and the passages that support them. We score each version on:
| Measure | What it tells you |
|---|---|
| Retrieval recall | Did the system find the passages that contain the answer? |
| Faithfulness | Does the answer stick to the retrieved passages, without adding claims? |
| Correctness | Is the answer right, judged against the expected answer? |
| Citation accuracy | Do the cited passages actually support the answer? |
| Latency | How long users wait for an answer |
| Cost per question | Model and infrastructure cost at your expected volume |
The evaluation set is re-run before every model upgrade or prompt change, so improvements are proven and regressions are caught.
Keeping hallucinations under control
- Answers are restricted to retrieved passages, and the system says when it does not know.
- Citations make every answer checkable.
- Questions outside the approved scope are declined or routed to a person.
- Low-confidence answers can require human review before they reach a customer.
Common RAG pitfalls we design around
- Stale or duplicated documents. Old versions produce confident wrong answers. We sync sources continuously and deduplicate before indexing.
- Poor chunking. Splitting text in the wrong place separates a rule from its exceptions. We chunk by document structure and keep headings with their content.
- Keyword-only or vector-only search. Each misses things the other finds. We combine semantic and keyword search, then re-rank.
- Ignoring permissions. An assistant that can see everything will leak something. Access rules are enforced at retrieval time.
- No evaluation set. Without one, every change is a guess. We build the evaluation set before the first prototype.
- Answering when it should not. We set the system to say it does not know, or to route the question to a person, when the evidence is weak.
- Unbounded cost. Retrieving too much context makes every answer slow and expensive. We tune how much is retrieved against the evaluation set.
Proof from our own work
- InvoiceAI, the document-intelligence platform we operate, reads and validates invoices and reconciles them with an audit trail.
- For our client EatSafe Legal, compliance software for Indian restaurants, we built AI-assisted review of aggregator contracts.
See custom AI development for the wider range of AI systems we build.
Security, access and deployment
- Documents, indexes and logs live in your cloud account.
- Access rules are applied at retrieval time, so users only search what they are allowed to read.
- Personal data can be redacted before indexing, in line with India's Digital Personal Data Protection Act, 2023.
- We use model-provider settings that keep your data out of their training where available, and design for data residency in India when required.
What drives the cost of a RAG project
- The number of sources and how messy they are (scans, duplicates, outdated versions).
- Access-control complexity across those sources.
- How accurate answers must be, which sets the size of the evaluation set and the amount of review.
- Expected question volume, which drives model and infrastructure cost.
Every engagement is priced from a written scope after discovery, so you approve a fixed price before the build starts.
FAQ
RAG and LLM development: common questions.
What is RAG (retrieval-augmented generation)?
RAG is a technique where an AI system first searches a set of trusted documents for passages relevant to a question, then gives those passages to a large language model to write the answer, with citations back to the sources. It lets the model answer from your own, current information instead of its general training data.
Should we use RAG or fine-tuning?
Use RAG when answers depend on facts in your documents, especially facts that change. Use fine-tuning when you need the model to follow a specific format, style or task pattern that prompting cannot achieve. Many production systems combine both: RAG for knowledge, light fine-tuning or examples for behaviour.
How do you measure whether a RAG system gives correct answers?
We build an evaluation set of real questions with expected answers and the sources that support them. Each version is scored on retrieval (did it find the right passages), faithfulness (does the answer stick to them), correctness, latency and cost per question.
Can a RAG system respect who is allowed to see which documents?
Yes. Access rules from your existing systems are applied at retrieval time, so a user's question only searches documents that user is permitted to read.
Which document types can you use?
Common sources include PDFs, Word and Google documents, wiki and knowledge-base pages, support tickets, emails, spreadsheets and database records. Scanned documents are converted with OCR before indexing.
Where is our data stored?
In your own cloud account. We design for data residency, including India, when your sector requires it, and use model-provider settings that keep your data out of their training where available.