Available Now · 90+ Readymade Solutions

RAG Development Company

pgvectorPineconeQdrantWeaviateHybrid SearchCitations

Miracuves builds RAG systems that answer from your own documents and data, not from what a model remembers. We ingest and clean your sources, index them in a vector database, retrieve with hybrid search and reranking, and return answers that cite the exact passage - filtered by what each user is allowed to open. You own 100% of the source code.

Reviewed on ClutchReady-made platforms from $2,199Need the whole LLM app? See LLM development

  • Cited Answers
  • Permission-Aware Search
  • 100% Source Ownership
  • NDA Day One
6 daysReady-made platform delivery
$3,699RAG pipeline base, from
2-8 wksCustom RAG build
100%Source code ownership
Retrieval engineers online now
CitationsHybrid searchRerankingAccess control9,000+ deliveredNDA day one
  • Web · Slack · APIOne retrieval index behind every channel
  • Permission-awareUsers only see what they may open
  • Faithfulness scoredTested on your real questions
  • 6 DaysReady-made platform, brief to live
  • Cited answersEach claim linked to its passage
  • Grounded Answers

    From your files, not model memory

  • NDA Day One

    Signed before documents are shared

  • Full Source Code

    Pipeline, index config and eval set

  • 60-Day Support

    Re-index and retrieval tuning after launch

  • 100% IP Ownership

    Embeddings and index are yours

  • Reviewed on Clutch

    Third-party client reviews

More than 6,000+ Companies Trust us Worldwide
In short

Miracuves is a RAG development company. We build retrieval augmented generation systems that search a company's own documents and data, pass the relevant passages to an LLM and return answers with citations, filtered by each user's permissions and kept fresh as files change. Custom RAG builds take 2-8 weeks, a ready-made AI chat platform ships in 6 working days, and you own 100% of the source code.

Our RAG Approach

How Miracuves builds retrieval augmented generation - grounded in your own documents

A language model only knows what it was trained on. Retrieval augmented generation fixes that at question time: the system searches your manuals, policies, tickets and records, hands the few passages that matter to the model, and the model answers from those passages with a citation to each one. When a policy changes, you re-index the file; nothing is retrained.

Most failed RAG projects fail at retrieval, not at the model. Scanned PDFs lose their tables, chunks cut a rule in half, pure vector search misses part numbers and exact terms. So Miracuves spends the effort where answers are won: clean ingestion, chunking that follows your document structure, hybrid search with a reranker, and an access filter applied before anything reaches the model.

Who this service is built for: Teams whose answers already exist in documents nobody can find fast enough - support teams working from help centers and release notes, operations staff checking SOPs, legal and compliance teams searching contracts and policies, sales engineers answering security questionnaires, and SaaS products that want a knowledge base AI assistant for their own customers. If you need the whole application around the model - agents, tools, fine-tuning, multi-model routing - that is our LLM development service. If you need a guided conversation with fixed flows rather than answers from documents, that is chatbot development.

  • Ingestion and cleaning: PDFs, scans, tables, wikis and tickets parsed with structure kept, duplicates removed
  • Chunking by document structure: headings, clauses and table rows stay intact, each chunk carries its source and page
  • Hybrid search with reranking: vector and keyword results merged, then a cross-encoder picks the passages worth sending
  • Permission-aware retrieval: access lists from your identity provider applied inside the search query, not after it
  • Faithfulness evaluation: every release scored on your own questions before it reaches users
9,000+Projects delivered since 2010
3,900+Apps published by Miracuves
90+Ready-made solutions to start from
6 daysReady-made platform delivery
2-8wMiracuves custom RAG build timelines
100%Source code ownership
IngestPDFs, wikis, tickets, tables
RetrieveHybrid search + reranking
CiteAnswers linked to sources

Why RAG at Miracuves

  • Vector storespgvector · Pinecone · Qdrant · Weaviate
  • RetrievalHybrid search + reranking
  • Access controlEnforced inside the query
  • CitationsOn every answer
  • FreshnessChanged files re-indexed automatically
  • Your documentsKept out of model training
Example engagement: Policy and procedure assistant, a typical 4-6 week build
"An operations team keeps 900 pages of SOPs across SharePoint and scanned binders. The build: OCR and layout parsing for the scans, clause-level chunks, pgvector inside their existing Postgres, hybrid search with a reranker, and answers that open the cited page. Site managers see only their own region's procedures, and a nightly job re-indexes whatever changed that day."

Approach Comparison

RAG vs fine-tuning vs a long-context prompt - which grounds your answers?

Three ways to get a model to answer about your business. They differ most on freshness, citations and who is allowed to see what - which is where most enterprise RAG decisions are actually made.

MetricFive things that decide cost, speed and reach
Miracuves default

RAG by Miracuves

Hybrid search + rerank + citations

Fine-tuned model

Knowledge trained into weights

Long-context prompt

Whole documents pasted in

01Fresh knowledge
Re-index a fileUpdated or deleted documents change answers the same day
Retrain to updateNew facts need a new training run
Paste againOnly as current as what you send each time
02Citations
Every answerEach claim links to the passage and page it came from
NoneKnowledge is blended into the weights
Hard to verifyThe model may quote or may paraphrase loosely
03Access control
Per user, at query timeChunks carry access lists from your identity provider
None inside the modelAnyone who can query can surface what it learned
ManualYou decide what to paste for whom
04Cost per query
A few passagesOnly the top reranked chunks are sent to the model
Low to runTraining and re-training runs cost up front
High at scaleEvery token of every document, every question
05Best for
Large, changing document setsSupport, policy, contract and product knowledge
Style and formatHouse tone, fixed outputs, domain vocabulary
Small, one-off jobsA few files and occasional questions
Choose RAG if…

Your answers live in documents that change · users must see a source · different people may read different files · the corpus is too big to paste into a prompt.

Consider an alternative if…

You need a fixed style or output format more than facts (fine-tuning) · the model must take actions across systems (AI agents) · you need the full application around the model. See LLM Development →

RAG guide

What to know before you hire a RAG development company

The questions buyers ask us before a RAG project starts - how the pipeline works, when fine-tuning is the better call, which vector database to pick, and what it costs to build and run.

How does a RAG pipeline work, step by step?

A RAG system has two halves. Indexing runs ahead of time; answering runs on every question.

  • Ingestion and cleaning: connectors pull files from your sources, parsers keep tables and headings, boilerplate and duplicates are dropped
  • Chunking: documents are split along their own sections, and each chunk keeps its source, page, version and access list
  • Embeddings: each chunk becomes a vector and is stored in a vector database next to a keyword index
  • Retrieval: a question runs vector and keyword search together, limited to chunks the user may read, then a reranker keeps the best few
  • Generation: the model answers only from those passages and cites each one, or says the sources do not cover it

RAG or fine-tuning - which one does my use case need?

Choose RAG when the knowledge changes, when users need to see where an answer came from, or when different people may read different documents. Updating a RAG system means re-indexing a file; updating a fine-tuned model means another training run, and it still cannot tell you its source or hide one department's data from another.

Fine-tuning is the better choice when the problem is behavior rather than knowledge: a fixed output format, a house writing style, or domain shorthand that prompting keeps getting wrong. The two combine well - a fine-tuned model that writes in your format, fed by RAG with current facts. If you need that wider application work, it sits with our LLM development team.

Which vector database should we use: pgvector, Pinecone, Qdrant or Weaviate?

Start with what you already run. If your data lives in PostgreSQL, pgvector keeps vectors, metadata and access rules in one database with one backup, and handles most company knowledge bases comfortably. Pinecone suits teams that want a fully managed index with no servers to operate. Qdrant is open-source, fast at filtering on metadata such as tenant and access group, and can be self-hosted. Weaviate ships hybrid keyword and vector search built in.

The choice matters less than people expect: retrieval quality is decided by parsing, chunking and reranking. We keep the store behind one interface in the code, so changing it later means re-loading the index, not rewriting the application.

How do you stop users seeing answers from documents they may not open?

Permissions have to be enforced at retrieval, not at the answer. Every chunk is stored with the access list of the document it came from, taken from the source system or your identity provider through SSO groups. At question time the search itself is filtered to the groups the user belongs to, so a restricted passage is never fetched and can never leak into an answer or a citation.

Access changes are synced like content changes: when someone leaves a team or a file is re-shared, the next sync updates the chunk's access list. Before launch we run leak tests with users from each group asking questions whose answers sit in files they must not see.

How do you keep a RAG system accurate and up to date after launch?

Accuracy is measured, not assumed. Before launch we agree a golden set of real questions with approved answers and their source passages. Retrieval is scored separately (did the right passage come back) from generation (does the answer stay faithful to what was retrieved), so a failure can be traced to search or to the model. A release that drops below the agreed bar does not ship, and the same set is re-run when you change the embedding model or the LLM.

Freshness is an engineering job too. Connectors re-index only files that changed, remove deleted ones the same day, and keep version labels so an answer about release 5 never quotes the release 4 guide.

What does RAG development cost to build and to run?

Building: a RAG system on our pipeline base starts from $3,699 and takes 2-8 weeks. A custom RAG system typically runs $8,000-$25,000, also 2-8 weeks, with larger scopes quoted in writing first. If you need a full chat product as the front end, our ChatGPT clone ships in 6 working days from $2,799, with retrieval added as custom work.

Running: you pay once to embed each document (and again for changed ones), for vector database storage, and for model tokens on every answer. Sending five reranked passages instead of twenty, caching repeated questions and routing simple ones to a smaller model are the three levers we use to hold cost and latency down.

How do you evaluate a RAG development company?

Ask how they would measure your system, not how clever the demo is. A serious vendor will ask for sample documents and real questions before quoting, explain how permissions are enforced inside the search, show a retrieval and faithfulness report from a test set, and tell you what happens when a file is deleted.

Also check you are buying the right thing. RAG is the retrieval layer. For the whole LLM application around it, see LLM development; for scripted conversation flows, see chatbot development; to plug AI into systems you already run, see AI integration services; and if you are still deciding where AI fits at all, start with AI consulting.

Technical Architecture

How Miracuves engineers structure a RAG pipeline for production

Retrieval quality is decided long before a question is asked - at parsing, chunking and indexing. These are the decisions our engineers make on every RAG build.

  • 01

    Indexing - Parse, Clean, Chunk, Embed

    Documents are parsed with layout kept (tables as rows, headings as structure), cleaned of boilerplate and duplicates, then split along their own sections. Each chunk stores its source, page, version and access list, so a citation and a permission check are always possible later.

  • 02

    Retrieval - Hybrid Search, Filter, Rerank

    A query runs vector search and keyword (BM25) search together, restricted to chunks the user may read. The merged candidates go through a cross-encoder reranker and only the top few reach the model - fewer tokens, fewer distractions, better answers.

  • 03

    Freshness and Cost - Incremental Re-indexing, Caching

    Connectors watch your sources and re-embed only what changed, and deletions are removed from the index the same day. Repeated questions hit a semantic cache, and easy questions route to a smaller model, keeping latency and the monthly bill predictable.

What most RAG projects get wrong

Fixed-size chunks that split tables. Vector-only search that misses SKUs and clause numbers. No reranker. Permissions checked after retrieval, or never. No re-indexing, so answers quote last year's policy. No test set, so nobody can say whether a change helped. Each one is cheaper to design in than to retrofit.

retrieve.py - permission-aware hybrid search on pgvector
# Hybrid retrieval: pgvector similarity + Postgres full-text, one query# The access filter runs inside the search, so no forbidden chunk is ever fetchedfrom rag.embed import embedfrom rag.rerank import rerankSQL = """WITH dense AS (  SELECT id FROM chunks WHERE acl && %(groups)s  ORDER BY embedding <=> %(qvec)s LIMIT 40),sparse AS (  SELECT id FROM chunks WHERE acl && %(groups)s    AND tsv @@ websearch_to_tsquery(%(q)s)  ORDER BY ts_rank(tsv, websearch_to_tsquery(%(q)s)) DESC LIMIT 40)SELECT id, text, source_uri, page FROM chunksWHERE id IN (SELECT id FROM dense UNION SELECT id FROM sparse)"""def retrieve(db, question: str, groups: list[str], top_n: int = 6):    rows = db.execute(SQL, {"qvec": embed(question), "q": question, "groups": groups})    # Cross-encoder scores each candidate against the question, best first    best = rerank(question, rows)[:top_n]    # Source and page travel with every passage so the answer can cite it    return [{"n": i + 1, "text": r.text, "cite": f"{r.source_uri}#page={r.page}"}            for i, r in enumerate(best)]
Vector and keyword search in one Postgres query, restricted by the user's groups, then reranked. The same pattern runs on Pinecone, Qdrant or Weaviate using their metadata filters.

Our Service Models

Three ways Miracuves delivers your RAG system

Whether you need one knowledge base assistant or retrieval inside your own product, you work with Miracuves as a company: retrieval engineers, backend, QA and a project lead, accountable for the answer quality.

Most Popular
Customer app
Partner app
Admin
RAG Pipeline Base · Fixed Price

RAG Pipeline Delivery

Miracuves starts from its RAG pipeline base - ingestion, hybrid search, reranking, citations and an evaluation harness - and connects it to your sources, permissions and interface in 2-8 weeks. Source code fully yours.

  • From $3,699 · written quote before any work
  • Connectors for your document stores and wikis
  • Citations and permission filter configured
  • Evaluation set built from your real questions
  • Full source code · NDA · 60-day support
RAG PipelineIngestionRetrievalAnswerChunk + ACLHybrid + rerankCitations
Custom RAG Development · Scoped

Custom Enterprise RAG Build

For large or messy corpora, multiple tenants or strict access rules: retrieval designed around your data, from parsing strategy to vector store choice. Retrieval engineer, backend, QA and project lead.

  • Scope and price fixed in writing before the build
  • Chunking and embedding choices tested on your documents
  • Weekly demos on your own questions
  • Faithfulness and retrieval scores reported each sprint
  • Full source code · IP 100% yours
Wk 1
Wk 2
Wk 3
Wk 4
Ongoing Retainer · Monthly

Ongoing RAG Development

Miracuves keeps your retrieval layer sharp as the corpus grows: new sources, re-tuned chunking, model and embedding upgrades re-tested against your evaluation set, on a monthly retainer.

  • From $2,299/month - cancel with 2 weeks notice
  • Named Miracuves engineers on your index
  • Talk to the engineers directly, not an account manager
  • New connectors and sources added each cycle
  • Scales up or down with your document volume

Quality Standards

How Miracuves checks every RAG delivery before users rely on it

A RAG system is only as good as its worst retrieval. These gates check the index, the access filter and the answers on your own questions before anything is handed over.

  • Parsing checked on your hardest files - scans, tables, multi-column layoutsIngestion
  • Chunk boundaries follow document structure, each chunk carrying source and pageIndexing
  • Retrieval precision and recall measured on a labeled question setRetrieval
  • Faithfulness and answer relevancy scored on every releaseEvaluation
  • Access tests: a user without rights must never retrieve a restricted chunkSecurity
  • Deleted and updated files verified out of the indexFreshness
  • Latency and cost per answer tracked in productionDelivery

Enforced QA Gates

Our 6 Retrieval Quality Gates

Every change to parsing, chunking, embeddings, prompts or the model must clear all six gates before it reaches your users.

01

Golden Question Set Agreed

Before building, we agree 50 to a few hundred real questions with approved answers and their source passages. This set defines "correct" for your corpus.

02

Retrieval Scored Separately

Did the right passage come back in the top results? Retrieval is scored on its own, so a bad answer is traced to search or to generation, not guessed at.

03

Faithfulness Threshold Enforced

Answers are checked claim by claim against the cited passages. A release that drops below the agreed faithfulness bar does not ship.

04

Permission Leak Tests

Test users from each access group ask questions whose answers sit in restricted files. Any restricted chunk in the results fails the build.

05

Re-index Drill

We edit, add and delete source files and confirm the index reflects each change within the agreed window before handoff.

06

Post-Launch Review - 60 Days

Unanswered questions, low-rated answers and zero-result searches are reviewed through the 60-day support window and fed back into chunking and the question set.

Technology Stack

The RAG stack Miracuves ships with

Chosen per project: the vector store you can already run, the embedding model that scores best on your documents, and a generation model allowed to see your data.

pg
pgvectorVectors inside your Postgres
Pi
PineconeManaged serverless vector index
Qd
QdrantOpen-source · payload filtering
Wv
WeaviateBuilt-in hybrid search
OS
OpenSearchBM25 keyword retrieval
LI
LlamaIndexIngestion and index framework
LC
LangChainRetrieval chains · loaders
Un
UnstructuredPDF, scan and table parsing
Em
Embedding modelsHosted or open-source, tested per corpus
Rr
Cross-encoder rerankersSecond-pass relevance scoring
GC
GPT · Claude · Gemini · LlamaAnswer generation
Rg
RagasFaithfulness · context precision
Lf
LangfuseTracing · retrieval inspection
FA
FastAPIRetrieval and answer API
Rd
RedisSemantic cache · rate limits
Af
AirflowScheduled re-index jobs

Our Process

From scattered documents to cited answers - what happens and when

Every RAG engagement follows the same order, because retrieval has to be proven before generation is worth tuning. At each step you know which sources, access groups and sample questions we need from you, and what you get back. Custom builds run milestone-based with the same checkpoints.

  1. Step 01

    Brief & NDA

    Tell us which questions users ask and where the answers live. NDA signed before any document is shared.

  2. Step 02

    Corpus & Access Audit

    We sample your sources, map access groups and agree the golden question set. No payment before scope is agreed.

  3. Step 03

    Index & Retrieve

    Parsing, chunking and embeddings tuned until the right passages come back for your questions.

  4. Step 04

    Answer & Evaluate

    Generation with citations added, scored for faithfulness, permission leaks and latency.

  5. Step 05

    Launch & Re-index

    API and interface live, source connectors scheduled, 60 days of support and tuning.

NDA FirstBefore any document is shared
2-8 WeeksRAG pipeline base delivery
Before LaunchFaithfulness scored on your questions
60 DaysPost-launch support

Six days is for the ready-made platform only

The six working days cover a ready-made catalogue platform. A RAG layer is custom work of 2-8 weeks, and what moves it most sits on your side: read access to the document stores, the permission groups from your identity provider, and real questions with approved answers. We list them on the first call so you can gather them in parallel.

See what you provideFACT-005, audited quarterly

Transparent Pricing

What RAG development costs at Miracuves

A RAG quote is driven by your corpus, not by guesswork: how many sources, how messy the files are, how strict the access rules are. We publish the starting points and put the rest in writing before work starts.

Ready-Made Platform

$2,199 from

Catalogue floor · 6 day delivery · fixed price

  • Ready-made platforms start at this price; the ChatGPT clone is priced on its own page
  • Knowledge base managed from the admin panel
  • Your branding and white-label applied
  • Full source code on handoff
  • 60-day post-launch support
  • Retrieval with citations added as custom work
Start a Clone Project
Most Requested

Custom RAG System

Custom Quote

Scoped before build · milestone billing

  • Retrieval engineer + backend + QA + project lead
  • Parsing and chunking built for your documents
  • Permission-aware hybrid search with reranking
  • Faithfulness and retrieval scores every sprint
  • Full source code · complete IP transfer
  • Billed per milestone, paid after each delivery
Get a Scope & Quote

Ongoing RAG Development

$2,299 /mo

Monthly retainer · cancel with 2 weeks notice

  • Miracuves engineers assigned to your index
  • New sources, connectors and re-tuning
  • Model and embedding upgrades re-evaluated
  • Direct communication - no relay
  • Scales with your document volume
  • Index, embeddings and code remain 100% yours
Discuss Ongoing Work

Why Miracuves publishes pricesRAG budgets drift when nobody names the cost drivers early. We name them on the first call - document volume, scan quality, number of access groups and connectors - and explain which one moves your quote before you sign anything.

What affects RAG project cost at Miracuves

The ready-made platform price stays fixed when your scope matches the base product, such as our ChatGPT clone. Custom RAG work scales with: the number of source systems (SharePoint, Confluence, Google Drive, Zendesk, databases), document volume and quality (scanned PDFs, tables, handwritten forms), how fine-grained permissions are, languages, the size of the evaluation set, freshness targets for re-indexing, and whether embeddings and models run on hosted APIs or on your own infrastructure.

Typical RAG budget ranges

  • RAG pipeline basefrom $3,6992-8 weeks
  • Custom RAG system$8,000-$25,0002-8 weeks; larger scopes are quoted in writing before work starts
  • Ongoing retainerfrom $2,299/month for new sources, tuning and upgrades

Every quote is written before payment - no surprise invoices after kickoff.

Example engagement

What a typical RAG project looks like at Miracuves

An illustrative example of a typical project of this kind, with client details anonymized. Figures show what this kind of build targets, not a named client's results.

A B2B software company wants one assistant for its support and sales engineers. The answers are spread across a public help center, a Confluence wiki with restricted spaces, PDF admin guides for each release, and years of resolved tickets.

  1. 01

    The Challenge

    Engineers search four tools and still quote old release behavior. Some wiki spaces hold customer contracts that only account owners may read, so a plain search box over everything is not allowed.

  2. 02

    What This Kind of Build Delivers

    Connectors for the help center, Confluence and ticket system; version-aware chunking so each release guide is labeled; pgvector with hybrid search and a reranker; Confluence space permissions mirrored into every chunk; answers in Slack and the support console with a link to the cited page.

  3. 03

    What It Targets

    Answers that name the release they apply to, zero restricted chunks in permission tests, updated guides searchable the day they publish, and faithfulness above the threshold agreed on the golden question set before launch.

2-8 WeeksTypical build
Per userAccess filter
Same dayRe-index target
See Real Client Projects
Project Brief
  • Build typeCustom RAG on the pipeline base
  • Typical timeline2-8 weeks
  • SourcesHelp center · Confluence · PDFs · tickets
  • Retrievalpgvector · hybrid · reranker
  • ChannelsSlack · support console · API
  • Source code100% client-owned
CitedEvery answer
ZeroTarget permission leaks
60 daysSupport included

Client Reviews

What clients say about building with Miracuves

These clients did not buy a RAG project from us; they built platforms with Miracuves. One added its own path recommender on top of a course catalog, and two run a rights registry on a base we built that already handled documents and permissions - the same ground a retrieval system stands on. Read every testimonial on our client testimonials page.

Client testimonial
"Miracuves's MXFlix + MXLearn combination gave us the player, the catalog, the user progress tracking, and the quiz/assignment layer. We added our path-recommender, our cohort-management module, and a custom certificate engine."
EM
Eric J. MorinCEO & Speaker, Tower Leadership
Learning platform with its own path recommender over the course catalog
Client testimonial
"Copyright administration is mostly workflow: intake, examination, register, certificate. We had all of it on email and spreadsheets. Miracuves built the registry and the examiner workflow on a base that already handled documents and approvals, so we spent the time on the rules that are specific to how rights actually get registered. Live inside a month."
IC
International Copyright Organization TeamFounding team, International Copyright Organization
Copyright registry: intake, examination, register and certificate workflow
Client testimonial
"Once the register worked, we needed rights holders to be able to license and share what they had registered. The permissions model and the document handling were already there, so the new work was the licensing terms and the split logic. It shipped in weeks, not the quarter we had budgeted."
IS
International Copyright Organization TeamFounding team, Intercopy Share
Rights licensing and revenue split built on the registry
6,000+Clients served
3,900+Apps published
35+Industries served
Read All Reviews

Why Miracuves

Six places to check us before you ever call us

Each one is either run by someone else or open to anyone. Check them in any order; the whole list takes about a minute.

Why clients choose Miracuves

Three promises we would stake the company on

Every promise on this site rests on these three. Each one is something you can check, not something you have to take on trust.

  • 01People you can name

    Our leadership is public, with real LinkedIn profiles, not a stock-photo team page. A named team works your build and sends you progress on WhatsApp every working day.

    Meet the leadership
  • 02Proof over promises

    Every number we publish, pricing, timelines, project counts, is defined and sourced on a public facts ledger. If we can't back a claim, we don't make it.

    Read the facts ledger
  • 03A process with a deadline

    Ready-made platforms go from kickoff to live deployment in 6 working days, guaranteed: miss it for reasons on our side and we work free until launch. Custom builds get a fixed quote after a free feasibility study.

    Get a feasibility study

Frequently Asked

Questions about RAG development at Miracuves

Something not covered here? Ask on WhatsApp and you will usually have an answer within two hours.

Ask us directly
Which data sources can a RAG system connect to?

Most document stores and business tools: SharePoint, Google Drive, Confluence, Notion, Zendesk and other help desks, S3 buckets, websites, and relational databases. Each source gets a connector that pulls new and changed files on a schedule or on change events, keeps the original link for citations, and brings the source's own permissions along. Sources without an API can be loaded from exports while a connector is built.

How much does RAG development cost at Miracuves?

The ready-made platform price stays fixed when your scope matches the base product. A RAG system on our pipeline base starts from $3,699 and takes 2-8 weeks. A custom RAG system typically runs $8,000-$25,000 and takes 2-8 weeks; larger scopes are quoted in writing before work starts. Where it lands depends on the number of source systems, document volume and scan quality, permission rules, languages and evaluation depth. Ongoing RAG development is available from $2,299/month. Every quote is written before payment, with no surprise invoices after kickoff.

Can a RAG assistant read scanned PDFs, tables and diagrams?

Scans are run through OCR with layout detection, so columns, headers and footers are handled rather than read as one stream of text. Tables are extracted row by row and stored with their headers, which is what lets a question about one value find the right row. Diagrams and photos can be described by a vision model and indexed as text, with a link back to the original page.

What do we need to prepare before a RAG project starts?

Three things. Read access to the sources that hold the answers, or exports of them. The access groups from your identity provider, so retrieval can mirror who may read what. And 50 or more real questions with the answers your experts would give, which become the evaluation set. With those ready, scoping takes days rather than weeks.

Does our data get used to train the language model?

No. In RAG your documents stay in your index; only the few passages retrieved for a question are sent to the model, and we choose provider plans that exclude training on your data. Where documents may not leave your network, embeddings and generation run on open-weight models inside your cloud account. An NDA is signed before any document is shared.

Who owns the index, embeddings and code after delivery?

You do. Handoff includes the ingestion and retrieval code, connector configurations, the index schema, the evaluation set and its reports, prompts, and deployment scripts, with 100% source code ownership. The vector database runs in your account, so you can re-embed with a different model or move vector stores later without asking us.

Does RAG work in languages other than English?

Yes. Multilingual embedding models place a question in one language close to a passage in another, so a user can ask in Spanish and retrieve an English policy. We test retrieval per language on your evaluation set, add keyword search tuned for each language, and answer in the user's language while citing the original passage.

Can you fix an existing RAG prototype that gives wrong answers?

Often, yes. We first run your real questions through it and score retrieval and faithfulness separately, which usually shows the problem sits in search - chunks too large or too small, no keyword search, no reranker, stale content - rather than in the model. Fixes are scoped from that report, so you keep what works and replace only what fails.

Can RAG answer questions about data in our database?

For text stored in a database - product descriptions, case notes, ticket bodies - yes, it is indexed like any document. For numbers and totals, such as last quarter's revenue by region, retrieval is the wrong tool; the model should write a query against the database instead. Many assistants need both, and we route each question to the right path.

Get Started

Ready to ground your AI answers in your own documents?

Tell Miracuves where your answers live and who is allowed to see them. We will come back with the retrieval design, vector store choice and delivery timeline - in writing, before any commitment.

9,000+Projects since 2010
2-8 WeeksCustom RAG build
100%Source code yours
NDA FirstBefore any document is shared
Book a Free ConsultationContact & Brief Form

NDA signed before you share a single document

Page reviewed by the Miracuves AI engineering team · Last updated September 2026 · Clutch & Google Reviews

Disclaimer

Miracuves is an independent software development company. We are not affiliated with, connected to, sponsored by, or endorsed by any of the brands or platforms named on this page.

Why these names

Names of the form “Brand Clone” are used descriptively. It is how the software industry refers to building a platform with functionality comparable to a known service, and how clients search for it.

Who built this

The entire design and codebase of our products is built by our own team. Our products contain no code, design, graphics, or content originating from any third-party website or application.

Trademarks

All third-party names and marks listed on this page are the property of their respective owners, referenced solely to describe the category of software offered.