Skip to main content

Command Palette

Search for a command to run...

Building Deterministic Q&A Systems: A Serverless RAG Implementation on AWS

Updated
β€’14 min readβ€’View as Markdown

Let's be honest β€” plugging an LLM into your internal docs and hoping for the best is not a strategy. πŸ˜…

The model doesn't know your documentation. It wasn't trained on it. And fine-tuning every time someone updates a README? That's expensive and slow.

RAG fixes this. Instead of baking knowledge into weights, you retrieve it at query time. The right document chunks land in the context window, the model generates a grounded response, and you get citations pointing back to the actual source. Auditable. Updatable. No retraining required. 🎯

This post walks through a fully serverless RAG stack on AWS β€” S3, Bedrock Knowledge Bases, OpenSearch Serverless, and Bedrock Guardrails. No self-managed clusters. No embedding servers. Just managed primitives wired together with Boto3.


Table of Contents

  1. The Architecture of Retrieval

  2. Vector Store Configuration

  3. Programmatic Implementation

  4. Ensuring Response Integrity with Guardrails

  5. Conclusion: The Value of Managed AI Infrastructure


1. The Architecture of Retrieval

Two data flows. One shared set of services. Completely different directions. πŸ”€

The ingestion path runs asynchronously β€” documents go in, vectors come out, everything gets indexed.

The retrieval path runs synchronously at query time β€” user sends a question, the system finds relevant chunks, the LLM generates a grounded answer.

Unified Service Architecture Diagram β€” Dashed orange arrows indicate the asynchronous ingestion path; solid green arrows indicate the synchronous retrieval path. Guardrail enforcement is applied at the generation layer.

Unified Service Architecture Diagram β€” Dashed orange arrows indicate the asynchronous ingestion path; solid green arrows indicate the synchronous retrieval path. Guardrail enforcement is applied at the generation layer.

🧠 Amazon Bedrock

Bedrock is AWS's managed API layer for foundation models. You don't provision GPUs. You don't manage model weights. You just call an API endpoint and get inference back.

It supports models from Anthropic (Claude), Meta (Llama), Mistral, Cohere, and Amazon's own Titan family β€” all under one unified SDK. Swap models by changing a single ARN. No infrastructure changes.

For this stack, Bedrock does two jobs:

  • Embedding β€” converts raw text into vector representations using Titan Embeddings v2

  • Generation β€” takes retrieved context + user query and produces a grounded response via Claude 3.5 Sonnet

Think of it as the orchestration layer between your documents and your users. It handles the hard parts so you don't have to. 🎯

πŸ“š Bedrock Knowledge Bases β€” The Glue Layer

Knowledge Bases is a feature inside Bedrock that automates the entire RAG ingestion pipeline. Point it at an S3 bucket, pick an embedding model, choose a vector store β€” and it handles the rest.

Under the hood it:

  • Crawls your S3 bucket for new or updated documents

  • Chunks each document using your configured strategy (fixed-size, semantic, or default)

  • Embeds every chunk via the selected model (Titan Embeddings v2 in our case)

  • Upserts the resulting vectors into OpenSearch Serverless automatically

No custom Lambda. No Glue job. No hand-rolled ingestion script that breaks every time a PDF has a weird encoding. 😀

At query time, the RetrieveAndGenerate API talks directly to the Knowledge Base β€” so the same service that managed ingestion also manages retrieval. One config, both directions. That consistency is what keeps the vector space aligned between stored embeddings and query embeddings. πŸ”

Ingestion Path: S3 β†’ Bedrock Knowledge Base β†’ OpenSearch Serverless

Raw documents stored in an S3 bucket serve as the authoritative source.

A Bedrock Knowledge Base sync job polls the bucket, reads each S3 document, chunks it, runs it through Titan Embeddings v2 model, and writes 1,536-dimensional float32 vectors to OpenSearch Serverless via the AOSS bulk API.

Source URI, chunk index, and last-modified timestamp ride along as field-level metadata on every indexed document. No intermediate compute to manage. The data state transitions from raw byte stream in S3 β†’ tokenized chunk β†’ embedding vector β†’ indexed HNSW node. That's the whole pipeline. βœ…

This Zero-ETL pattern eliminates any bespoke ingestion pipeline.

Retrieval Path: Client β†’ Bedrock β†’ OpenSearch β†’ Bedrock (LLM) β†’ Client

The client fires a RetrieveAndGenerate call. Bedrock embeds the query using the same Titan model (vector space stays consistent), runs a k-NN lookup, pulls the top-5 chunks, builds the prompt, and sends it to Claude 3.5 Sonnet.

Back comes a JSON payload with the generated answer and citations linking to the exact S3 source documents. No hallucination hiding behind vague confidence scores. πŸ’ͺ

⚠️

IAM note: Both paths need a role with bedrock:InvokeModel, bedrock:Retrieve, aoss:APIAccessAll, and s3:GetObject. Attach it via instance profile to SageMaker or Cloud9 β€” no hardcoded keys, ever.


πŸ—„οΈ Vector Store (Amazon OpenSearch Serverless)

A vector store is a database built specifically for high-dimensional numerical arrays β€” not rows and columns, not key-value pairs. Just vectors, indexed for fast similarity search.

When Bedrock embeds a document chunk, it produces a 1,536-dimensional float array. That array encodes semantic meaning β€” two chunks about the same topic will produce vectors that point in the same direction, even if they use completely different words. That's what makes semantic search work. πŸ”

Amazon OpenSearch Serverless is the vector store used here. It runs the HNSW algorithm on top of the Faiss engine, which means:

  • Sub-linear query latency at high recall rates

  • No cluster to provision, scale, or babysit

  • Auto-sharding as your corpus grows

You write vectors in during ingestion. You query vectors out at inference time. That's the whole job. πŸ“¦

2. Vector Store Configuration

OpenSearch Serverless handles high-dimensional ANN search natively via the k-NN plugin. No cluster sizing. No manual shard rebalancing. No 3am pager alerts because someone ran a bad query. πŸ™

The integration with Bedrock Knowledge Bases is a direct service-to-service API call β€” no VPC proxy tier required.

Index Mapping

Create this index before running the first sync job. The dimension value must match the embedding model output exactly β€” and I mean exactly. More on that in a second. πŸ‘‡

{
  "settings": {
    "index": {
      "knn": true,
      "knn.algo_param.ef_search": 512
    }
  },
  "mappings": {
    "properties": {
      "bedrock-knowledge-base-default-vector": {
        "type": "knn_vector",
        "dimension": 1536,
        "method": {
          "name": "hnsw",
          "space_type": "cosinesimil",
          "engine": "faiss",
          "parameters": {
            "ef_construction": 512,
            "m": 16
          }
        }
      },
      "AMAZON_BEDROCK_TEXT_CHUNK": { "type": "text", "index": false },
      "AMAZON_BEDROCK_METADATA":   { "type": "text", "index": false }
    }
  }
}

Similarity Function Selection:

Cosine vs. Euclidean β€” Just Pick the Right One 🎯

Titan Embeddings v2 outputs unit-normalized vectors. Cosine similarity measures the angle between them β€” magnitude doesn't matter, direction does. That's exactly what you want for semantic search.

Euclidean distance is magnitude-sensitive. Wrong tool here. Use cosinesimil. Done.

Parameter Value Why
space_type cosinesimil Unit-normalized vectors β€” angle = semantic proximity
engine faiss Optimized for batch ANN on float32 vectors
m 16 HNSW graph connectivity β€” balances build time vs. recall
ef_construction 512 Higher = better recall, slower index build
ef_search 512 >98% recall on typical corpora
dimension 1536 Must match Titan Embeddings v2 output exactly

Chunking Strategy πŸ“„

Fixed-size chunking at 300 tokens with 20% overlap works well for structured technical docs. The overlap makes sure sentences that fall on a chunk boundary show up in both adjacent chunks. Skip the overlap and you'll miss retrieval hits right at the seams. Don't skip it. 😀

πŸ’₯ Heads up: If you swap embedding models later and the new one has a different output dimension, you can't migrate in-place. You'll recreate the index and re-ingest the whole corpus. Plan for this before you go to production.


3. Programmatic Implementation

πŸ§ͺ Amazon SageMaker Studio

SageMaker Studio is a managed JupyterLab environment running inside AWS. It's where you write and run the code in this post.

The reason it's recommended over a local setup isn't about features β€” it's about IAM. Every SageMaker Studio notebook runs with an attached execution role. Boto3 picks up those credentials automatically via the EC2 Instance Metadata Service (IMDS). No ~/.aws/credentials file. No access key rotation. No accidentally committing secrets to GitHub. 😬

It also comes with the AWS SDK pre-installed, direct network access to Bedrock and OpenSearch endpoints, and persistent storage for your notebooks.

If you'd rather work in a full IDE, AWS Cloud9 gives you the same IAM setup with a VS Code-style interface β€” useful if you're building a FastAPI or Streamlit service around the RAG pipeline.

Either way: get your execution role right first, everything else follows. πŸ”

Run all of this in Amazon SageMaker Studio (JupyterLab). πŸ§ͺ

The execution role attaches automatically via IMDS. Boto3 is pre-installed. No credential files, no SDK setup, no pip install rabbit holes on day one. If you'd rather build a FastAPI or Streamlit layer around it, AWS Cloud9 works identically for IAM β€” just a different editor.

Step 1 β€” Upload Documents to S3 with Metadata Tags

Metadata travels with the object and gets preserved as indexed field attributes in OpenSearch. Tag everything consistently from day one β€” you'll thank yourself later. 🏷️

# 01_upload_documents.py
import boto3
from datetime import datetime, timezone

s3 = boto3.client("s3")
BUCKET_NAME = "my-rag-knowledge-base-docs"

def upload_document(file_path: str, doc_type: str, version: str):
    object_key = file_path.split("/")[-1]
    with open(file_path, "rb") as f:
        s3.put_object(
            Bucket=BUCKET_NAME,
            Key=object_key,
            Body=f.read(),
            ContentType="application/pdf",
            Metadata={
                "doc_type":       doc_type,
                "schema_version": version,
                "ingested_at":    datetime.now(timezone.utc).isoformat(),
            }
        )
    print(f"Uploaded: s3://{BUCKET_NAME}/{object_key}")

Step 2 β€” Trigger Knowledge Base Sync

One API call. Bedrock handles chunking, embedding, and indexing.

# 02_sync_knowledge_base.py
import boto3

bedrock_agent = boto3.client("bedrock-agent", region_name="us-east-1")

KNOWLEDGE_BASE_ID = "KBID12345678"
DATA_SOURCE_ID    = "DSID87654321"

response = bedrock_agent.start_ingestion_job(
    knowledgeBaseId = KNOWLEDGE_BASE_ID,
    dataSourceId    = DATA_SOURCE_ID,
    description     = "Incremental sync β€” new documents batch 2025-07-15"
)

print("Ingestion job ID:", response["ingestionJob"]["ingestionJobId"])

Step 3 β€” Query with RetrieveAndGenerate πŸš€

This is the whole thing in one call. Embedding, retrieval, prompt assembly, generation. The sessionId keeps multi-turn context alive server-side β€” your client stays stateless.

# 03_retrieve_and_generate.py
import boto3
import json

bedrock_runtime = boto3.client(
    "bedrock-agent-runtime",
    region_name="us-east-1"
)

KNOWLEDGE_BASE_ID = "KBID12345678"
MODEL_ARN         = "arn:aws:bedrock:us-east-1::foundation-model/anthropic.claude-3-5-sonnet-20241022-v2:0"
GUARDRAIL_ID      = "GRDID99887766"
GUARDRAIL_VERSION = "DRAFT"

def query_knowledge_base(
    user_query: str,
    session_id: str | None = None
) -> dict:

    payload = {
        "input": { "text": user_query },
        "retrieveAndGenerateConfiguration": {
            "type": "KNOWLEDGE_BASE",
            "knowledgeBaseConfiguration": {
                "knowledgeBaseId": KNOWLEDGE_BASE_ID,
                "modelArn": MODEL_ARN,
                "retrievalConfiguration": {
                    "vectorSearchConfiguration": {
                        "numberOfResults": 5,
                        "overrideSearchType": "SEMANTIC"
                    }
                },
                "generationConfiguration": {
                    "guardrailConfiguration": {
                        "guardrailId": GUARDRAIL_ID,
                        "guardrailVersion": GUARDRAIL_VERSION
                    }
                }
            }
        }
    }

    if session_id:
        payload["sessionId"] = session_id

    response = bedrock_runtime.retrieve_and_generate(**payload)

    return {
        "session_id":     response["sessionId"],
        "generated_text": response["output"]["text"],
        "citations":      response.get("citations", [])
    }

# ── Run it ─────────────────────────────────────────────────────────────
result = query_knowledge_base(
    "What is the default timeout value for AWS Lambda functions?"
)

print(json.dumps(result, indent=2))

Expected JSON output:

The citation shows exactly which S3 document answered the question. Score of 0.9421 means the retrieval was tight. That's the system working as intended. πŸŽ‰

SageMaker Studio JupyterLab β€” Cell execution output. The upper cell shows theΒ Β Boto3 invocation. The output cell returns a structured JSON response containingΒ , theΒ Β for multi-turn state, andΒ Β identifying the source S3 object and chunk index that grounded the answer.

SageMaker Studio JupyterLab β€” Cell execution output. The upper cell shows theΒ RetrieveAndGenerateΒ Boto3 invocation. The output cell returns a structured JSON response containingΒ generated_text, theΒ session_idΒ for multi-turn state, andΒ citationsΒ identifying the source S3 object and chunk index that grounded the answer.

Inference Latency Breakdown ⚑

Stage Approx. Latency
Query tokenization + Titan Embeddings v2 ~30 ms
k-NN retrieval (100K vector collection) ~20–50 ms
Claude 3.5 Sonnet generation (~200 tokens) ~800–2,000 ms
End-to-end P95 ~1.2–2.5 seconds

Most of the time is generation. The retrieval is fast. If your P95 is slow, look at your output token count first. πŸ”Ž


4. Ensuring Response Integrity with Guardrails

Good retrieval doesn't guarantee good output. The LLM can still go off-script β€” introduce facts not in the retrieved chunks, expose PII that slipped through, or drift into topics it has no business answering. 😬

Bedrock Guardrails wraps the generation step with a declarative enforcement layer. PII detection, content filtering, topic boundaries, and grounding checks β€” configured once, applied on every call.

Setting Up a Guardrail via Boto3

# 04_create_guardrail.py
import boto3

bedrock = boto3.client("bedrock", region_name="us-east-1")

guardrail = bedrock.create_guardrail(
    name        = "rag-qa-guardrail-v1",
    description = "PII redaction + topic boundaries for the docs Q&A system",

    # Detected PII gets replaced with typed placeholder tokens.
    # [EMAIL], [NAME], etc. β€” not stored, not logged, gone. πŸ”’
    sensitiveInformationPolicyConfig = {
        "piiEntitiesConfig": [
            { "type": "EMAIL",          "action": "ANONYMIZE" },
            { "type": "PHONE",          "action": "ANONYMIZE" },
            { "type": "AWS_ACCESS_KEY", "action": "BLOCK"     },
            { "type": "NAME",           "action": "ANONYMIZE" },
        ]
    },

    # Per-category harm thresholds. Breach the level β†’ request blocked.
    contentPolicyConfig = {
        "filtersConfig": [
            { "type": "HATE",          "inputStrength": "HIGH",   "outputStrength": "HIGH"   },
            { "type": "VIOLENCE",      "inputStrength": "MEDIUM", "outputStrength": "MEDIUM" },
            { "type": "PROMPT_ATTACK", "inputStrength": "HIGH",   "outputStrength": "NONE"   },
        ]
    },

    # Hard topic boundary. Model won't engage regardless of context. 🚫
    topicPolicyConfig = {
        "topicsConfig": [
            {
                "name": "financial-advice",
                "definition": "Questions requesting investment, trading, or financial guidance",
                "examples": ["Should I buy AWS stock?", "What is the ROI of this service?"],
                "type": "DENY"
            }
        ]
    },

    blockedInputMessaging   = "This query falls outside the system's defined scope.",
    blockedOutputsMessaging = "The generated response was blocked by the content policy."
)

print("Guardrail ARN:", guardrail["guardrailArn"])

PiiEntitiesConfig β€” ANONYMIZE vs. BLOCK

ANONYMIZE swaps the entity with a token ([EMAIL], [NAME]). The original value never hits your logs or your client. Substitution happens in-memory inside the Guardrail pipeline.

BLOCK kills the whole response. Use it for AWS_ACCESS_KEY, CREDIT_DEBIT_CARD_NUMBER, anything you absolutely cannot have surfacing in generated text. No half-measures there. πŸ›‘

Grounding Score β€” Catching Hallucinations Quantitatively

This one's underrated. Guardrails can compute a semantic similarity score between the generated response and the retrieved context.

If the model introduces a claim that isn't in the retrieved chunks β€” which is exactly what hallucination is β€” and the grounding score drops below your threshold, the response is blocked. No vibes-based review. Quantitative enforcement. πŸ“

πŸ”– Version your guardrails: Call bedrock.create_guardrail_version() after creation to lock a DRAFT into a numbered version ("1", "2", etc.). Point production at the version number, not DRAFT. That way tuning in dev doesn't accidentally nuke live traffic. Ask me how I know. πŸ˜…


5. Conclusion: The Value of Managed AI Infrastructure

Here's the short version: the whole stack runs without a single self-managed server. πŸ–οΈ

OpenSearch Serverless auto-provisions shards as query throughput or index size grows. 1 query or 1,000 concurrent queries β€” same config, no operator touch. Manual cluster sharding is someone else's problem.

Bedrock routes inference across a managed model fleet. No reserved instances, no capacity planning. Scale hits your account quota, not an infrastructure ceiling.

S3 holds your documents with eleven-nines durability and no capacity limit. Incremental sync jobs are idempotent β€” unchanged documents don't get re-embedded. No redundant inference spend. πŸ’Έ

IAM Policy β€” Minimum Viable Scope

{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "BedrockInference",
      "Effect": "Allow",
      "Action": [
        "bedrock:InvokeModel",
        "bedrock:RetrieveAndGenerate",
        "bedrock:Retrieve"
      ],
      "Resource": [
        "arn:aws:bedrock:us-east-1::foundation-model/*",
        "arn:aws:bedrock:us-east-1:<ACCOUNT_ID>:knowledge-base/KBID12345678"
      ]
    },
    {
      "Sid": "OpenSearchAccess",
      "Effect": "Allow",
      "Action": "aoss:APIAccessAll",
      "Resource": "arn:aws:aoss:us-east-1:<ACCOUNT_ID>:collection/<COLLECTION_ID>"
    },
    {
      "Sid": "S3ReadDocuments",
      "Effect": "Allow",
      "Action": [ "s3:GetObject", "s3:ListBucket" ],
      "Resource": [
        "arn:aws:s3:::my-rag-knowledge-base-docs",
        "arn:aws:s3:::my-rag-knowledge-base-docs/*"
      ]
    }
  ]
}

Scope it tight. Principle of least privilege isn't optional. πŸ”

Before You Go to Prod β€” Three Things to Consider πŸ“‹

  • Version your corpus. Tag S3 objects with corpus_version. For major updates, spin up a new Knowledge Base + vector collection rather than re-syncing in-place. Zero-downtime cutover + rollback path. Worth the extra five minutes.

  • Measure retrieval quality. MRR or NDCG against a held-out eval set. If scores drop after a corpus update, your chunking strategy or embedding model is the first thing to check.

  • Tag everything for cost attribution. Bedrock charges per 1K tokens (generation) and per 1K vectors (retrieval). Slap project, environment, and team tags on every resource and use Cost Explorer to track per-query spend over time.

The entire thing runs from a single SageMaker Studio notebook with a scoped execution role. No servers to babysit. No clusters to tune. Just working, citeable, auditable Q&A on your own documents. πŸš€

More from this blog

C

Chandradeo Arya's Tech Blog

13 posts

DevOps & Cloud Instructor, Curriculum Author, Solutions & AI Architect