Codebase study guide
PolicyLens
An AI web app that reads an insurance policy PDF and answers questions about it in plain English. This guide explains how every part works, shows where data moves, and lists the questions an interviewer is likely to ask.
Read chapters 1 to 7 first. They explain the product and the two main flows. Chapters 13 and 14 list where the code differs from the README and which parts are weak. Read them before an interview, because a good interviewer will find these points. Chapter 15 holds practice questions. Click a question to show the answer.
1 · Overview
What the product does, in simple words.
PolicyLens helps a person understand an insurance policy. Policy PDFs are long and full of legal terms. The user uploads a PDF. The system reads it, saves it in a searchable form, and makes a short summary. The user then asks questions in a chat window. The chat assistant is named IRIS.
The 30-second pitch
“PolicyLens is a RAG application. RAG means retrieval-augmented generation. The system splits a policy into clause-sized chunks and stores each chunk as a vector in Postgres with pgvector. When the user asks a question, the system finds the best chunks, re-ranks them, and sends only those chunks to an LLM. The LLM answers from the chunks only, and the answer streams to the screen.”
Why RAG and not another method?
- Pasting the whole PDF into the LLM costs many tokens, is slow, and the model can miss details in the middle of a long text.
- Fine-tuning a model teaches it style, not facts. Every user has a different policy. You cannot train a model for each upload.
- RAG sends only the 3 most relevant chunks. It is cheap and fast. It can show the source of each answer. It also keeps one user’s policy apart from another’s.
The two main flows
- Ingest (write path). PDF → parse → clean → chunk → embed → store → summary.
- Query (read path). Question → preprocess → embed → vector search → re-rank → build context → LLM → streamed answer.
2 · System map
Which program talks to which.
What each program owns
| Program | Folder | Job |
|---|---|---|
| Frontend | frontend/ | Single-page app. Shows landing page, login, upload, dashboard and chat. Stores the JWT in localStorage. Calls only the Node API. |
| Node backend | backend/ | Sign-up and login. Checks the JWT. Saves the PDF in Supabase Storage. Keeps the policy list and chat history. Forwards ingest and query calls to Python. |
| Python RAG API | api/ + rag_engine/ | Parses the PDF. Chunks, embeds and stores the text. Searches, re-ranks and calls the LLM. Makes the summary. |
| Supabase | supabase_migrations/ | Postgres tables, the pgvector index, the search function, the file bucket and Google OAuth. |
The AI libraries (PyMuPDF, sentence-transformers, tiktoken) are Python tools. The product logic (users, files, history) was easy to build in Express. Node is the “front door” and Python is the “engine room”. The cost is one extra network hop and two codebases to deploy.
3 · Data model
Five tables. Run the SQL files 001 to 005 in order.
policy_id links policies, chunks, summaries and history. The database does not enforce these links (dashed). Solid arrows are real foreign keys to users.How a policy gets its ID
Node creates the ID at upload: policy_id = "<user_id>_<random uuid>". The same string is the file name in Storage (<user_id>/<policy_id>.pdf), the key in policies, and the filter value inside every chunk’s metadata. This is how one policy stays apart from another.
What is stored with each chunk
Each row in policy_chunks has the chunk text, a 768-number embedding, and a metadata JSON object. The JSON comes from ChunkMetadata.to_supabase_dict().
| Field | Meaning |
|---|---|
policy_id | The policy this chunk belongs to. Used as the search filter. |
chunk_id, chunk_index | A random UUID and the order of the chunk in the document. |
source_file | Name of the uploaded file. |
section_name, section_number | The heading, for example “SECTION 3 - EXCLUSIONS”, and the code “SECTION 3”. |
page_number | Estimated page. See chapter 14 for why it is not exact. |
clause_type | One of: coverage, exclusion, definition, deductible, limit, endorsement, schedule, general_condition, unknown. |
coverage_category | fire, flood, theft, liability, medical, vehicle, property, life, travel, or empty. |
deductible_related, limit_related, endorsement_flag, table_chunk | True/false flags set by keyword checks. |
token_count | Size of the chunk in tokens (tiktoken cl100k_base). |
The search function in SQL
CREATE FUNCTION match_policy_chunks(query_embedding vector(768),
match_count int,
filter jsonb DEFAULT '{}')
...
SELECT id, content, metadata,
1 - (embedding <=> query_embedding) AS similarity
FROM policy_chunks
WHERE metadata @> filter -- JSON "contains" check
ORDER BY embedding <=> query_embedding -- cosine distance, smallest first
LIMIT match_count;
<=>is the pgvector cosine distance. Similarity = 1 − distance.metadata @> filterkeeps only rows whose JSON contains the filter. A filter of{"policy_id":"X"}returns chunks of policy X only.- The HNSW index (
m = 16,ef_construction = 64) makes the nearest-neighbour search fast. A GIN index onmetadataspeeds up the filter.
The file 001_vector_store.sql creates vector(768). Its comments and supabase_migrations/README.md say 1536 or 1024. Trust the SQL: the column is 768 wide. The storage bucket policy-pdfs is not created by any migration. You must create it by hand.
4 · Sign-in
Two ways to log in. Both end with a JWT that Node issues.
Email and password
- The browser sends
POST /api/auth/signupwith email, password and name. Login uses/api/auth/login. - Node hashes the password with bcrypt (cost 10) and inserts a row in
users. On login, Node compares the password with the stored hash. - Node signs a JWT with
JWT_SECRET. The payload is{ id, email }. It expires in 7 days. - The browser saves the token in
localStorage["token"].api.jsaddsAuthorization: Bearer …to every request.
Google sign-in
How Node checks a token
The file middleware/auth.js tries two checks in this order.
- Verify the token with
jwt.verify(token, JWT_SECRET). If it works, setreq.userand continue. - If it fails, ask Supabase:
supabase.auth.getUser(token). If Supabase accepts it, setreq.user = { id, email }. - If both fail, return
401.
Other rules in the sign-in code
- Personal e-mail only.
utils/personalEmail.jsallows 10 domains (gmail, hotmail, outlook, yahoo, icloud, live, msn, protonmail, me, mac). The browser enforces this. The server does not. - Password strength meter and 8-character minimum are in the browser.
- Forgot-password screen is a stub. A
TODOcomment says the real flow is not built. - OAuth loading gate.
App.jsxshows “Completing sign-in…” while it reads the URL hash, so the page does not flash the login screen.
5 · Upload and ingest
From a PDF file to searchable chunks.
Step by step, across the services
Node side (backend/src/routes/ingest.js)
- Receive.
multerkeeps the file in memory. Node rejects a file with no name that ends in.pdf. - Name it.
policy_id = <user_id>_<uuid>. - Store it. Upload to the bucket
policy-pdfsat<user_id>/<policy_id>.pdf. Make a signed URL valid for 1 year. - Record it. Insert a row in
policieswithstatus = 'processing'. - Forward it. Post the file to Python
/ingest/uploadwithaxios. Do notawaitit. If it fails, Node only logs the error. - Answer. Return
{ status: 'processing', policy_id, … }at once.
Python side (api/routes/ingest.py)
- Check that the file name ends in
.pdf. Write the file to a temp path. - If the policy already exists and
overwriteis false, delete the temp file and returnskipped. - Add a FastAPI
BackgroundTask. Returnprocessing. - The task runs
IngestionService.ingest(). It always deletes the temp file at the end. - If the result is
successorskipped, it starts a daemon thread for the summary.
The pipeline inside IngestionService
status_tracker.update_status() writes these values in IngestionService. The browser reads them through the status endpoint.Stage 1 · Parse (ingestion/pdf_loader.py)
pymupdf4llm.to_markdown(..., page_chunks=True)returns one markdown text per page.- The loader adds a marker
---PAGE_START:n---before each page and joins the pages. - This replaced LlamaParse (a cloud parser that took 60 to 120 s). The local parser takes about 0.2 s. Set
USE_LLAMAPARSE=trueto switch back. The old code is still in the file. - A scanned PDF with no text layer gives empty text. The local parser has no OCR.
Stage 2 · Clean (ingestion/cleaner.py)
- Page map. Before cleaning,
extract_page_map()records which page each text position belongs to. - Remove noise. Delete lines such as “Page 3 of 20”, a lone number, “Confidential”, a web address, a copyright line, the page markers and rows of
___or***. - Fix hyphenation. Join words split at a line end (
insur-⏎ance→insurance). - Fix spacing. Reduce three or more blank lines to one blank line. Remove trailing spaces.
- Mark headings. Put
##before lines that start with SECTION, PART, ARTICLE or CHAPTER plus a number.
Stage 3 · Chunk (chunking/clause_chunker.py)
This is the most special part of the project. A normal splitter cuts text every N characters and can cut a clause in half. This chunker follows the structure of a policy.
Chunk size by clause type (tokens, from config/constants.py)
| Clause type | Target | Max | Overlap |
|---|---|---|---|
| coverage | 512 | 768 | 64 |
| exclusion | 384 | 512 | 64 |
| definition | 256 | 384 | 32 |
| deductible | 384 | 512 | 64 |
| limit | 384 | 512 | 64 |
| endorsement | 512 | 768 | 64 |
| schedule | 512 | 1,024 | 0 |
| general_condition / unknown | 512 | 768 | 64 |
Definitions are short and exact, so they get small chunks. Schedules are tables, so they get a large limit and no overlap. The “target” value is stored in the config, but the code only reads max and overlap.
How the clause type is found
The code tests keyword lists in a fixed order: exclusion → definition → deductible → schedule → endorsement → limit → coverage → general_condition. It checks the heading first. If nothing matches, it checks the first 600 characters of the body. Exclusion comes first because an exclusion text often contains the word “cover”. A comment in the code warns not to change this order.
What the other metadata flags mean
coverage_category: the first category (fire, flood, theft, …) whose keyword appears in the text.deductible_related: the text contains “deductible”, “excess” or “retention”.limit_related: the text contains “limit”, “maximum”, “sum insured” or “aggregate”.- Token counts use
tiktokenwith thecl100k_baseencoding.
Stage 4 · Embed (embeddings/jina_embedder.py)
- The factory
get_embedder()readsEMBEDDING_PROVIDER. The default isjina. The optionlocalloads a BGE model with sentence-transformers. The optionsopenaiandcohereraiseNotImplementedError. - Jina v3 gives 1024 numbers by default. The code asks for
dimensions: 768. Jina v3 supports shortened vectors (Matryoshka training). This keeps the vector the same size as the database column. - The code sends 100 texts per API call and runs up to 5 calls at the same time. The order of results is kept.
- A retry decorator tries up to 3 times with waits of 1 s, 2 s, 4 s.
- For a question,
embed_query()sends one text.
Stage 5 · Store (vector_store/supabase_store.py)
- Each row is
{ content, embedding, metadata }. The code inserts 50 rows per request, one request after another. policy_exists()andget_policy_chunk_count()count rows wheremetadata->>policy_idequals the ID.- With
overwrite=true, the service deletes the old chunks first. - After the insert, the service sets the status to ready at once. It does not wait for the summary.
How the status endpoint decides
GET /ingest/status/{id} counts chunks in the database. If the count is more than zero, it says ready and 100. If not, it returns the in-memory value from status_tracker. If there is no value, it returns not_found.
status_tracker is a Python dictionary in memory. It is lost when the server restarts. It is not shared between worker processes. There is no failed state. If ingest crashes, the browser waits for ever.
6 · Auto summary
A structured overview made after the upload.
- The ingest task starts a daemon thread that runs
SummaryService.generate(policy_id). The dashboard does not wait for it. - The service searches with a fixed query: “policy overview benefits coverage exclusions surrender death benefit premiums”. It takes the top 10 chunks. There is no re-ranking here.
- It builds a context of up to 3,000 estimated tokens.
- It asks the LLM (max 4,096 tokens) to return only JSON, written in plain English. The prompt sets rules: short sentences, no jargon, explain unavoidable terms in brackets, never copy legal text.
- It removes ``` code fences if the LLM added them, then parses the JSON. If parsing fails, the result has
parse_errorand nothing is saved. - It upserts the result in
policy_summaries(keypolicy_id).
The JSON fields
| Field | Type |
|---|---|
policy_name, policy_type, insurer, uin | text (uin may be null) |
key_benefits, exclusions, important_conditions | list of full sentences |
death_benefit, survival_benefit, surrender_value | 2–3 sentences each |
loan_facility, tax_benefit | 2–3 sentences, or null |
free_look_period | one sentence |
How the browser gets it
- When ingest is ready, the dashboard calls
GET /api/policies/:id/summaryevery 6 s, up to 20 times (2 minutes). A 404 means “not yet”. - While it waits, the page shows a skeleton and the text “this takes about 30–40 s”.
- About 3 to 5 seconds after the summary shows, a small prompt invites the user to open the chat. It hides after 12 s.
The first version made the user wait for the summary before the dashboard opened. The team moved it to a separate thread. The user now sees the dashboard as soon as chunks are stored, and chat works even before the summary exists.
The summary prompt asks for death benefit, survival benefit and surrender value. This suits life insurance. For a motor or health policy, many fields will be empty or null.
7 · Ask a question
From a typed question to a streamed answer.
Step by step
- Preprocess (
retrieval/query_preprocessor.py). Trim spaces. If the text starts with a question word (what, is, are, does, can, how …) and has no “?”, add one and capitalise the first letter. - Pick filters. Always
policy_id. Then simple keyword rules: “flood / water / storm / hurricane” →coverage_category = flood; fire words →fire; theft words →theft; “deductible / excess” →deductible_related = true; “limit / maximum / cap” →limit_related = true. - Embed the cleaned question into one vector (768 numbers).
- Search. Call the SQL function
match_policy_chunkswith the vector,k = 8and the filter. If it returns nothing, repeat with only{policy_id}. - Re-rank (
retrieval/reranker.py). A cross-encoder (cross-encoder/ms-marco-TinyBERT-L-2-v2) scores each pair [question, chunk]. Sort by score. Keep 3. - Build context (
retrieval/context_builder.py). Write each chunk as a block with “Source n”, section name, clause type and score. Estimate tokens as words × 1.3. Stop before the total passes 3,000. Always keep at least one block. - Build the prompt. A system prompt plus a user prompt (below).
- Call the LLM and stream the tokens to the client.
The system prompt (summary of its 7 rules)
- Persona: “PolicyDecoder”, a friendly insurance helper. (The UI calls it IRIS.)
- Use simple, plain English. Short paragraphs.
- Answer only from the given context. Never guess.
- If the answer is not there, say: “I couldn’t find this in your policy document. You may want to contact your insurer directly.”
- Name the section briefly, for example “As per Section 3.1…”. One reference is enough.
- For yes/no questions, start with Yes or No, then explain.
- Aim for 3 to 5 sentences. Never speculate.
The user prompt template
INSURANCE POLICY CONTEXT:
--- Source 1 ---
Section: SECTION 3 - EXCLUSIONS
Clause Type: exclusion
Relevance: 0.912
<chunk text>
...
---
QUESTION: Is flood damage covered?
POLICY ID: <policy_id>
Please answer the question based strictly on the policy context
provided above. Cite specific sections where applicable.
How streaming works end to end
chat_history at the end.What the response contains
POST /query/stream(used by the UI) returns plain text only. It does not return sources.POST /query/(not used by the UI) returns{ answer, policy_id, source_count, sources[] }. Each source hassection,clause_type,page_number,highlight_text(the full chunk),relevance_scoreand a 200-charactersnippet. The design intends the UI to highlight the passage in a PDF viewer. The UI does not do this yet.- Each question is independent. The prompt holds no earlier messages, so follow-up questions such as “and what about the limit?” lose their context.
The API default is k = 8. The service keeps top_k_rerank = 3. The README says “re-rank to 5” and “about 3,400 tokens”. The code is the truth: 8 in, 3 out, 3,000 tokens at most.
8 · Frontend
React 19, Vite 7, Tailwind 4, lucide icons, Supabase JS.
App.jsx keeps appState and mirrors it in the URL hash (#login, #dboard …) with history.pushState. It does not use react-router, even though the package is installed.Main files
| File | Job |
|---|---|
App.jsx | Screen switch. Dark-mode flag. Listens to Supabase onAuthStateChange, blocks non-personal e-mail, swaps a Google session for a Node JWT. |
api.js | fetchApi() adds the Bearer token, sets JSON headers (not for FormData), parses JSON or text, and throws on errors. streamApi() reads a response body chunk by chunk and calls onToken. The base URL is VITE_API_URL, else http://localhost:4000/api on localhost, else /api. |
Hero, Navbar, Footer, About | Landing page. |
Auth/Login.jsx | Login, sign-up and forgot-password views. Password strength bar. Google button. |
Dashboard/UploadModal.jsx | Drag-and-drop. Accepts PDF only, 50 MB maximum. Shows four stages and a progress bar. Polls status every 2 s. |
Dashboard/Dboard.jsx | Main screen (about 1,000 lines). Sidebar of past policies. Summary cards. A floating, resizable chat widget with one message list per policy. |
Dashboard/Chatbot.jsx | Full-page chat route. Can export the chat as a PDF with jsPDF (loaded from a CDN). Without a file prop it queries the policy ID "test", so it works only as a demo. |
Timers in the browser
| What | Every | Stops when |
|---|---|---|
| Ingest status poll | 2 s | status is ready (no time limit) |
| Summary poll | 6 s | summary found, or 20 tries |
| Chat prompt bubble | 3–5 s delay, then 12 s visible | chat opens |
Details worth knowing
- React StrictMode is off. A comment in
main.jsxsays it ran the upload effect twice and sent two uploads.UploadModalalso has three guard fixes (acompletedRef, a ref for the callback, and a narrow dependency list). The cleaner fix is an effect that is safe to run twice. - Theme. Colours are two JavaScript objects,
LIGHTandDARK, used as inline styles. The landing page uses Tailwinddark:classes. - Chat state lives in a map
{ [policy_id]: messages[] }in memory. A page reload clears it. The history API exists, but the dashboard does not read it. - History list comes from
GET /api/policies. It shows the file name, the date and, once loaded, the policy name from the summary. - The file picker accepts
.pdf,.doc,.docx, but the server accepts PDF only.
9 · API tables
Every route in the code. The README lists some paths wrongly.
Node backend (port 4000). All paths start with /api.
| Method | Path | Auth | What it does |
|---|---|---|---|
| GET | /health | none | Liveness text. |
| POST | /auth/signup | none | Create a user. Return token + user. |
| POST | /auth/login | none | Check password. Return token + user. |
| POST | /auth/google | none | Find or create a user by e-mail. Return token. See chapter 14. |
| POST | /ingest/upload | JWT | Multipart field file. Store PDF, add policy row, forward to Python. |
| GET | /ingest/status/:policy_id | JWT | Ask Python for progress. When ready, update the policy row. |
| GET | /policies | JWT | List the user’s policies, newest first. |
| DELETE | /policies/:policy_id | JWT | Delete the file and the policy row only. |
| GET | /policies/:policy_id/summary | none | Read the stored summary. 404 if not ready. |
| POST | /query | JWT | Proxy to Python. Save question, answer and sources. |
| POST | /query/stream | JWT | Proxy the stream. Save the answer at the end. |
| GET | /history | JWT | Query parameters: policy_id, q (search text), page, limit. |
| GET / DELETE | /history/:id | JWT | One entry by row ID. |
| DELETE | /history?policy_id=… | JWT | Clear all entries, or those of one policy. |
Python API (port 7860 on HF, 8000 locally)
| Method | Path | What it does |
|---|---|---|
| POST | /ingest/upload | Multipart file, policy_id, overwrite. Start background ingest. |
| POST | /ingest/ | JSON with pdf_url. This route passes the URL to the loader as a file path, so it fails. Treat it as unfinished. |
| GET | /ingest/status/{id} | Chunk count, status, progress, message. |
| POST | /ingest/summary/{id} | Make or remake the summary now (slow, waits for the LLM). |
| GET | /ingest/summary/{id} | Read the stored summary. |
| POST | /query/ | { question, policy_id, k = 8 } → answer + sources. |
| POST | /query/stream | Same input. Returns text/plain tokens. |
| GET | /health/ | Checks Supabase. Returns ok or degraded. |
10 · Code patterns
The design ideas you can name in an interview.
| Pattern | Where | Why it helps |
|---|---|---|
| Abstract base + factory | BaseEmbedder / get_embedder(), BaseLLM / get_llm(), BaseVectorStore / get_vector_store() | Change provider with an environment variable. Add a new provider without touching the pipeline. |
| Lazy imports | Inside factory functions and service constructors | Heavy libraries (torch, Gemini SDK) load only when used. Server starts faster. |
| Load once at start-up | lifespan() in api/main.py builds one QueryService in app.state | The cross-encoder model loads once, not for every request. |
| Retry with backoff | utils/retry.py @with_retry | Handles short API failures. LLM: 4 tries, 5 s start, ×2. Jina and Supabase: 3 tries, 1 s start, ×2. |
| Singleton status store | StatusTracker | One shared dictionary for progress. Simple, but not durable. |
| Fire and forget | Node → Python upload call; Python summary thread | Keeps the request fast. The caller polls for the result. |
| Typed config | pydantic-settings Settings with @lru_cache | All keys come from .env. One cached object. |
| Typed metadata | ChunkMetadata (Pydantic) with to_supabase_dict() | One schema for every chunk. Converts enums, dates and null to JSON-safe values. |
| Proxy / BFF | Node routes query.js, ingest.js | The browser sees one API. Keys and the Python URL stay on the server. |
get_llm() supports gemini, groq, kimi and openai (the last one re-uses the Kimi class with an OpenAI client). The setting LLM_MODEL is only printed in a log. Each provider reads its own model setting (GROQ_MODEL, GEMINI_MODEL, KIMI_MODEL). The files local_llm.py, openai_llm.py, openai_embedder.py, query_request.py and query_response.py are empty.
11 · Config and deploy
Where each part runs and which settings it needs.
| Part | Runs on | How it gets there |
|---|---|---|
| Python RAG API | Hugging Face Space devjhawar/policylens-rag-api (Docker, port 7860) | GitHub Action sync-to-hf.yml: on each push to main, huggingface_hub.upload_folder sends the repo (not .git, .github, .env, *.zip). The README header (sdk: docker, app_port: 7860) is the Space config. |
| Node backend | Its own Docker image (backend/Dockerfile, Node 18 Alpine, non-root user, port 7860) | Prepared for Hugging Face Spaces. npm ci --only=production, then npm start. |
| Frontend | Vercel (the examples use policylens-ai.vercel.app) | VITE_API_URL points to the Node API. VITE_BASE_PATH sets the Vite base path. |
| Database / files / Google OAuth | Supabase project | Run SQL 001–005. Create the policy-pdfs bucket. Enable the Google provider. |
Python .env
| Variable | Default / note |
|---|---|
SUPABASE_URL, SUPABASE_SERVICE_KEY | Required. Use the service-role key. |
LLAMA_CLOUD_API_KEY | Required by the settings class even when not used. Set any text. |
LLM_PROVIDER | groq (also gemini, kimi, openai) |
GROQ_API_KEY, GROQ_MODEL | Model default openai/gpt-oss-20b |
GEMINI_API_KEY, MOONSHOT_API_KEY, OPENAI_API_KEY | Only for those providers. |
EMBEDDING_PROVIDER, JINA_API_KEY | jina (or local). Jina key required for Jina. |
USE_LLAMAPARSE | false. Set true for the cloud parser. |
CORS_ORIGINS | * by default. Set real origins in production. |
DEBUG | false |
Node backend/.env
PORT (4000), SUPABASE_URL, SUPABASE_SERVICE_KEY, JWT_SECRET, PYTHON_API_URL, CORS_ORIGINS.
Frontend .env
VITE_API_URL, VITE_SUPABASE_URL, VITE_SUPABASE_ANON_KEY.
Run it on your machine
- Run SQL files
001to005in the Supabase SQL editor. Create the bucketpolicy-pdfs. - Python:
pip install -r requirements.txt, fill.env, thenpython -m uvicorn api.main:app --port 8000 --reload. - Node:
cd backend && npm install, fillbackend/.env, thennpm run dev. - Frontend:
cd frontend && npm install && npm run dev(port 5173).
The Dockerfile caches the model ms-marco-MiniLM-L-6-v2 at build time, but the code loads ms-marco-TinyBERT-L-2-v2. The server therefore downloads the real model at start-up. The HEALTHCHECK calls /, which has no route, and docker-compose.yml maps port 8000 while the image listens on 7860. Fix these three before relying on them.
12 · Tests
What exists and what is missing.
| File | What it checks |
|---|---|
test_chunker.py | Chunks are made, none are empty, policy IDs and indexes are right, exclusion and coverage types are found, table chunks are flagged, token counts stay in limits, page map is used. |
test_retrieval.py | Query preprocessing and filters. The retriever calls embed and search, and falls back on empty results. Context builder format. |
test_vector_store.py | Insert, search, delete and exists, with a mocked Supabase client. |
test_integration.py | Full query flow and ingestion skip/overwrite with mocks. Response format, highlight fields, page map. |
test_evaluation.py | Small checks on a sample policy: chunk quality, query cleaning, context token limit, at most 5 sources. |
test_e2e.py | A live script (not a pytest file). Needs real keys. Ingests a sample, asks 3 questions, deletes it. |
Run them with python -m pytest rag_engine/tests/ -v. I read these files but did not run them.
There are no tests for the Node backend or the React app. There is no measure of answer quality (recall, faithfulness). The “evaluation” file checks plumbing, not retrieval accuracy.
13 · README vs code
The code is the truth. Know these differences before someone else finds them.
| Topic | README says | Code does |
|---|---|---|
| LLM | Kimi (kimi-k2.5) | Default provider is Groq, openai/gpt-oss-20b. Kimi is legacy. |
| Re-ranker | BGE reranker | cross-encoder/ms-marco-TinyBERT-L-2-v2 |
| Re-rank size | top 5, context about 3,400 tokens | top 3, context up to 3,000 estimated tokens |
| Summary input | top 15 chunks | top 10 chunks, 3,000-token context |
| Chunker file | chunker.py | clause_chunker.py and table_chunker.py |
| Auth paths | /auth/signup, /auth/login | /api/auth/…, plus /api/auth/google |
| History API | POST /api/history, DELETE /api/history/:policy_id | No POST. The query routes save history. Delete is by row ID, or by ?policy_id=. |
| Python version | 3.11 | pyproject.toml says ≥ 3.12. The Dockerfile uses 3.11. |
| Health text | — | Reports model “BAAI/bge-base-en-v1.5 + kimi-k2.5”. Both are out of date. |
| Embedding size | 768 (Jina) | True, by shortening Jina v3. SQL comments and the migrations README still say 1536 / 1024. |
| Unused settings | — | MIN_CONFIDENCE_THRESHOLD, RERANKER_MIN_SCORE, MMR_LAMBDA_MULT, DEFAULT_TOP_K are defined and never read. |
| Extra files | — | api.zip (old copy of api/) and a root node_modules/ are committed. |
14 · Risks and bugs
Found by reading the code. Rated by how much harm they can do.
Items 1 to 4 can expose user data. If you present this project, say you know about them and say how you would fix them. That shows good judgement.
| # | Level | Problem | Fix |
|---|---|---|---|
| 1 | high | POST /api/auth/google trusts the e-mail in the request body. Anyone can send any e-mail and get a 7-day JWT for that account. | Send the Supabase access token. Node calls supabase.auth.getUser(token) and uses the e-mail from the result. |
| 2 | high | GET /api/policies/:id/summary has no login check. A comment says “bypass user check for debugging”. | Add authMiddleware and check that the policy belongs to req.user.id. |
| 3 | high | No owner check on query and status routes. Any logged-in user who knows a policy_id can ask questions about it. The Python API has no login at all and is public on Hugging Face. | Check ownership in Node. Add a shared secret header between Node and Python. Limit CORS. |
| 4 | high | frontend/.env.example has a value that looks like a Google OAuth client secret (it starts with GOCSPX-) in the anon-key line. The Supabase project URL is also hard-coded in supabaseClient.js. | Treat the secret as leaked and rotate it. Use a placeholder in the example file. |
| 5 | medium | No failed state. If ingest crashes, or Node’s forward call fails, the row stays “processing” and the browser polls with no time limit. | Catch errors, set status failed, return it, and add a poll timeout. |
| 6 | medium | Progress is kept in process memory only. | Store status in a database table or Redis. |
| 7 | medium | The store retries the whole add_chunks call. A failure at batch 3 repeats batches 1 and 2 and creates duplicates. A part-way failure also leaves chunks behind, so the status endpoint says “ready” for a half-stored policy. | Retry per batch. Delete the policy’s chunks on failure. Or insert in one transaction. |
| 8 | medium | Deleting a policy removes only the file and the policies row. Chunks, summary and chat history stay in the database. | Also delete rows in policy_chunks, policy_summaries, chat_history. |
| 9 | medium | Keyword filters can hide good chunks. A chunk has only one coverage_category (the first match), and the filter needs an exact match. The fallback runs only when there are zero rows. | Use filters only to boost, or drop them. Test retrieval with and without filters. |
| 10 | medium | Chat has no memory. Sources never reach the UI because the stream route returns text only. | Rewrite follow-ups with earlier messages. Send sources as a final JSON line or use SSE. |
| 11 | medium | JWT is stored in localStorage (XSS can read it). The personal-e-mail rule and password rules exist only in the browser. No rate limit. The history search text goes into a PostgREST .or() string. | Use httpOnly cookies or short tokens with refresh. Repeat checks on the server. Add express-rate-limit. Escape the search text. |
| 12 | low | Page numbers are approximate. Each chunk gets the page found for the previous section, and the page map uses positions from before cleaning. | Compute the page from the start of the current section on the raw text. |
| 13 | low | rag_engine/main.py has from __future__ import annotations after other code. Python rejects this, so the CLI cannot start. | Move the import to the top. |
| 14 | low | Jina v3 supports a task setting (query vs passage). The code does not set it. | Pass retrieval.passage for chunks and retrieval.query for questions, then compare results. |
| 15 | low | The summary prompt suits life insurance. The coverage categories suit property insurance. Health and motor policies fit neither well. | Detect the policy type first, then choose a prompt. |
| 16 | low | Scanned PDFs return no text. The .pdf check is case-sensitive. The upload form accepts Word files. POST /ingest/ (URL mode) fails. | Add OCR or the LlamaParse fallback. Use lower-case checks. Remove the dead options. |
15 · Interview Q&A
Say the answers in your own words. Answers use “I” and “we”. Change them to match the part you personally built.
Explain PolicyLens in 30 seconds.basics
PolicyLens lets a user upload an insurance policy PDF and ask questions about it in plain English. The system splits the PDF into clause-sized chunks and stores them as vectors in Postgres with pgvector. For each question it finds the best chunks, re-ranks them, and gives them to an LLM that answers only from that text. The answer streams to the screen. It also makes a structured summary on upload.
Why RAG? Why not give the full PDF to the LLM, or fine-tune a model?basics · rag
A policy can be 50 pages. Sending all of it on every question costs many tokens, adds delay, and the model can miss details in the middle. Fine-tuning changes how a model writes, not what it knows, and we cannot train a model for every upload.
With RAG we send only the 3 most relevant chunks. It is cheaper and faster. We can show which section the answer came from. Each user’s data stays in the database and is filtered by policy ID.
Why do you have both a Node and a Python backend?basics · backend
The AI work needs Python libraries: PyMuPDF for PDFs, sentence-transformers for the re-ranker, tiktoken for token counts. The product work (users, file storage, history) was quick to build in Express. Node is the single public API for the browser. Python is an internal engine that Node calls.
The trade-off is one more network hop, two deployments, and two places where auth can go wrong. A smaller team could put everything in FastAPI. I would consider that if I started again.
Walk me through what happens when I upload a PDF.basics
- The browser posts the file to Node. Node checks the JWT and makes a policy ID.
- Node saves the file in Supabase Storage and adds a row to
policies. - Node forwards the file to Python without waiting, and answers the browser at once.
- Python runs a background task: parse with pymupdf4llm, clean, chunk by clause, embed with Jina in batches of 100, insert into pgvector in batches of 50.
- The browser polls the status every 2 seconds and shows a progress bar. When chunks exist, the status is ready.
- A separate thread then makes the summary with the LLM. The dashboard polls for it every 6 seconds.
Walk me through what happens when I ask a question.basics
- The browser posts
{ question, policy_id }to Node. Node addsk = 8and forwards it to Python. - Python cleans the question, picks metadata filters, and embeds it with Jina.
- pgvector returns the 8 nearest chunks for that policy (cosine distance).
- A cross-encoder scores each chunk against the question. We keep the top 3.
- We build a context block, add the system prompt, and call the LLM with streaming.
- Tokens flow back through FastAPI and Node to the browser. Node saves the full answer in
chat_historywhen the stream ends.
How does your chunking work, and why is it “clause-aware”?rag
A fixed-size splitter can cut a clause in the middle, and a half clause can reverse its meaning (“…does not cover” without the rest). Insurance text has structure: sections, numbered clauses, tables. My chunker splits at SECTION, PART, ARTICLE, CHAPTER and “##” headings first. A section that fits the size limit stays whole. A larger one is split at clause markers such as 3.1 or (a), with a 3-line overlap. Only if a piece is still too large do we split by tokens.
Each chunk also gets metadata: section name, clause type, coverage category, flags for deductible and limit, page and token count.
Why different chunk sizes for different clause types?rag
Different clauses have different shapes. A definition is short and exact, so a 384-token maximum keeps it focused. Coverage and endorsement clauses are longer and need more room (768). A schedule is a table, so it gets 1,024 tokens and no overlap. Smaller chunks give sharper matches. Larger chunks keep more context. The clause type lets us choose the right balance for each part.
How do you handle tables?rag
If 30% or more of the lines in a section hold a “|”, the section goes to the table chunker. It finds the header row and the separator row, groups the data rows into batches up to 1,024 tokens, and repeats the header in every batch. This way a row like “Fire | $500” never loses its column names. These chunks are marked as type schedule with table_chunk = true.
What is an embedding? Why 768 dimensions?rag
An embedding is a list of numbers that represents the meaning of a text. Texts with similar meaning have vectors that point in similar directions, so we can search by meaning instead of by exact words. “Is fire covered?” can match a chunk that says “loss caused by fire, smoke or explosion”.
Jina v3 gives 1024 numbers by default. It supports shortened vectors, so we ask for 768. This matches the database column and uses less storage. Smaller vectors lose a small amount of accuracy but search faster. If I changed the size, I would need to change the column and re-embed all chunks.
What does pgvector do? What are cosine distance and HNSW?rag
pgvector adds a vector column type and distance operators to Postgres. Cosine distance (<=>) measures the angle between two vectors. A small angle means similar meaning. The SQL function returns 1 − distance as the similarity.
An exact search compares the query with every row. HNSW is an index that links each vector to near neighbours in layers. A search walks the graph and finds close matches much faster, with a small chance of missing the true best one. I used m = 16 and ef_construction = 64, which are common defaults.
Why a re-ranker? Why retrieve 8 and keep 3?rag
Vector search uses a bi-encoder: the question and each chunk are turned into vectors separately. This is fast but rough. A cross-encoder reads the question and a chunk together and gives one relevance score. It is more accurate but slower, so we run it only on the 8 candidates.
We keep 3 because the context must be short. Less text means lower cost, lower delay and less noise for the LLM. The numbers 8 and 3 were reduced from 8 and 5 to cut response time. I would tune them with a test set.
How do you reduce hallucination?rag
- The system prompt says: answer only from the context, never guess, and use a fixed sentence when the answer is missing.
- Temperature is low (0.3).
- The context holds only the 3 best chunks, so there is less unrelated text.
- The non-stream API returns the source chunks so a person can check them.
The limit: nothing checks the answer after generation. A next step is a faithfulness check, where a second model confirms that each claim appears in the sources.
How do you keep one user’s policy separate from another’s?rag sec
Every chunk stores policy_id in its metadata. Every search passes a filter { policy_id }, and the SQL uses metadata @> filter. The policy ID starts with the user ID and a random UUID, so it is hard to guess.
I should be honest about the gap: the server does not yet check that the logged-in user owns that policy ID. The fix is a lookup in the policies table before every query. I would also move to row-level security in Supabase.
How do you measure answer quality?rag
Today the tests check the pipeline: chunk sizes, filters, formats. They do not measure quality. To measure it I would build a golden set of 30 to 50 questions with the correct answer and the correct source chunk for several real policies. Then I would track recall@k (is the right chunk in the top 8?), MRR, and answer faithfulness (using a judge model or manual review). I would run it when I change the chunker, the embedding model or the prompts.
What happens with a scanned PDF?rag
pymupdf4llm reads the text layer. A scan has none, so the text is empty and nothing useful is stored. The project keeps the old LlamaParse path behind USE_LLAMAPARSE=true, which does OCR in the cloud but takes 60 to 120 seconds. A better design detects an empty text layer and runs OCR only for those files.
Your chat does not remember earlier messages. How would you add memory?rag
Right now each question is searched alone, so “and what is the limit?” has no subject. I would send the last few messages with the request. Then I would do query rewriting: ask a small model to turn the follow-up into a stand-alone question (“What is the fire coverage limit?”) before embedding. The rewritten question goes to search. The recent messages also go in the final prompt.
What is a weakness of your metadata filters?rag
The filters are keyword rules. A question with “flood” adds coverage_category = flood. But a chunk has only one category, the first match, so a chunk about fire and flood may be tagged “fire” and be excluded. The fallback runs only if the search returns zero rows. A safer design uses the filter as a boost, or runs a search with and without it and merges the results.
What if a PDF contains text that tries to give the LLM orders (prompt injection)?rag · sec
The policy text is untrusted input. The model has no tools, so it cannot take actions, which limits the damage. The worst case is a misleading answer. I would wrap the context in clear delimiters, tell the model that text inside them is data and not instructions, and scan the output for odd content. For a bigger system I would also log sources for audit.
How does the summary feature work?rag
After ingest, a thread runs a fixed search query (“policy overview benefits coverage exclusions surrender death benefit premiums”), takes the 10 best chunks, and asks the LLM for a JSON object with fields such as key benefits, exclusions, death benefit and free-look period, in plain English. The code strips code fences, parses the JSON and saves it in policy_summaries. If parsing fails, it saves nothing. The prompt suits life insurance best.
Why background tasks and polling? Why not wait for the result?backend
Ingest takes seconds to minutes: parse, many embedding calls, many inserts. A normal HTTP request would risk a timeout, and the user would see nothing. With a background task the upload returns fast and the UI shows real progress. Polling is simple and works on any host.
The alternatives are server-sent events or WebSockets (push instead of poll) and a job queue such as Celery, RQ or Cloud Tasks. A queue also survives restarts and allows retries, which the current in-process tasks do not.
What design patterns did you use?backend
Abstract base class plus factory for the LLM, embedder and vector store, so I can switch provider with an environment variable. Lazy imports for heavy libraries. A start-up hook that loads the query service once. A retry decorator with exponential backoff. A singleton status tracker. A backend-for-frontend proxy in Node. Pydantic for settings and chunk metadata.
How does retry with backoff work in your code?backend
The @with_retry decorator calls the function, and if it raises, waits and tries again. The wait grows each time: 1 s, 2 s, 4 s for Jina and Supabase (3 tries); 5 s, 10 s, 20 s for LLM calls (4 tries). This helps with short rate limits and network errors. One known problem: on the Supabase insert, a retry repeats batches that already succeeded. I would retry per batch instead.
Why is the status tracker in memory? What breaks?backend
It was the simplest thing that worked on one server. It breaks on restart (progress is lost), with several workers (each has its own dictionary), and it has no failed state. The fix is a status column or table in the database, or Redis. The chunk count already serves as a durable “ready” signal, which is why the endpoint checks it first.
Why Supabase pgvector and not Pinecone or Chroma?backend
One database holds users, policies, chat history and vectors. I can filter vectors with normal JSON and SQL conditions, join them with other data, and use one free-tier service. There are fewer moving parts. A dedicated vector database can scale to many millions of vectors with more tuning options. For this size, Postgres is enough.
Why Jina API instead of a local embedding model?backend · rag
The Python service runs on a free Hugging Face Space with a small CPU. A local model needs torch and a large download, and it embeds slowly on CPU. The Jina API is fast, has a free tier (about 1 million tokens per month) and gives strong quality. The cost is a network call, a third party that sees the text, and a limit on usage. The code keeps a local BGE option for offline use.
Why can you swap the LLM so easily?backend
All LLM classes follow the BaseLLM interface: complete, stream and the same with message lists. Groq and Kimi use the OpenAI-compatible API, so they share most of the code. The factory picks the class from LLM_PROVIDER. The query service never knows which provider it uses. Gemini needs a small adapter to convert message formats.
How do you show streaming text in React?frontend
I use fetch and read response.body.getReader() in a loop. Each chunk is decoded with TextDecoder and passed to a callback. The callback updates state: text = text + token for the AI message with a known ID. React re-renders and the bubble grows. I used a plain text stream, not SSE, to keep it simple.
Why a hash-based state router and not react-router?frontend
The app has only five screens and no nested routes. A state variable plus history.pushState with a hash works on any static host without server rewrites. It also let us handle the OAuth redirect, which returns tokens in the URL hash. The cost is manual work for deep links and no route-level code splitting. react-router is installed but unused. If the app grows, I would switch to it.
Why did you remove React StrictMode? Was that the right fix?frontend
In development StrictMode runs effects twice. The upload effect then sent two uploads. I removed StrictMode and added guards (a ref flag, a ref for the callback, a narrow dependency list). It works, but the better fix is an effect that is safe to run twice: for example, start the upload from a button click handler instead of an effect, and cancel with an AbortController on cleanup.
How do the upload progress bar and the summary loading work?frontend
After the upload call returns a policy ID, a timer calls the status endpoint every 2 seconds and sets the bar to the progress number from the server. When status is ready, the app moves to the dashboard. Then another timer asks for the summary every 6 seconds, up to 20 times, while a skeleton shows. The chat button is already usable at that point.
How does authentication work?security
Email and password: bcrypt hashes the password, and Node returns a JWT signed with a secret, valid for 7 days. Google: the browser signs in through Supabase, then asks Node for its own JWT. Every protected route runs a middleware that verifies the JWT, and falls back to checking a Supabase token. The browser stores the token in localStorage and sends it as a Bearer header.
What are the biggest security problems in the project?security
- The Google login route issues a JWT for any e-mail without checking the Supabase token. Fix: verify the token with Supabase on the server.
- The summary route has no login check. Fix: add the middleware and an owner check.
- No ownership check on query and status, and the Python API is public. Fix: check owner in Node, add a shared secret, restrict CORS.
- An example env file holds what looks like a Google client secret. Fix: rotate it and remove it from the repo history.
I would also add rate limits, server-side validation, and delete all data for a policy when the user deletes it.
JWT in localStorage or a cookie?security
localStorage is easy, but any script on the page (an XSS bug) can read the token. An httpOnly cookie cannot be read by scripts, but then you need CSRF protection (SameSite and a CSRF token). The most robust choice is a short-lived access token in memory plus a refresh token in an httpOnly cookie. We used localStorage for speed of development, and I know the risk.
Insurance documents are private. How do you handle privacy?security · ops
Today: files are in a private Supabase bucket and shared by signed URL (valid for 1 year). Chunk text goes to Jina, and question plus chunks go to the LLM provider. So data leaves our system. I would shorten URL life, tell users clearly, offer delete-everything, remove personal data before sending where possible, and choose providers with no-training and no-retention terms. For strict customers, I would use the local embedding model and a self-hosted LLM.
Where is the bottleneck? How would you scale it?scale
PDF parsing is fast (0.2 s). The slow parts are external: the embedding calls during ingest, and the LLM calls (the summary takes 30 to 40 s). For many users I would: move ingest to a job queue with separate workers; run several API instances behind a load balancer; keep status in a shared store; cache embeddings of common questions; add connection pooling for Postgres; and add rate limits per user. The re-ranker runs on CPU, so I would watch its latency and could move it to a GPU or a hosted API.
What does it cost to run?ops
The prototype uses free tiers: Hugging Face Space, Vercel, Supabase, Jina (about 1M tokens per month) and Groq. One upload costs a few embedding calls and one LLM call for the summary. One question costs one embedding call and one LLM call. The main cost risk is the LLM at scale. Short context (3 chunks) keeps tokens low.
How do you deploy it?ops
The Python API is a Docker image on a Hugging Face Space. A GitHub Action uploads the repo to the Space on every push to main. The Node backend has its own Dockerfile. The React app deploys on Vercel with an environment variable for the API URL. Secrets are environment variables, never committed. Known gaps: the Dockerfile caches a different re-ranker than the code loads, the health check points to a missing route, and there is no automatic test step in the pipeline.
Tell me about a performance improvement you made.story
PDF parsing. The first version used LlamaParse, a cloud service that took 60 to 120 seconds per file. I replaced it with pymupdf4llm, which runs locally in about 0.2 seconds, and made the output format identical (same page markers) so the chunker did not change. I kept the old path behind a flag as a rollback. Other changes: summary moved to its own thread so the dashboard opens sooner; embeddings sent in parallel batches; retrieval reduced from 8-and-5 to 8-and-3 to cut latency.
What was a hard bug you fixed?story
Duplicate uploads. In development, React ran the upload effect twice and the same file was sent twice. I found it by watching the network tab and the double log lines. I fixed it with a ref guard and by removing a changing callback from the dependency list, and I left comments in the code explaining each fix. I now prefer to start uploads from a click handler, not an effect.
What would you do next if you had two more weeks?story
- Fix the security items: verified Google login, owner checks, a Node-to-Python secret, secret rotation.
- Add a failed state, a durable status store and a job queue.
- Build an evaluation set and track recall@k and faithfulness.
- Add conversation memory with query rewriting.
- Show sources in the chat and jump to the page in the PDF (the API already returns page and highlight text).
- Support more policy types with a detected type and matching prompts.
What would you do differently if you started again?story
I would write the evaluation set first, so every choice (chunk size, k, re-ranker) has a number behind it. I would use one backend language, or define the contract between Node and Python early. I would design auth and ownership checks at the start instead of adding them later. And I would keep the README in step with the code, because it already differs in several places.
16 · Numbers to know
Short facts for fast recall.
| Item | Value |
|---|---|
| PDF parse time | about 0.2 s (LlamaParse was 60–120 s) |
| Chunks per 20 pages | about 31 |
| Chunk max size | 384 to 1,024 tokens, by clause type |
| Table-heavy rule | 30% of lines contain “|” |
| Embedding size | 768 numbers (Jina v3 shortened from 1024) |
| Embedding batch / threads | 100 texts per call, 5 calls at once |
| Insert batch | 50 rows per call |
| HNSW settings | m = 16, ef_construction = 64, cosine |
| Search / keep | 8 candidates → 3 chunks |
| Context limit | 3,000 estimated tokens (words × 1.3) |
| LLM settings | temperature 0.3, max 4,096 output tokens |
| Summary input | 10 chunks, 3,000-token context |
| Summary time | about 30–40 s |
| Ingest poll / summary poll | 2 s / 6 s (20 tries) |
| JWT life · bcrypt cost | 7 days · 10 |
| Signed file URL life | 1 year |
| Upload limit (browser only) | 50 MB, PDF |
| Node → Python upload timeout | 15 s |
| Retry (Jina, Supabase / LLM) | 3 tries from 1 s / 4 tries from 5 s, ×2 each time |
17 · Glossary
Terms in simple words.
- RAG
- Retrieval-augmented generation. Find text first, then let an LLM write an answer from that text.
- LLM
- Large language model. A program that writes text, for example gpt-oss, Gemini or Kimi.
- Chunk
- A small piece of the document, such as one clause, that is stored and searched alone.
- Token
- A small piece of text (part of a word) that models count. About 1 token is ¾ of an English word.
- Embedding
- A list of numbers that stands for the meaning of a text.
- Vector search
- Finding the stored embeddings that are closest to the question embedding.
- Cosine similarity
- A score for how much two vectors point the same way. 1 means the same direction.
- pgvector
- A Postgres add-on that stores vectors and searches them.
- HNSW
- An index that makes nearest-neighbour search fast by walking a graph of close points.
- Bi-encoder
- Turns the question and the chunk into vectors separately. Fast. Used for the first search.
- Cross-encoder
- Reads the question and the chunk together and scores them. Slower, more exact. Used to re-rank.
- Re-ranking
- Sorting the first results again with a better model.
- Matryoshka embedding
- A vector trained so its first N numbers still work alone. This lets Jina v3 give 768 instead of 1024.
- Metadata filter
- A rule that limits a search to rows with given values, such as one
policy_id. - Streaming
- Sending the answer in small parts while the model writes it.
- Polling
- The client asks the server again and again until a job is done.
- Background task
- Work that continues after the HTTP answer is sent.
- JWT
- JSON Web Token. A signed string that proves who the user is.
- bcrypt
- A slow password-hashing method. It makes stolen hashes hard to crack.
- Deductible
- The part of a claim the policyholder pays first.
- Exclusion
- A thing the policy does not cover.
- Free-look period
- Days after purchase when the user can cancel and get a refund.
- Surrender value
- Money the user gets if they cancel a life policy early.
- IRIS
- The name of the chat assistant in the UI. (The system prompt calls it “PolicyDecoder”.)