Codebase study guide

PolicyLens

An AI web app that reads an insurance policy PDF and answers questions about it in plain English. This guide explains how every part works, shows where data moves, and lists the questions an interviewer is likely to ask.

Repo: arpanz/policylens Read at commit: 9a5db38 Style: Simplified Technical English
How to use this guide

Read chapters 1 to 7 first. They explain the product and the two main flows. Chapters 13 and 14 list where the code differs from the README and which parts are weak. Read them before an interview, because a good interviewer will find these points. Chapter 15 holds practice questions. Click a question to show the answer.

1 · Overview

What the product does, in simple words.

PolicyLens helps a person understand an insurance policy. Policy PDFs are long and full of legal terms. The user uploads a PDF. The system reads it, saves it in a searchable form, and makes a short summary. The user then asks questions in a chat window. The chat assistant is named IRIS.

The 30-second pitch

“PolicyLens is a RAG application. RAG means retrieval-augmented generation. The system splits a policy into clause-sized chunks and stores each chunk as a vector in Postgres with pgvector. When the user asks a question, the system finds the best chunks, re-ranks them, and sends only those chunks to an LLM. The LLM answers from the chunks only, and the answer streams to the screen.”

Programs3 servicesReact UI, Node API, Python RAG API
DatabaseSupabasePostgres + pgvector + Storage
PDF parserpymupdf4llmlocal, about 0.2 s
EmbeddingsJina v3768 numbers per chunk
Re-rankerCross-encodersentence-transformers
LLM (default)Groq gpt-oss-20bGemini, Kimi, OpenAI optional

Why RAG and not another method?

  • Pasting the whole PDF into the LLM costs many tokens, is slow, and the model can miss details in the middle of a long text.
  • Fine-tuning a model teaches it style, not facts. Every user has a different policy. You cannot train a model for each upload.
  • RAG sends only the 3 most relevant chunks. It is cheap and fast. It can show the source of each answer. It also keeps one user’s policy apart from another’s.

The two main flows

  1. Ingest (write path). PDF → parse → clean → chunk → embed → store → summary.
  2. Query (read path). Question → preprocess → embed → vector search → re-rank → build context → LLM → streamed answer.

2 · System map

Which program talks to which.

Browser React 19 · Vite 7 Tailwind 4 Home · Login Upload · Dashboard IRIS chat widget JWT in localStorage Node backend Express 5 · port 4000 /api/auth/* /api/ingest/* /api/query[/stream] /api/policies* /api/history* checks JWT on each call bcrypt · multer · axios Python RAG API FastAPI · port 7860 (HF) /ingest/* /query/ /query/stream /health/ parse · chunk · embed search · re-rank · LLM no login check here Supabase Postgres + pgvector users policies policy_chunks policy_summaries chat_history Storage bucket policy-pdfs Auth: Google OAuth Jinaembeddings GroqLLM REST + JWT HTTP proxy axios vectors summaries users · PDF files · policy rows · chat history Google sign-in (Supabase JS client) embed chat
The browser only talks to Node. Node holds the login and the product data. Python holds the AI work. Both use the same Supabase project. Dashed box = third-party API. Dashed arrow = sign-in that skips Node until the token exchange.

What each program owns

ProgramFolderJob
Frontendfrontend/Single-page app. Shows landing page, login, upload, dashboard and chat. Stores the JWT in localStorage. Calls only the Node API.
Node backendbackend/Sign-up and login. Checks the JWT. Saves the PDF in Supabase Storage. Keeps the policy list and chat history. Forwards ingest and query calls to Python.
Python RAG APIapi/ + rag_engine/Parses the PDF. Chunks, embeds and stores the text. Searches, re-ranks and calls the LLM. Makes the summary.
Supabasesupabase_migrations/Postgres tables, the pgvector index, the search function, the file bucket and Google OAuth.
Why two backends?

The AI libraries (PyMuPDF, sentence-transformers, tiktoken) are Python tools. The product logic (users, files, history) was easy to build in Express. Node is the “front door” and Python is the “engine room”. The cost is one extra network hop and two codebases to deploy.

3 · Data model

Five tables. Run the SQL files 001 to 005 in order.

users id uuid PK email unique password_hash name created_at policies id uuid PK user_id FK → users policy_id text, unique filename storage_path storage_url (signed, 1 year) status processing | ready | failed chunk_count created_at chat_history id uuid PK user_id FK → users policy_id text question answer sources jsonb source_count created_at policy_chunks id bigserial PK content text metadata jsonb (has policy_id) embedding vector(768) created_at policy_summaries policy_id text PK summary jsonb created_at 1 user : many 1 user : many policy_id, no FK policy_id
One text key joins the data. policy_id links policies, chunks, summaries and history. The database does not enforce these links (dashed). Solid arrows are real foreign keys to users.

How a policy gets its ID

Node creates the ID at upload: policy_id = "<user_id>_<random uuid>". The same string is the file name in Storage (<user_id>/<policy_id>.pdf), the key in policies, and the filter value inside every chunk’s metadata. This is how one policy stays apart from another.

What is stored with each chunk

Each row in policy_chunks has the chunk text, a 768-number embedding, and a metadata JSON object. The JSON comes from ChunkMetadata.to_supabase_dict().

FieldMeaning
policy_idThe policy this chunk belongs to. Used as the search filter.
chunk_id, chunk_indexA random UUID and the order of the chunk in the document.
source_fileName of the uploaded file.
section_name, section_numberThe heading, for example “SECTION 3 - EXCLUSIONS”, and the code “SECTION 3”.
page_numberEstimated page. See chapter 14 for why it is not exact.
clause_typeOne of: coverage, exclusion, definition, deductible, limit, endorsement, schedule, general_condition, unknown.
coverage_categoryfire, flood, theft, liability, medical, vehicle, property, life, travel, or empty.
deductible_related, limit_related, endorsement_flag, table_chunkTrue/false flags set by keyword checks.
token_countSize of the chunk in tokens (tiktoken cl100k_base).

The search function in SQL

CREATE FUNCTION match_policy_chunks(query_embedding vector(768),
                                    match_count int,
                                    filter jsonb DEFAULT '{}')
...
SELECT id, content, metadata,
       1 - (embedding <=> query_embedding) AS similarity
FROM policy_chunks
WHERE metadata @> filter            -- JSON "contains" check
ORDER BY embedding <=> query_embedding   -- cosine distance, smallest first
LIMIT match_count;
  • <=> is the pgvector cosine distance. Similarity = 1 − distance.
  • metadata @> filter keeps only rows whose JSON contains the filter. A filter of {"policy_id":"X"} returns chunks of policy X only.
  • The HNSW index (m = 16, ef_construction = 64) makes the nearest-neighbour search fast. A GIN index on metadata speeds up the filter.
Caution

The file 001_vector_store.sql creates vector(768). Its comments and supabase_migrations/README.md say 1536 or 1024. Trust the SQL: the column is 768 wide. The storage bucket policy-pdfs is not created by any migration. You must create it by hand.

4 · Sign-in

Two ways to log in. Both end with a JWT that Node issues.

Email and password

  1. The browser sends POST /api/auth/signup with email, password and name. Login uses /api/auth/login.
  2. Node hashes the password with bcrypt (cost 10) and inserts a row in users. On login, Node compares the password with the stored hash.
  3. Node signs a JWT with JWT_SECRET. The payload is { id, email }. It expires in 7 days.
  4. The browser saves the token in localStorage["token"]. api.js adds Authorization: Bearer … to every request.

Google sign-in

Browser Supabase Auth Node backend Supabase DB 1. signInWithOAuth('google') 2. redirect back, token in URL #hash 3. App.jsx checks the e-mail domain 4. POST /api/auth/google { email, name } 5. find user, or insert one 6. user row 7. app JWT, valid 7 days 8. save token in localStorage Gap at step 4: Node does not verify the Supabase token. Any caller can post any e-mail and receive a JWT.
Google login is a token exchange. The browser proves identity to Supabase. The app then needs its own JWT, so it posts the e-mail to Node. The sequence is correct in shape, but step 4 has no proof of identity.

How Node checks a token

The file middleware/auth.js tries two checks in this order.

  1. Verify the token with jwt.verify(token, JWT_SECRET). If it works, set req.user and continue.
  2. If it fails, ask Supabase: supabase.auth.getUser(token). If Supabase accepts it, set req.user = { id, email }.
  3. If both fail, return 401.

Other rules in the sign-in code

  • Personal e-mail only. utils/personalEmail.js allows 10 domains (gmail, hotmail, outlook, yahoo, icloud, live, msn, protonmail, me, mac). The browser enforces this. The server does not.
  • Password strength meter and 8-character minimum are in the browser.
  • Forgot-password screen is a stub. A TODO comment says the real flow is not built.
  • OAuth loading gate. App.jsx shows “Completing sign-in…” while it reads the URL hash, so the page does not flash the login screen.

5 · Upload and ingest

From a PDF file to searchable chunks.

Step by step, across the services

BrowserNodeSupabasePythonJina / LLM POST /api/ingest/upload save PDF · add policies row forward PDF (not awaited, 15 s limit) { policy_id, status: processing } background task: parse · clean · chunk embed, 100 chunks per call insert chunks, 50 per call status = ready (100%) loop every 2 seconds until status is ready GET /ingest/status/:id GET /ingest/status/:id { status, progress, message } same JSON
The upload call returns in about one second. Node does not wait for the AI work. The browser polls for progress instead. After the status is ready, a separate thread makes the summary (chapter 6).

Node side (backend/src/routes/ingest.js)

  1. Receive. multer keeps the file in memory. Node rejects a file with no name that ends in .pdf.
  2. Name it. policy_id = <user_id>_<uuid>.
  3. Store it. Upload to the bucket policy-pdfs at <user_id>/<policy_id>.pdf. Make a signed URL valid for 1 year.
  4. Record it. Insert a row in policies with status = 'processing'.
  5. Forward it. Post the file to Python /ingest/upload with axios. Do not await it. If it fails, Node only logs the error.
  6. Answer. Return { status: 'processing', policy_id, … } at once.

Python side (api/routes/ingest.py)

  1. Check that the file name ends in .pdf. Write the file to a temp path.
  2. If the policy already exists and overwrite is false, delete the temp file and return skipped.
  3. Add a FastAPI BackgroundTask. Return processing.
  4. The task runs IngestionService.ingest(). It always deletes the temp file at the end.
  5. If the result is success or skipped, it starts a daemon thread for the summary.

The pipeline inside IngestionService

PDFuploaded filesaved as tempfile Parsepymupdf4llmmarkdown +page markers CleanDocumentCleanerremove noise,fix headings ChunkClauseChunkerabout 31 chunksper 20 pages EmbedJina v3 API768 numbersper chunk StoreSupabasepolicy_chunks50 rows per call SummaryLLM, own threadpolicy_summariesafter “ready” status_tracker progress value 5 · start 10 · old data cleared 15 · parsing 60 · chunks ready 65 · embedding 85 · embedded 90 · storing 100 · ready The summary thread writes 96 and 99 afterwards, but the status endpoint reports “ready” as soon as chunks exist.
The progress bar in the UI is real. status_tracker.update_status() writes these values in IngestionService. The browser reads them through the status endpoint.

Stage 1 · Parse (ingestion/pdf_loader.py)

  • pymupdf4llm.to_markdown(..., page_chunks=True) returns one markdown text per page.
  • The loader adds a marker ---PAGE_START:n--- before each page and joins the pages.
  • This replaced LlamaParse (a cloud parser that took 60 to 120 s). The local parser takes about 0.2 s. Set USE_LLAMAPARSE=true to switch back. The old code is still in the file.
  • A scanned PDF with no text layer gives empty text. The local parser has no OCR.

Stage 2 · Clean (ingestion/cleaner.py)

  1. Page map. Before cleaning, extract_page_map() records which page each text position belongs to.
  2. Remove noise. Delete lines such as “Page 3 of 20”, a lone number, “Confidential”, a web address, a copyright line, the page markers and rows of ___ or ***.
  3. Fix hyphenation. Join words split at a line end (insur-⏎ance → insurance).
  4. Fix spacing. Reduce three or more blank lines to one blank line. Remove trailing spaces.
  5. Mark headings. Put ## before lines that start with SECTION, PART, ARTICLE or CHAPTER plus a number.

Stage 3 · Chunk (chunking/clause_chunker.py)

This is the most special part of the project. A normal splitter cuts text every N characters and can cut a clause in half. This chunker follows the structure of a policy.

Cleaned markdown text Split into sections at SECTION · PART · ARTICLE · CHAPTER · “## ” Table-heavy? 30% of lines have “|” yes TableChunker repeat header row in each chunk at most 1,024 tokens per chunk clause type = schedule no Detect clause type heading first, then first 600 characters Fits the limit? max 384 to 1,024 tokens yes Keep as one chunk no Split at clause markers 3.1 · (a) · (i) — 3 lines of overlap Still too big? Split by tokens window = max tokens, overlap 0 to 64 Attach metadata page · clause type · category · flags · tokens
Three rules drive the chunker: keep a clause whole when it fits, split on clause numbers when it does not, and treat tables on their own so the header row stays with its data.

Chunk size by clause type (tokens, from config/constants.py)

Clause typeTargetMaxOverlap
coverage51276864
exclusion38451264
definition25638432
deductible38451264
limit38451264
endorsement51276864
schedule5121,0240
general_condition / unknown51276864

Definitions are short and exact, so they get small chunks. Schedules are tables, so they get a large limit and no overlap. The “target” value is stored in the config, but the code only reads max and overlap.

How the clause type is found

The code tests keyword lists in a fixed order: exclusion → definition → deductible → schedule → endorsement → limit → coverage → general_condition. It checks the heading first. If nothing matches, it checks the first 600 characters of the body. Exclusion comes first because an exclusion text often contains the word “cover”. A comment in the code warns not to change this order.

What the other metadata flags mean

  • coverage_category: the first category (fire, flood, theft, …) whose keyword appears in the text.
  • deductible_related: the text contains “deductible”, “excess” or “retention”.
  • limit_related: the text contains “limit”, “maximum”, “sum insured” or “aggregate”.
  • Token counts use tiktoken with the cl100k_base encoding.

Stage 4 · Embed (embeddings/jina_embedder.py)

  • The factory get_embedder() reads EMBEDDING_PROVIDER. The default is jina. The option local loads a BGE model with sentence-transformers. The options openai and cohere raise NotImplementedError.
  • Jina v3 gives 1024 numbers by default. The code asks for dimensions: 768. Jina v3 supports shortened vectors (Matryoshka training). This keeps the vector the same size as the database column.
  • The code sends 100 texts per API call and runs up to 5 calls at the same time. The order of results is kept.
  • A retry decorator tries up to 3 times with waits of 1 s, 2 s, 4 s.
  • For a question, embed_query() sends one text.

Stage 5 · Store (vector_store/supabase_store.py)

  • Each row is { content, embedding, metadata }. The code inserts 50 rows per request, one request after another.
  • policy_exists() and get_policy_chunk_count() count rows where metadata->>policy_id equals the ID.
  • With overwrite=true, the service deletes the old chunks first.
  • After the insert, the service sets the status to ready at once. It does not wait for the summary.

How the status endpoint decides

GET /ingest/status/{id} counts chunks in the database. If the count is more than zero, it says ready and 100. If not, it returns the in-memory value from status_tracker. If there is no value, it returns not_found.

Caution

status_tracker is a Python dictionary in memory. It is lost when the server restarts. It is not shared between worker processes. There is no failed state. If ingest crashes, the browser waits for ever.

6 · Auto summary

A structured overview made after the upload.

  1. The ingest task starts a daemon thread that runs SummaryService.generate(policy_id). The dashboard does not wait for it.
  2. The service searches with a fixed query: “policy overview benefits coverage exclusions surrender death benefit premiums”. It takes the top 10 chunks. There is no re-ranking here.
  3. It builds a context of up to 3,000 estimated tokens.
  4. It asks the LLM (max 4,096 tokens) to return only JSON, written in plain English. The prompt sets rules: short sentences, no jargon, explain unavoidable terms in brackets, never copy legal text.
  5. It removes ``` code fences if the LLM added them, then parses the JSON. If parsing fails, the result has parse_error and nothing is saved.
  6. It upserts the result in policy_summaries (key policy_id).

The JSON fields

FieldType
policy_name, policy_type, insurer, uintext (uin may be null)
key_benefits, exclusions, important_conditionslist of full sentences
death_benefit, survival_benefit, surrender_value2–3 sentences each
loan_facility, tax_benefit2–3 sentences, or null
free_look_periodone sentence

How the browser gets it

  • When ingest is ready, the dashboard calls GET /api/policies/:id/summary every 6 s, up to 20 times (2 minutes). A 404 means “not yet”.
  • While it waits, the page shows a skeleton and the text “this takes about 30–40 s”.
  • About 3 to 5 seconds after the summary shows, a small prompt invites the user to open the chat. It hides after 12 s.
Design choice

The first version made the user wait for the summary before the dashboard opened. The team moved it to a separate thread. The user now sees the dashboard as soon as chunks are stored, and chat works even before the summary exists.

Caution

The summary prompt asks for death benefit, survival benefit and surrender value. This suits life insurance. For a motor or health policy, many fields will be empty or null.

7 · Ask a question

From a typed question to a streamed answer.

Questiontyped in thechat box Preprocessadd “?”pick filters EmbedJina queryvector Searchpgvectork = 8 Re-rankcross-encoderkeep 3 Contextabout 3,000tokens at most LLMGroq, streamtemperature 0.3 Answertokens appearin the UI text text + filter{policy_id…} 768 numbers 8 × content,metadata, score 3 × same +rerank_score one text block system + userprompt text/plainstream If the search returns 0 rows, search again with only the policy_id filter.
Wide first, narrow later. The vector search is fast but rough, so it returns 8 candidates. The slower cross-encoder reads question and chunk together and keeps the best 3. The LLM sees only those 3.

Step by step

  1. Preprocess (retrieval/query_preprocessor.py). Trim spaces. If the text starts with a question word (what, is, are, does, can, how …) and has no “?”, add one and capitalise the first letter.
  2. Pick filters. Always policy_id. Then simple keyword rules: “flood / water / storm / hurricane” → coverage_category = flood; fire words → fire; theft words → theft; “deductible / excess” → deductible_related = true; “limit / maximum / cap” → limit_related = true.
  3. Embed the cleaned question into one vector (768 numbers).
  4. Search. Call the SQL function match_policy_chunks with the vector, k = 8 and the filter. If it returns nothing, repeat with only {policy_id}.
  5. Re-rank (retrieval/reranker.py). A cross-encoder (cross-encoder/ms-marco-TinyBERT-L-2-v2) scores each pair [question, chunk]. Sort by score. Keep 3.
  6. Build context (retrieval/context_builder.py). Write each chunk as a block with “Source n”, section name, clause type and score. Estimate tokens as words × 1.3. Stop before the total passes 3,000. Always keep at least one block.
  7. Build the prompt. A system prompt plus a user prompt (below).
  8. Call the LLM and stream the tokens to the client.

The system prompt (summary of its 7 rules)

  • Persona: “PolicyDecoder”, a friendly insurance helper. (The UI calls it IRIS.)
  • Use simple, plain English. Short paragraphs.
  • Answer only from the given context. Never guess.
  • If the answer is not there, say: “I couldn’t find this in your policy document. You may want to contact your insurer directly.”
  • Name the section briefly, for example “As per Section 3.1…”. One reference is enough.
  • For yes/no questions, start with Yes or No, then explain.
  • Aim for 3 to 5 sentences. Never speculate.

The user prompt template

INSURANCE POLICY CONTEXT:
--- Source 1 ---
Section: SECTION 3 - EXCLUSIONS
Clause Type: exclusion
Relevance: 0.912
<chunk text>
...
---

QUESTION: Is flood damage covered?

POLICY ID: <policy_id>

Please answer the question based strictly on the policy context
provided above. Cite specific sections where applicable.

How streaming works end to end

LLM APIsends smallpieces of text Pythonstream_query()yields each token FastAPIStreamingResponsetext/plain Noderes.write(chunk)and keeps a copy Browserfetch reader +TextDecoder React statetext = text + tokenbubble grows stream ends:insert chat_history
Nothing waits for the full answer. Each layer passes tokens on as soon as it gets them. Node also keeps a copy so it can save the whole answer in chat_history at the end.

What the response contains

  • POST /query/stream (used by the UI) returns plain text only. It does not return sources.
  • POST /query/ (not used by the UI) returns { answer, policy_id, source_count, sources[] }. Each source has section, clause_type, page_number, highlight_text (the full chunk), relevance_score and a 200-character snippet. The design intends the UI to highlight the passage in a PDF viewer. The UI does not do this yet.
  • Each question is independent. The prompt holds no earlier messages, so follow-up questions such as “and what about the limit?” lose their context.
Which numbers are real?

The API default is k = 8. The service keeps top_k_rerank = 3. The README says “re-rank to 5” and “about 3,400 tokens”. The code is the truth: 8 in, 3 out, 3,000 tokens at most.

8 · Frontend

React 19, Vite 7, Tailwind 4, lucide icons, Supabase JS.

home login dboard upload chatbot Get started log in sign-up ok ready / cancel Upload button Open IRIS Back Log out
The router is a state variable. App.jsx keeps appState and mirrors it in the URL hash (#login, #dboard …) with history.pushState. It does not use react-router, even though the package is installed.

Main files

FileJob
App.jsxScreen switch. Dark-mode flag. Listens to Supabase onAuthStateChange, blocks non-personal e-mail, swaps a Google session for a Node JWT.
api.jsfetchApi() adds the Bearer token, sets JSON headers (not for FormData), parses JSON or text, and throws on errors. streamApi() reads a response body chunk by chunk and calls onToken. The base URL is VITE_API_URL, else http://localhost:4000/api on localhost, else /api.
Hero, Navbar, Footer, AboutLanding page.
Auth/Login.jsxLogin, sign-up and forgot-password views. Password strength bar. Google button.
Dashboard/UploadModal.jsxDrag-and-drop. Accepts PDF only, 50 MB maximum. Shows four stages and a progress bar. Polls status every 2 s.
Dashboard/Dboard.jsxMain screen (about 1,000 lines). Sidebar of past policies. Summary cards. A floating, resizable chat widget with one message list per policy.
Dashboard/Chatbot.jsxFull-page chat route. Can export the chat as a PDF with jsPDF (loaded from a CDN). Without a file prop it queries the policy ID "test", so it works only as a demo.

Timers in the browser

WhatEveryStops when
Ingest status poll2 sstatus is ready (no time limit)
Summary poll6 ssummary found, or 20 tries
Chat prompt bubble3–5 s delay, then 12 s visiblechat opens

Details worth knowing

  • React StrictMode is off. A comment in main.jsx says it ran the upload effect twice and sent two uploads. UploadModal also has three guard fixes (a completedRef, a ref for the callback, and a narrow dependency list). The cleaner fix is an effect that is safe to run twice.
  • Theme. Colours are two JavaScript objects, LIGHT and DARK, used as inline styles. The landing page uses Tailwind dark: classes.
  • Chat state lives in a map { [policy_id]: messages[] } in memory. A page reload clears it. The history API exists, but the dashboard does not read it.
  • History list comes from GET /api/policies. It shows the file name, the date and, once loaded, the policy name from the summary.
  • The file picker accepts .pdf,.doc,.docx, but the server accepts PDF only.

9 · API tables

Every route in the code. The README lists some paths wrongly.

Node backend (port 4000). All paths start with /api.

MethodPathAuthWhat it does
GET/healthnoneLiveness text.
POST/auth/signupnoneCreate a user. Return token + user.
POST/auth/loginnoneCheck password. Return token + user.
POST/auth/googlenoneFind or create a user by e-mail. Return token. See chapter 14.
POST/ingest/uploadJWTMultipart field file. Store PDF, add policy row, forward to Python.
GET/ingest/status/:policy_idJWTAsk Python for progress. When ready, update the policy row.
GET/policiesJWTList the user’s policies, newest first.
DELETE/policies/:policy_idJWTDelete the file and the policy row only.
GET/policies/:policy_id/summarynoneRead the stored summary. 404 if not ready.
POST/queryJWTProxy to Python. Save question, answer and sources.
POST/query/streamJWTProxy the stream. Save the answer at the end.
GET/historyJWTQuery parameters: policy_id, q (search text), page, limit.
GET / DELETE/history/:idJWTOne entry by row ID.
DELETE/history?policy_id=…JWTClear all entries, or those of one policy.

Python API (port 7860 on HF, 8000 locally)

MethodPathWhat it does
POST/ingest/uploadMultipart file, policy_id, overwrite. Start background ingest.
POST/ingest/JSON with pdf_url. This route passes the URL to the loader as a file path, so it fails. Treat it as unfinished.
GET/ingest/status/{id}Chunk count, status, progress, message.
POST/ingest/summary/{id}Make or remake the summary now (slow, waits for the LLM).
GET/ingest/summary/{id}Read the stored summary.
POST/query/{ question, policy_id, k = 8 } → answer + sources.
POST/query/streamSame input. Returns text/plain tokens.
GET/health/Checks Supabase. Returns ok or degraded.

10 · Code patterns

The design ideas you can name in an interview.

PatternWhereWhy it helps
Abstract base + factoryBaseEmbedder / get_embedder(), BaseLLM / get_llm(), BaseVectorStore / get_vector_store()Change provider with an environment variable. Add a new provider without touching the pipeline.
Lazy importsInside factory functions and service constructorsHeavy libraries (torch, Gemini SDK) load only when used. Server starts faster.
Load once at start-uplifespan() in api/main.py builds one QueryService in app.stateThe cross-encoder model loads once, not for every request.
Retry with backoffutils/retry.py @with_retryHandles short API failures. LLM: 4 tries, 5 s start, ×2. Jina and Supabase: 3 tries, 1 s start, ×2.
Singleton status storeStatusTrackerOne shared dictionary for progress. Simple, but not durable.
Fire and forgetNode → Python upload call; Python summary threadKeeps the request fast. The caller polls for the result.
Typed configpydantic-settings Settings with @lru_cacheAll keys come from .env. One cached object.
Typed metadataChunkMetadata (Pydantic) with to_supabase_dict()One schema for every chunk. Converts enums, dates and null to JSON-safe values.
Proxy / BFFNode routes query.js, ingest.jsThe browser sees one API. Keys and the Python URL stay on the server.
Provider notes

get_llm() supports gemini, groq, kimi and openai (the last one re-uses the Kimi class with an OpenAI client). The setting LLM_MODEL is only printed in a log. Each provider reads its own model setting (GROQ_MODEL, GEMINI_MODEL, KIMI_MODEL). The files local_llm.py, openai_llm.py, openai_embedder.py, query_request.py and query_response.py are empty.

11 · Config and deploy

Where each part runs and which settings it needs.

PartRuns onHow it gets there
Python RAG APIHugging Face Space devjhawar/policylens-rag-api (Docker, port 7860)GitHub Action sync-to-hf.yml: on each push to main, huggingface_hub.upload_folder sends the repo (not .git, .github, .env, *.zip). The README header (sdk: docker, app_port: 7860) is the Space config.
Node backendIts own Docker image (backend/Dockerfile, Node 18 Alpine, non-root user, port 7860)Prepared for Hugging Face Spaces. npm ci --only=production, then npm start.
FrontendVercel (the examples use policylens-ai.vercel.app)VITE_API_URL points to the Node API. VITE_BASE_PATH sets the Vite base path.
Database / files / Google OAuthSupabase projectRun SQL 001–005. Create the policy-pdfs bucket. Enable the Google provider.

Python .env

VariableDefault / note
SUPABASE_URL, SUPABASE_SERVICE_KEYRequired. Use the service-role key.
LLAMA_CLOUD_API_KEYRequired by the settings class even when not used. Set any text.
LLM_PROVIDERgroq (also gemini, kimi, openai)
GROQ_API_KEY, GROQ_MODELModel default openai/gpt-oss-20b
GEMINI_API_KEY, MOONSHOT_API_KEY, OPENAI_API_KEYOnly for those providers.
EMBEDDING_PROVIDER, JINA_API_KEYjina (or local). Jina key required for Jina.
USE_LLAMAPARSEfalse. Set true for the cloud parser.
CORS_ORIGINS* by default. Set real origins in production.
DEBUGfalse

Node backend/.env

PORT (4000), SUPABASE_URL, SUPABASE_SERVICE_KEY, JWT_SECRET, PYTHON_API_URL, CORS_ORIGINS.

Frontend .env

VITE_API_URL, VITE_SUPABASE_URL, VITE_SUPABASE_ANON_KEY.

Run it on your machine

  1. Run SQL files 001 to 005 in the Supabase SQL editor. Create the bucket policy-pdfs.
  2. Python: pip install -r requirements.txt, fill .env, then python -m uvicorn api.main:app --port 8000 --reload.
  3. Node: cd backend && npm install, fill backend/.env, then npm run dev.
  4. Frontend: cd frontend && npm install && npm run dev (port 5173).
Caution

The Dockerfile caches the model ms-marco-MiniLM-L-6-v2 at build time, but the code loads ms-marco-TinyBERT-L-2-v2. The server therefore downloads the real model at start-up. The HEALTHCHECK calls /, which has no route, and docker-compose.yml maps port 8000 while the image listens on 7860. Fix these three before relying on them.

12 · Tests

What exists and what is missing.

FileWhat it checks
test_chunker.pyChunks are made, none are empty, policy IDs and indexes are right, exclusion and coverage types are found, table chunks are flagged, token counts stay in limits, page map is used.
test_retrieval.pyQuery preprocessing and filters. The retriever calls embed and search, and falls back on empty results. Context builder format.
test_vector_store.pyInsert, search, delete and exists, with a mocked Supabase client.
test_integration.pyFull query flow and ingestion skip/overwrite with mocks. Response format, highlight fields, page map.
test_evaluation.pySmall checks on a sample policy: chunk quality, query cleaning, context token limit, at most 5 sources.
test_e2e.pyA live script (not a pytest file). Needs real keys. Ingests a sample, asks 3 questions, deletes it.

Run them with python -m pytest rag_engine/tests/ -v. I read these files but did not run them.

Gaps

There are no tests for the Node backend or the React app. There is no measure of answer quality (recall, faithfulness). The “evaluation” file checks plumbing, not retrieval accuracy.

13 · README vs code

The code is the truth. Know these differences before someone else finds them.

TopicREADME saysCode does
LLMKimi (kimi-k2.5)Default provider is Groq, openai/gpt-oss-20b. Kimi is legacy.
Re-rankerBGE rerankercross-encoder/ms-marco-TinyBERT-L-2-v2
Re-rank sizetop 5, context about 3,400 tokenstop 3, context up to 3,000 estimated tokens
Summary inputtop 15 chunkstop 10 chunks, 3,000-token context
Chunker filechunker.pyclause_chunker.py and table_chunker.py
Auth paths/auth/signup, /auth/login/api/auth/…, plus /api/auth/google
History APIPOST /api/history, DELETE /api/history/:policy_idNo POST. The query routes save history. Delete is by row ID, or by ?policy_id=.
Python version3.11pyproject.toml says ≥ 3.12. The Dockerfile uses 3.11.
Health text—Reports model “BAAI/bge-base-en-v1.5 + kimi-k2.5”. Both are out of date.
Embedding size768 (Jina)True, by shortening Jina v3. SQL comments and the migrations README still say 1536 / 1024.
Unused settings—MIN_CONFIDENCE_THRESHOLD, RERANKER_MIN_SCORE, MMR_LAMBDA_MULT, DEFAULT_TOP_K are defined and never read.
Extra files—api.zip (old copy of api/) and a root node_modules/ are committed.

14 · Risks and bugs

Found by reading the code. Rated by how much harm they can do.

Warning · fix before real users

Items 1 to 4 can expose user data. If you present this project, say you know about them and say how you would fix them. That shows good judgement.

#LevelProblemFix
1highPOST /api/auth/google trusts the e-mail in the request body. Anyone can send any e-mail and get a 7-day JWT for that account.Send the Supabase access token. Node calls supabase.auth.getUser(token) and uses the e-mail from the result.
2highGET /api/policies/:id/summary has no login check. A comment says “bypass user check for debugging”.Add authMiddleware and check that the policy belongs to req.user.id.
3highNo owner check on query and status routes. Any logged-in user who knows a policy_id can ask questions about it. The Python API has no login at all and is public on Hugging Face.Check ownership in Node. Add a shared secret header between Node and Python. Limit CORS.
4highfrontend/.env.example has a value that looks like a Google OAuth client secret (it starts with GOCSPX-) in the anon-key line. The Supabase project URL is also hard-coded in supabaseClient.js.Treat the secret as leaked and rotate it. Use a placeholder in the example file.
5mediumNo failed state. If ingest crashes, or Node’s forward call fails, the row stays “processing” and the browser polls with no time limit.Catch errors, set status failed, return it, and add a poll timeout.
6mediumProgress is kept in process memory only.Store status in a database table or Redis.
7mediumThe store retries the whole add_chunks call. A failure at batch 3 repeats batches 1 and 2 and creates duplicates. A part-way failure also leaves chunks behind, so the status endpoint says “ready” for a half-stored policy.Retry per batch. Delete the policy’s chunks on failure. Or insert in one transaction.
8mediumDeleting a policy removes only the file and the policies row. Chunks, summary and chat history stay in the database.Also delete rows in policy_chunks, policy_summaries, chat_history.
9mediumKeyword filters can hide good chunks. A chunk has only one coverage_category (the first match), and the filter needs an exact match. The fallback runs only when there are zero rows.Use filters only to boost, or drop them. Test retrieval with and without filters.
10mediumChat has no memory. Sources never reach the UI because the stream route returns text only.Rewrite follow-ups with earlier messages. Send sources as a final JSON line or use SSE.
11mediumJWT is stored in localStorage (XSS can read it). The personal-e-mail rule and password rules exist only in the browser. No rate limit. The history search text goes into a PostgREST .or() string.Use httpOnly cookies or short tokens with refresh. Repeat checks on the server. Add express-rate-limit. Escape the search text.
12lowPage numbers are approximate. Each chunk gets the page found for the previous section, and the page map uses positions from before cleaning.Compute the page from the start of the current section on the raw text.
13lowrag_engine/main.py has from __future__ import annotations after other code. Python rejects this, so the CLI cannot start.Move the import to the top.
14lowJina v3 supports a task setting (query vs passage). The code does not set it.Pass retrieval.passage for chunks and retrieval.query for questions, then compare results.
15lowThe summary prompt suits life insurance. The coverage categories suit property insurance. Health and motor policies fit neither well.Detect the policy type first, then choose a prompt.
16lowScanned PDFs return no text. The .pdf check is case-sensitive. The upload form accepts Word files. POST /ingest/ (URL mode) fails.Add OCR or the LlamaParse fallback. Use lower-case checks. Remove the dead options.

15 · Interview Q&A

Say the answers in your own words. Answers use “I” and “we”. Change them to match the part you personally built.

Explain PolicyLens in 30 seconds.basics

PolicyLens lets a user upload an insurance policy PDF and ask questions about it in plain English. The system splits the PDF into clause-sized chunks and stores them as vectors in Postgres with pgvector. For each question it finds the best chunks, re-ranks them, and gives them to an LLM that answers only from that text. The answer streams to the screen. It also makes a structured summary on upload.

Why RAG? Why not give the full PDF to the LLM, or fine-tune a model?basics · rag

A policy can be 50 pages. Sending all of it on every question costs many tokens, adds delay, and the model can miss details in the middle. Fine-tuning changes how a model writes, not what it knows, and we cannot train a model for every upload.

With RAG we send only the 3 most relevant chunks. It is cheaper and faster. We can show which section the answer came from. Each user’s data stays in the database and is filtered by policy ID.

Why do you have both a Node and a Python backend?basics · backend

The AI work needs Python libraries: PyMuPDF for PDFs, sentence-transformers for the re-ranker, tiktoken for token counts. The product work (users, file storage, history) was quick to build in Express. Node is the single public API for the browser. Python is an internal engine that Node calls.

The trade-off is one more network hop, two deployments, and two places where auth can go wrong. A smaller team could put everything in FastAPI. I would consider that if I started again.

Walk me through what happens when I upload a PDF.basics
  • The browser posts the file to Node. Node checks the JWT and makes a policy ID.
  • Node saves the file in Supabase Storage and adds a row to policies.
  • Node forwards the file to Python without waiting, and answers the browser at once.
  • Python runs a background task: parse with pymupdf4llm, clean, chunk by clause, embed with Jina in batches of 100, insert into pgvector in batches of 50.
  • The browser polls the status every 2 seconds and shows a progress bar. When chunks exist, the status is ready.
  • A separate thread then makes the summary with the LLM. The dashboard polls for it every 6 seconds.
Walk me through what happens when I ask a question.basics
  • The browser posts { question, policy_id } to Node. Node adds k = 8 and forwards it to Python.
  • Python cleans the question, picks metadata filters, and embeds it with Jina.
  • pgvector returns the 8 nearest chunks for that policy (cosine distance).
  • A cross-encoder scores each chunk against the question. We keep the top 3.
  • We build a context block, add the system prompt, and call the LLM with streaming.
  • Tokens flow back through FastAPI and Node to the browser. Node saves the full answer in chat_history when the stream ends.
How does your chunking work, and why is it “clause-aware”?rag

A fixed-size splitter can cut a clause in the middle, and a half clause can reverse its meaning (“…does not cover” without the rest). Insurance text has structure: sections, numbered clauses, tables. My chunker splits at SECTION, PART, ARTICLE, CHAPTER and “##” headings first. A section that fits the size limit stays whole. A larger one is split at clause markers such as 3.1 or (a), with a 3-line overlap. Only if a piece is still too large do we split by tokens.

Each chunk also gets metadata: section name, clause type, coverage category, flags for deductible and limit, page and token count.

Why different chunk sizes for different clause types?rag

Different clauses have different shapes. A definition is short and exact, so a 384-token maximum keeps it focused. Coverage and endorsement clauses are longer and need more room (768). A schedule is a table, so it gets 1,024 tokens and no overlap. Smaller chunks give sharper matches. Larger chunks keep more context. The clause type lets us choose the right balance for each part.

How do you handle tables?rag

If 30% or more of the lines in a section hold a “|”, the section goes to the table chunker. It finds the header row and the separator row, groups the data rows into batches up to 1,024 tokens, and repeats the header in every batch. This way a row like “Fire | $500” never loses its column names. These chunks are marked as type schedule with table_chunk = true.

What is an embedding? Why 768 dimensions?rag

An embedding is a list of numbers that represents the meaning of a text. Texts with similar meaning have vectors that point in similar directions, so we can search by meaning instead of by exact words. “Is fire covered?” can match a chunk that says “loss caused by fire, smoke or explosion”.

Jina v3 gives 1024 numbers by default. It supports shortened vectors, so we ask for 768. This matches the database column and uses less storage. Smaller vectors lose a small amount of accuracy but search faster. If I changed the size, I would need to change the column and re-embed all chunks.

What does pgvector do? What are cosine distance and HNSW?rag

pgvector adds a vector column type and distance operators to Postgres. Cosine distance (<=>) measures the angle between two vectors. A small angle means similar meaning. The SQL function returns 1 − distance as the similarity.

An exact search compares the query with every row. HNSW is an index that links each vector to near neighbours in layers. A search walks the graph and finds close matches much faster, with a small chance of missing the true best one. I used m = 16 and ef_construction = 64, which are common defaults.

Why a re-ranker? Why retrieve 8 and keep 3?rag

Vector search uses a bi-encoder: the question and each chunk are turned into vectors separately. This is fast but rough. A cross-encoder reads the question and a chunk together and gives one relevance score. It is more accurate but slower, so we run it only on the 8 candidates.

We keep 3 because the context must be short. Less text means lower cost, lower delay and less noise for the LLM. The numbers 8 and 3 were reduced from 8 and 5 to cut response time. I would tune them with a test set.

How do you reduce hallucination?rag
  • The system prompt says: answer only from the context, never guess, and use a fixed sentence when the answer is missing.
  • Temperature is low (0.3).
  • The context holds only the 3 best chunks, so there is less unrelated text.
  • The non-stream API returns the source chunks so a person can check them.

The limit: nothing checks the answer after generation. A next step is a faithfulness check, where a second model confirms that each claim appears in the sources.

How do you keep one user’s policy separate from another’s?rag sec

Every chunk stores policy_id in its metadata. Every search passes a filter { policy_id }, and the SQL uses metadata @> filter. The policy ID starts with the user ID and a random UUID, so it is hard to guess.

I should be honest about the gap: the server does not yet check that the logged-in user owns that policy ID. The fix is a lookup in the policies table before every query. I would also move to row-level security in Supabase.

How do you measure answer quality?rag

Today the tests check the pipeline: chunk sizes, filters, formats. They do not measure quality. To measure it I would build a golden set of 30 to 50 questions with the correct answer and the correct source chunk for several real policies. Then I would track recall@k (is the right chunk in the top 8?), MRR, and answer faithfulness (using a judge model or manual review). I would run it when I change the chunker, the embedding model or the prompts.

What happens with a scanned PDF?rag

pymupdf4llm reads the text layer. A scan has none, so the text is empty and nothing useful is stored. The project keeps the old LlamaParse path behind USE_LLAMAPARSE=true, which does OCR in the cloud but takes 60 to 120 seconds. A better design detects an empty text layer and runs OCR only for those files.

Your chat does not remember earlier messages. How would you add memory?rag

Right now each question is searched alone, so “and what is the limit?” has no subject. I would send the last few messages with the request. Then I would do query rewriting: ask a small model to turn the follow-up into a stand-alone question (“What is the fire coverage limit?”) before embedding. The rewritten question goes to search. The recent messages also go in the final prompt.

What is a weakness of your metadata filters?rag

The filters are keyword rules. A question with “flood” adds coverage_category = flood. But a chunk has only one category, the first match, so a chunk about fire and flood may be tagged “fire” and be excluded. The fallback runs only if the search returns zero rows. A safer design uses the filter as a boost, or runs a search with and without it and merges the results.

What if a PDF contains text that tries to give the LLM orders (prompt injection)?rag · sec

The policy text is untrusted input. The model has no tools, so it cannot take actions, which limits the damage. The worst case is a misleading answer. I would wrap the context in clear delimiters, tell the model that text inside them is data and not instructions, and scan the output for odd content. For a bigger system I would also log sources for audit.

How does the summary feature work?rag

After ingest, a thread runs a fixed search query (“policy overview benefits coverage exclusions surrender death benefit premiums”), takes the 10 best chunks, and asks the LLM for a JSON object with fields such as key benefits, exclusions, death benefit and free-look period, in plain English. The code strips code fences, parses the JSON and saves it in policy_summaries. If parsing fails, it saves nothing. The prompt suits life insurance best.

Why background tasks and polling? Why not wait for the result?backend

Ingest takes seconds to minutes: parse, many embedding calls, many inserts. A normal HTTP request would risk a timeout, and the user would see nothing. With a background task the upload returns fast and the UI shows real progress. Polling is simple and works on any host.

The alternatives are server-sent events or WebSockets (push instead of poll) and a job queue such as Celery, RQ or Cloud Tasks. A queue also survives restarts and allows retries, which the current in-process tasks do not.

What design patterns did you use?backend

Abstract base class plus factory for the LLM, embedder and vector store, so I can switch provider with an environment variable. Lazy imports for heavy libraries. A start-up hook that loads the query service once. A retry decorator with exponential backoff. A singleton status tracker. A backend-for-frontend proxy in Node. Pydantic for settings and chunk metadata.

How does retry with backoff work in your code?backend

The @with_retry decorator calls the function, and if it raises, waits and tries again. The wait grows each time: 1 s, 2 s, 4 s for Jina and Supabase (3 tries); 5 s, 10 s, 20 s for LLM calls (4 tries). This helps with short rate limits and network errors. One known problem: on the Supabase insert, a retry repeats batches that already succeeded. I would retry per batch instead.

Why is the status tracker in memory? What breaks?backend

It was the simplest thing that worked on one server. It breaks on restart (progress is lost), with several workers (each has its own dictionary), and it has no failed state. The fix is a status column or table in the database, or Redis. The chunk count already serves as a durable “ready” signal, which is why the endpoint checks it first.

Why Supabase pgvector and not Pinecone or Chroma?backend

One database holds users, policies, chat history and vectors. I can filter vectors with normal JSON and SQL conditions, join them with other data, and use one free-tier service. There are fewer moving parts. A dedicated vector database can scale to many millions of vectors with more tuning options. For this size, Postgres is enough.

Why Jina API instead of a local embedding model?backend · rag

The Python service runs on a free Hugging Face Space with a small CPU. A local model needs torch and a large download, and it embeds slowly on CPU. The Jina API is fast, has a free tier (about 1 million tokens per month) and gives strong quality. The cost is a network call, a third party that sees the text, and a limit on usage. The code keeps a local BGE option for offline use.

Why can you swap the LLM so easily?backend

All LLM classes follow the BaseLLM interface: complete, stream and the same with message lists. Groq and Kimi use the OpenAI-compatible API, so they share most of the code. The factory picks the class from LLM_PROVIDER. The query service never knows which provider it uses. Gemini needs a small adapter to convert message formats.

How do you show streaming text in React?frontend

I use fetch and read response.body.getReader() in a loop. Each chunk is decoded with TextDecoder and passed to a callback. The callback updates state: text = text + token for the AI message with a known ID. React re-renders and the bubble grows. I used a plain text stream, not SSE, to keep it simple.

Why a hash-based state router and not react-router?frontend

The app has only five screens and no nested routes. A state variable plus history.pushState with a hash works on any static host without server rewrites. It also let us handle the OAuth redirect, which returns tokens in the URL hash. The cost is manual work for deep links and no route-level code splitting. react-router is installed but unused. If the app grows, I would switch to it.

Why did you remove React StrictMode? Was that the right fix?frontend

In development StrictMode runs effects twice. The upload effect then sent two uploads. I removed StrictMode and added guards (a ref flag, a ref for the callback, a narrow dependency list). It works, but the better fix is an effect that is safe to run twice: for example, start the upload from a button click handler instead of an effect, and cancel with an AbortController on cleanup.

How do the upload progress bar and the summary loading work?frontend

After the upload call returns a policy ID, a timer calls the status endpoint every 2 seconds and sets the bar to the progress number from the server. When status is ready, the app moves to the dashboard. Then another timer asks for the summary every 6 seconds, up to 20 times, while a skeleton shows. The chat button is already usable at that point.

How does authentication work?security

Email and password: bcrypt hashes the password, and Node returns a JWT signed with a secret, valid for 7 days. Google: the browser signs in through Supabase, then asks Node for its own JWT. Every protected route runs a middleware that verifies the JWT, and falls back to checking a Supabase token. The browser stores the token in localStorage and sends it as a Bearer header.

What are the biggest security problems in the project?security
  1. The Google login route issues a JWT for any e-mail without checking the Supabase token. Fix: verify the token with Supabase on the server.
  2. The summary route has no login check. Fix: add the middleware and an owner check.
  3. No ownership check on query and status, and the Python API is public. Fix: check owner in Node, add a shared secret, restrict CORS.
  4. An example env file holds what looks like a Google client secret. Fix: rotate it and remove it from the repo history.

I would also add rate limits, server-side validation, and delete all data for a policy when the user deletes it.

JWT in localStorage or a cookie?security

localStorage is easy, but any script on the page (an XSS bug) can read the token. An httpOnly cookie cannot be read by scripts, but then you need CSRF protection (SameSite and a CSRF token). The most robust choice is a short-lived access token in memory plus a refresh token in an httpOnly cookie. We used localStorage for speed of development, and I know the risk.

Insurance documents are private. How do you handle privacy?security · ops

Today: files are in a private Supabase bucket and shared by signed URL (valid for 1 year). Chunk text goes to Jina, and question plus chunks go to the LLM provider. So data leaves our system. I would shorten URL life, tell users clearly, offer delete-everything, remove personal data before sending where possible, and choose providers with no-training and no-retention terms. For strict customers, I would use the local embedding model and a self-hosted LLM.

Where is the bottleneck? How would you scale it?scale

PDF parsing is fast (0.2 s). The slow parts are external: the embedding calls during ingest, and the LLM calls (the summary takes 30 to 40 s). For many users I would: move ingest to a job queue with separate workers; run several API instances behind a load balancer; keep status in a shared store; cache embeddings of common questions; add connection pooling for Postgres; and add rate limits per user. The re-ranker runs on CPU, so I would watch its latency and could move it to a GPU or a hosted API.

What does it cost to run?ops

The prototype uses free tiers: Hugging Face Space, Vercel, Supabase, Jina (about 1M tokens per month) and Groq. One upload costs a few embedding calls and one LLM call for the summary. One question costs one embedding call and one LLM call. The main cost risk is the LLM at scale. Short context (3 chunks) keeps tokens low.

How do you deploy it?ops

The Python API is a Docker image on a Hugging Face Space. A GitHub Action uploads the repo to the Space on every push to main. The Node backend has its own Dockerfile. The React app deploys on Vercel with an environment variable for the API URL. Secrets are environment variables, never committed. Known gaps: the Dockerfile caches a different re-ranker than the code loads, the health check points to a missing route, and there is no automatic test step in the pipeline.

Tell me about a performance improvement you made.story

PDF parsing. The first version used LlamaParse, a cloud service that took 60 to 120 seconds per file. I replaced it with pymupdf4llm, which runs locally in about 0.2 seconds, and made the output format identical (same page markers) so the chunker did not change. I kept the old path behind a flag as a rollback. Other changes: summary moved to its own thread so the dashboard opens sooner; embeddings sent in parallel batches; retrieval reduced from 8-and-5 to 8-and-3 to cut latency.

What was a hard bug you fixed?story

Duplicate uploads. In development, React ran the upload effect twice and the same file was sent twice. I found it by watching the network tab and the double log lines. I fixed it with a ref guard and by removing a changing callback from the dependency list, and I left comments in the code explaining each fix. I now prefer to start uploads from a click handler, not an effect.

What would you do next if you had two more weeks?story
  1. Fix the security items: verified Google login, owner checks, a Node-to-Python secret, secret rotation.
  2. Add a failed state, a durable status store and a job queue.
  3. Build an evaluation set and track recall@k and faithfulness.
  4. Add conversation memory with query rewriting.
  5. Show sources in the chat and jump to the page in the PDF (the API already returns page and highlight text).
  6. Support more policy types with a detected type and matching prompts.
What would you do differently if you started again?story

I would write the evaluation set first, so every choice (chunk size, k, re-ranker) has a number behind it. I would use one backend language, or define the contract between Node and Python early. I would design auth and ownership checks at the start instead of adding them later. And I would keep the README in step with the code, because it already differs in several places.

16 · Numbers to know

Short facts for fast recall.

ItemValue
PDF parse timeabout 0.2 s (LlamaParse was 60–120 s)
Chunks per 20 pagesabout 31
Chunk max size384 to 1,024 tokens, by clause type
Table-heavy rule30% of lines contain “|”
Embedding size768 numbers (Jina v3 shortened from 1024)
Embedding batch / threads100 texts per call, 5 calls at once
Insert batch50 rows per call
HNSW settingsm = 16, ef_construction = 64, cosine
Search / keep8 candidates → 3 chunks
Context limit3,000 estimated tokens (words × 1.3)
LLM settingstemperature 0.3, max 4,096 output tokens
Summary input10 chunks, 3,000-token context
Summary timeabout 30–40 s
Ingest poll / summary poll2 s / 6 s (20 tries)
JWT life · bcrypt cost7 days · 10
Signed file URL life1 year
Upload limit (browser only)50 MB, PDF
Node → Python upload timeout15 s
Retry (Jina, Supabase / LLM)3 tries from 1 s / 4 tries from 5 s, ×2 each time

17 · Glossary

Terms in simple words.

RAG
Retrieval-augmented generation. Find text first, then let an LLM write an answer from that text.
LLM
Large language model. A program that writes text, for example gpt-oss, Gemini or Kimi.
Chunk
A small piece of the document, such as one clause, that is stored and searched alone.
Token
A small piece of text (part of a word) that models count. About 1 token is ¾ of an English word.
Embedding
A list of numbers that stands for the meaning of a text.
Vector search
Finding the stored embeddings that are closest to the question embedding.
Cosine similarity
A score for how much two vectors point the same way. 1 means the same direction.
pgvector
A Postgres add-on that stores vectors and searches them.
HNSW
An index that makes nearest-neighbour search fast by walking a graph of close points.
Bi-encoder
Turns the question and the chunk into vectors separately. Fast. Used for the first search.
Cross-encoder
Reads the question and the chunk together and scores them. Slower, more exact. Used to re-rank.
Re-ranking
Sorting the first results again with a better model.
Matryoshka embedding
A vector trained so its first N numbers still work alone. This lets Jina v3 give 768 instead of 1024.
Metadata filter
A rule that limits a search to rows with given values, such as one policy_id.
Streaming
Sending the answer in small parts while the model writes it.
Polling
The client asks the server again and again until a job is done.
Background task
Work that continues after the HTTP answer is sent.
JWT
JSON Web Token. A signed string that proves who the user is.
bcrypt
A slow password-hashing method. It makes stolen hashes hard to crack.
Deductible
The part of a claim the policyholder pays first.
Exclusion
A thing the policy does not cover.
Free-look period
Days after purchase when the user can cancel and get a refund.
Surrender value
Money the user gets if they cancel a life policy early.
IRIS
The name of the chat assistant in the UI. (The system prompt calls it “PolicyDecoder”.)