// the find
supabase-community/chatgpt-your-files
Production-ready MVP for securely chatting with your documents using pgvector
A step-by-step Supabase workshop that builds a document Q&A app: upload a markdown file, split it by heading, embed each section with gte-small in an Edge Function, then chat over the matches with OpenAI. It suits developers who want to see pgvector and row-level security wired together end to end, not anyone who needs something to deploy as-is.
- RLS covers every layer. Storage object policies, documents, and document_sections all key off auth.uid(), and the Edge Functions forward the caller's Authorization header so those policies apply instead of being bypassed with a service-role key.
- Embedding runs as a trigger chain. An after-insert statement trigger batches the new section IDs and fires pg_net requests at the embed function, so the upload path does not wait on model inference.
- Query and document embeddings come from the same gte-small model, one in the browser via Transformers.js and one in the Edge Function. The vectors are normalized and indexed with inner product, and the README explains why that gives the same ranking as cosine distance, which most RAG demos skip.
- The workshop is split into git tags per feature, so you can jump to any stage and diff just that change.
- The description says production-ready, but the README admits that if the embed function fails, sections are left with null embeddings. Retry is listed as something to build later, so those sections stay unsearchable until someone notices.
- Chunking is by markdown heading only. A long section becomes one row with one vector, so retrieval gets coarse on big sections and on documents without headings. The parser is built for markdown and nothing else.
- The model and dimension are hardcoded in several places: the vector(384) column, the match function signature, and the model name on both client and server. Changing embedding models means a migration and re-embedding every row, and the README does not say so.
- There is no upload size limit, rate limit, or per-user quota on the chat endpoint, and the OpenAI key sits in a local .env file. That is fine for a workshop, but anyone who can sign up can spend your OpenAI budget.