// the find
Azure-Samples/serverless-chat-langchainjs
Build your own serverless AI Chat with Retrieval-Augmented-Generation using LangChain.js, TypeScript and Azure
A Microsoft reference sample for building a RAG chatbot with LangChain.js, running entirely on Azure serverless (Static Web Apps + Functions) with Cosmos DB for NoSQL as the vector store. Aimed at JS/TS devs who want a working end-to-end RAG pipeline on Azure rather than a bare LangChain snippet.
Covers the whole pipeline, not just the chat endpoint: PDF ingestion, chunking, embedding, vector search, streaming responses with citations and follow-up questions. Ships a genuinely free local dev path via Ollama (llama3.1 + nomic-embed-text) so you can run the full stack without an Azure OpenAI key. Using Cosmos DB's native vector search means one less service to provision instead of bolting on a separate vector DB. Includes CodeTour files that walk through ingestion, vector storage, and retrieval step by step — actual teaching material, not just a README wall of text.
Tightly coupled to azd and Bicep-provisioned Azure resources; swapping Cosmos DB for another vector store means rewriting the retrieval layer, not just changing a connection string. Sample data is three synthetic PDFs (terms of service, privacy policy, support guide) — real-world ingestion problems like scanned pages, tables, or large files aren't exercised. The README admits local Ollama models sometimes fail to follow the citation/formatting instructions correctly, so the free local path isn't functionally equivalent to the Azure OpenAI path. Chat history is per-user but there's no discussion of auth or tenant isolation beyond that — anything past a single-org demo needs that built separately.