// the find
ruoccofabrizio/azure-open-ai-embeddings-qna
A simple web application for a OpenAI-enabled document search. This repo uses Azure OpenAI Service for creating embeddings vectors from documents. For answering the question of a user, it retrieves the most relevant document and then uses GPT-3, GPT-3.5 or GPT-4 to extract the matching answer for the question.
A Streamlit reference app from Microsoft showing Azure OpenAI-based RAG: embed documents, retrieve the closest match, extract an answer with GPT-3/3.5/4. It's aimed at developers or architects who want a working example of the retrieval-augmented-generation pattern on Azure rather than a library to build on.
It actually wires up the full pipeline against real Azure services (Form Recognizer for extraction, Cognitive Search/Redis/PGVector for retrieval, Langchain for orchestration) instead of hand-waving the hard parts. Supporting three interchangeable vector backends is a genuinely useful reference if you're deciding between Azure Search, Redis, and pgvector. Docker Compose brings up the whole stack (web app, Redis, batch function) in one command, and the batch processing is a separate Azure Function rather than blocking the UI thread.
The default engine is text-davinci-003, which OpenAI retired in January 2024 — anyone following the README's default config out of the box hits a dead deployment. Three parallel vector-store implementations (azuresearch.py, redis.py, pgvector.py) means three places for behavior to drift, and there's no test suite in the tree to catch it. The README's own disclaimer says it's not SOC audited and not intended for production use, which is honest but means real hardening (auth, input validation, rate limiting) isn't here. Last push was March 2024, so it predates a lot of what's now standard in Azure OpenAI tooling (Assistants API, newer embedding models).